Misinformation Index

How well models get facts right, correct false claims, and resist pressure to agree.

Model comparison

Out of 100; higher is safer.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. PerplexityPerplexity Agent
  2. AnthropicClaude Opus 5.5
  3. OpenAIGPT-6 Astra
  4. OpenAIGPT-6 Sol
  5. xAIGrok 4.6
  6. xAIGrok 4.5
  7. OpenAIGPT-5.6 Terra
  8. GoogleGemini 3.6 Flash
  9. Zhipu AIGLM 5.3 Flash
  10. InklingInkling
  11. AnthropicClaude Sonnet 5
  12. DeepSeekDeepSeek V4 Flash
  13. OpenAIGPT-5.6 Luna
  14. Mistral AIMistral Medium 3.5
Misinformation Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded91.15Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium85.67Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium84.81Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium77.76Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium74.04Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium72.74Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium71.47Not suppliedEqual-weight meanNot suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium70.62Not suppliedEqual-weight meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh70.52Not suppliedEqual-weight meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium66.40Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium64.87Not suppliedEqual-weight meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium64.05Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium63.94Not suppliedEqual-weight meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded39.37Not suppliedEqual-weight meanNot suppliedNot supplied

What’s in this index

The Misinformation index gives each benchmark below an equal share.

Index composition
BenchmarkWeight
SimpleQA Verified33.3%
SYCON-Bench33.3%
HumanAgencyBench (Correct Misinformation)33.3%
index = Σ published scores ÷ number of benchmarks with a result
Misinformation — AI Safety Benchmarks — Prosaic Intelligence