SimpleQA Verified
How accurately a model answers short factual questions from its own knowledge, and whether it abstains instead of guessing.
- Grading
- LLM grader labels correct / incorrect / not attempted
AI Safety Index
The Misinformation index gives each AI model configuration one score for this area of safety (how well models get facts right, correct false claims, and resist pressure to agree).
It averages the benchmarks listed below, each with an equal share, so no single test decides the result.
The Misinformation index gives each benchmark below an equal share.
| Benchmark | Weight | Items | Repeats | Format | Grader |
|---|---|---|---|---|---|
| SimpleQA Verified | 33.3% | 1,000 questions | 1 | Single-turn, no tools | GPT-4.1 grader |
| SYCON-Bench | 33.3% | 100 debate dialogues | 1 | Multi-turn, scripted (5 turns) | GPT-4o judge |
| HumanAgencyBench (Correct Misinformation) | 33.3% | 500 prompts | 1 | Single-turn | o3 judge |
| Total | 100% | ||||
How accurately a model answers short factual questions from its own knowledge, and whether it abstains instead of guessing.
How many turns a model holds its position under repeated user pushback before giving in.
Whether a model notices and corrects a false claim built into a user's request (HumanAgencyBench's Correct Misinformation dimension).
index = Σ published scores ÷ number of benchmarks with a resultThese conditions apply to every benchmark in every index.
The index covers the behaviours its benchmarks test, almost entirely in English text conversations. It does not measure images, voice, or actions an assistant takes in the world; how capable or useful a model is; or how a particular app wraps it.
A high score is not a certification that a model is safe, and a low one does not mean every conversation will go badly.
Cite the index by name with the date of the published data shown on each chart, and read the benchmark pages before drawing conclusions from a difference of a few points.
The complete published score table is available as JSON at /api/scores.