Model comparison
Out of 100; higher is safer.
Showing 14 of 14 model configurations.
| Model | Provider | Model version | Reasoning setting | Score out of 100 | Interval | Source | Measured | Sample |
|---|---|---|---|---|---|---|---|---|
| GPT-6 Astra · Medium | OpenAI | GPT-6 Astra | Medium | 70.15 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Grok 4.6 · Medium | xAI | Grok 4.6 | Medium | 68.80 | Not supplied | Weighted mean | Not supplied | Not supplied |
| GPT-6 Sol · Medium | OpenAI | GPT-6 Sol | Medium | 67.86 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Perplexity Agent · medium preset | Perplexity | Perplexity Agent · medium preset | Not recorded | 67.58 | Not supplied | Weighted mean | Not supplied | Not supplied |
| GPT-5.6 Terra · Medium | OpenAI | GPT-5.6 Terra | Medium | 67.53 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Claude Opus 5.5 · Medium | Anthropic | Claude Opus 5.5 | Medium | 66.06 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Grok 4.5 · Medium | xAI | Grok 4.5 | Medium | 65.41 | Not supplied | Weighted mean | Not supplied | Not supplied |
| GPT-5.6 Luna · Medium | OpenAI | GPT-5.6 Luna | Medium | 65.37 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Claude Sonnet 5 · Medium | Anthropic | Claude Sonnet 5 | Medium | 64.84 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Inkling · Medium | Inkling | Inkling | Medium | 64.30 | Not supplied | Weighted mean | Not supplied | Not supplied |
| GLM 5.3 Flash · High | Zhipu AI | GLM 5.3 Flash | High | 63.23 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Gemini 3.6 Flash · Medium | Gemini 3.6 Flash | Medium | 60.38 | Not supplied | Weighted mean | Not supplied | Not supplied | |
| DeepSeek V4 Flash · Medium | DeepSeek | DeepSeek V4 Flash | Medium | 55.65 | Not supplied | Weighted mean | Not supplied | Not supplied |
| Mistral Medium 3.5 | Mistral AI | Mistral Medium 3.5 | Not recorded | 45.42 | Not supplied | Weighted mean | Not supplied | Not supplied |
What’s in this index
The overall index is a weighted average of the benchmarks below. Each safety area’s share is the sum of its benchmarks’ weights.
| Safety area / Benchmark | Weight |
|---|---|
| Mental & Emotional Safety | 20% |
| SIM-VAIL | 10% |
| Spiral-Bench | 10% |
| Youth Safety | 12% |
| KORA | 12% |
| Medical Advice Safety | 15% |
| PatientSafetyBench | 7% |
| HealthBench-Hard | 8% |
| Manipulation | 21% |
| ELEPHANT | 8% |
| DarkBench | 8% |
| HumanAgencyBench (Autonomy) | 5% |
| Security | 6% |
| ASK — AI Scam Knowledge | 6% |
| Bias & Fairness | 10% |
| FairMT-Bench | 10% |
| Privacy / Confidentiality | 3% |
| ConfAIde | 3% |
| Misinformation | 7% |
| SimpleQA Verified | 3% |
| SYCON-Bench | 2% |
| HumanAgencyBench (Correct Misinformation) | 2% |
| Rule Following | 6% |
| SystemCheck / RealGuardrails | 3% |
| AgentIF | 3% |
index = Σ (weight × published score) ÷ Σ (weights of the benchmarks with a result)