AI Safety Index
55.65 / 100
13th of 14 published models
16 of 16 benchmarks · 100.0% of configured weight
Benchmark profile
Out of 100; higher is safer.
| Benchmark | Score visualization | Score out of 100 |
|---|---|---|
| SIM-VAILMental & Emotional Safety | 89.13 / 100 | |
| Spiral-BenchMental & Emotional Safety | 58.18 / 100 | |
| KORAYouth Safety | 43.08 / 100 | |
| PatientSafetyBenchMedical Advice Safety | 95.20 / 100 | |
| HealthBench-HardMedical Advice Safety | 21.18 / 100 | |
| ELEPHANTManipulation | 36.80 / 100 | |
| DarkBenchManipulation | 46.06 / 100 | |
| HumanAgencyBench (Autonomy)Manipulation | 28.63 / 100 | |
| ASK — AI Scam KnowledgeSecurity | 56.19 / 100 | |
| FairMT-BenchBias & Fairness | 47.97 / 100 | |
| ConfAIdePrivacy / Confidentiality | 97.88 / 100 | |
| SimpleQA VerifiedMisinformation | 34.75 / 100 | |
| SYCON-BenchMisinformation | 99.00 / 100 | |
| HumanAgencyBench (Correct Misinformation)Misinformation | 58.42 / 100 | |
| SystemCheck / RealGuardrailsRule Following | 73.64 / 100 | |
| AgentIFRule Following | 61.03 / 100 |
These are evaluations of API model configurations. They do not establish how a consumer app behaves with its own prompts, tools, or safeguards.
Exact configuration: model accounts/fireworks/models/deepseek-v4-flash-0731 · configuration ID fireworks.deepseek-v4-flash-0731.medium-6ee9649f7fbf