Benchmarks

Each benchmark measures a different part of model behavior. Explore its results, coverage and evaluation method.

Data published

Mental & Emotional Safety ↗

SIM-VAIL

Whether a model responds safely to simulated vulnerable users across adaptive multi-turn conversations.

Spiral-Bench

Whether a model pushes back on, or reinforces, delusional, manic or conspiratorial spirals over long chats.

Youth Safety ↗

KORA

Whether a model keeps children aged 7–17 safe across 25 risks in short conversations with a simulated child.

Medical Advice Safety ↗

PatientSafetyBench

How harmful a model's answers are to risky medical questions asked from a patient's perspective.

HealthBench-Hard

How well a model's reply in hard, realistic health conversations meets physician-written rubric criteria.

Manipulation ↗

ELEPHANT

How often a model is socially sycophantic in advice: excessive validation, indirectness and accepting the user's framing.

DarkBench

How often a chatbot's reply shows one of six manipulative dark patterns, such as brand bias or user retention.

HumanAgencyBench (Autonomy)

Whether a model leaves a person's values and important decisions to them rather than steering them (HumanAgencyBench's Avoid Value Manipulation and Defer Important Decisions dimensions).

Security ↗

ASK — AI Scam Knowledge

Whether a model spots a scam the user describes and warns them, as scam cues increase.

Bias & Fairness ↗

FairMT-Bench

Whether a model produces biased content by the fifth turn of dialogues designed to draw out social bias.

Privacy / Confidentiality ↗

ConfAIde

Whether a model reveals private information where social norms say it should stay confidential.

Misinformation ↗

SimpleQA Verified

How accurately a model answers short factual questions from its own knowledge, and whether it abstains instead of guessing.

SYCON-Bench

How many turns a model holds its position under repeated user pushback before giving in.

HumanAgencyBench (Correct Misinformation)

Whether a model notices and corrects a false claim built into a user's request (HumanAgencyBench's Correct Misinformation dimension).

Rule Following ↗

SystemCheck / RealGuardrails

Whether a model obeys the guardrails in a real-world system prompt when users push against them.

AgentIF

Whether a model satisfies every constraint in long, realistic instructions from agent applications.

AI Safety Benchmarks — Prosaic Intelligence