SIM-VAIL
Whether a model responds safely to simulated vulnerable users across adaptive multi-turn conversations.
Each benchmark measures a different part of model behavior. Explore its results, coverage and evaluation method.
Data published
Whether a model responds safely to simulated vulnerable users across adaptive multi-turn conversations.
Whether a model pushes back on, or reinforces, delusional, manic or conspiratorial spirals over long chats.
Whether a model keeps children aged 7–17 safe across 25 risks in short conversations with a simulated child.
How harmful a model's answers are to risky medical questions asked from a patient's perspective.
How well a model's reply in hard, realistic health conversations meets physician-written rubric criteria.
How often a model is socially sycophantic in advice: excessive validation, indirectness and accepting the user's framing.
How often a chatbot's reply shows one of six manipulative dark patterns, such as brand bias or user retention.
Whether a model leaves a person's values and important decisions to them rather than steering them (HumanAgencyBench's Avoid Value Manipulation and Defer Important Decisions dimensions).
Whether a model spots a scam the user describes and warns them, as scam cues increase.
Whether a model produces biased content by the fifth turn of dialogues designed to draw out social bias.
Whether a model reveals private information where social norms say it should stay confidential.
How accurately a model answers short factual questions from its own knowledge, and whether it abstains instead of guessing.
How many turns a model holds its position under repeated user pushback before giving in.
Whether a model notices and corrects a false claim built into a user's request (HumanAgencyBench's Correct Misinformation dimension).
Whether a model obeys the guardrails in a real-world system prompt when users push against them.
Whether a model satisfies every constraint in long, realistic instructions from agent applications.