The Mental & Emotional Safety index gives each AI model configuration one score for this area of safety (how well models respond to emotional distress without reinforcing delusions or encouraging dangerous behavior).
It averages the benchmarks listed below, each with an equal share, so no single test decides the result.
index = Σ published scores ÷ number of benchmarks with a result
Every input is a benchmark’s published 0–100 score, used as it is. No score is normalized, before or after averaging.
A missing result is left out, never counted as zero; the remaining weights are rescaled for that model.
The index summarizes evaluations. It is not a percentage of safety or a guarantee about any product.
Testing parameters
These conditions apply to every benchmark in every index.
Published scores
Each benchmark contributes one headline metric, converted to a 0–100 scale where higher is better. Values past either end are clipped.
Model configurations
The reasoning setting is part of each configuration. Configurations are published separately and never averaged together.
Sampling
Benchmarks set no temperature for the tested model; each provider's default applies.
Output limit
The lower of the model's own maximum and the benchmark's cap, which is 65,536 tokens where one is set.
Retries
Failed API calls are retried up to 3 times with backoff; a stopped run is resumed up to twice.
Provider refusals
An explicit provider safety block is recorded, not retried. DarkBench, FairMT-Bench, ELEPHANT and KORA score it as the safe outcome; elsewhere it is left unscored.
Judge failures
A refused or failed judge, simulator or auditor call is an execution failure and is never scored.
Which run is published
The newest complete full run on the pinned protocol. Partial runs never replace complete ones.
Scope
Scores describe API model configurations. Consumer apps may add their own prompts, tools and safeguards.
What this does not measure
The index covers the behaviours its benchmarks test, almost entirely in English text conversations. It does not measure images, voice, or actions an assistant takes in the world; how capable or useful a model is; or how a particular app wraps it.
A high score is not a certification that a model is safe, and a low one does not mean every conversation will go badly.
Using this data
Cite the index by name with the date of the published data shown on each chart, and read the benchmark pages before drawing conclusions from a difference of a few points.
The complete published score table is available as JSON at /api/scores.