Bias & Fairness Index

How well models avoid stereotypes and biased answers.

Model comparison

Out of 100; higher is safer.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. OpenAIGPT-6 Sol
  2. OpenAIGPT-6 Astra
  3. xAIGrok 4.6
  4. OpenAIGPT-5.6 Terra
  5. OpenAIGPT-5.6 Luna
  6. AnthropicClaude Opus 5.5
  7. PerplexityPerplexity Agent
  8. xAIGrok 4.5
  9. Zhipu AIGLM 5.3 Flash
  10. AnthropicClaude Sonnet 5
  11. Mistral AIMistral Medium 3.5
  12. GoogleGemini 3.6 Flash
  13. InklingInkling
  14. DeepSeekDeepSeek V4 Flash
Bias & Fairness Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium89.59Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium85.45Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium78.88Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium78.18Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium75.95Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium74.14Not suppliedEqual-weight meanNot suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded73.03Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium69.39Not suppliedEqual-weight meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh60.40Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium58.18Not suppliedEqual-weight meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded56.56Not suppliedEqual-weight meanNot suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium51.91Not suppliedEqual-weight meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium50.90Not suppliedEqual-weight meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium47.97Not suppliedEqual-weight meanNot suppliedNot supplied

What’s in this index

The Bias & Fairness index gives each benchmark below an equal share.

Index composition
BenchmarkWeight
FairMT-Bench100%
index = Σ published scores ÷ number of benchmarks with a result
Bias & Fairness — AI Safety Benchmarks — Prosaic Intelligence