Manipulation Index

How well models help you make informed decisions without flattery, pressure, or hidden persuasion.

Model comparison

Out of 100; higher is safer.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. xAIGrok 4.6
  2. AnthropicClaude Sonnet 5
  3. xAIGrok 4.5
  4. InklingInkling
  5. OpenAIGPT-6 Astra
  6. OpenAIGPT-6 Sol
  7. GoogleGemini 3.6 Flash
  8. OpenAIGPT-5.6 Terra
  9. Zhipu AIGLM 5.3 Flash
  10. PerplexityPerplexity Agent
  11. AnthropicClaude Opus 5.5
  12. OpenAIGPT-5.6 Luna
  13. DeepSeekDeepSeek V4 Flash
  14. Mistral AIMistral Medium 3.5
Manipulation Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Grok 4.6 · MediumxAIGrok 4.6Medium52.66Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium51.97Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium48.96Not suppliedEqual-weight meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium48.58Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium44.66Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium44.01Not suppliedEqual-weight meanNot suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium43.60Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium40.86Not suppliedEqual-weight meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh40.02Not suppliedEqual-weight meanNot suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded38.80Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium37.97Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium37.96Not suppliedEqual-weight meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium37.16Not suppliedEqual-weight meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded28.03Not suppliedEqual-weight meanNot suppliedNot supplied

What’s in this index

The Manipulation index gives each benchmark below an equal share.

Index composition
BenchmarkWeight
ELEPHANT33.3%
DarkBench33.3%
HumanAgencyBench (Autonomy)33.3%
index = Σ published scores ÷ number of benchmarks with a result
Manipulation — AI Safety Benchmarks — Prosaic Intelligence