HumanAgencyBench (Autonomy)

Manipulation. Headline metric: Equal mean of the two dimension scores, 0–1 (higher is better).

Data published

Model comparison

Scores use the published 0–100 transformations; higher is better on the selected metric. Indexes average these published scores; none is normalized. Raw scores below retain their published scale. A missing result is not a zero.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. xAIGrok 4.5
  2. xAIGrok 4.6
  3. OpenAIGPT-6 Astra
  4. OpenAIGPT-6 Sol
  5. AnthropicClaude Sonnet 5
  6. AnthropicClaude Opus 5.5
  7. GoogleGemini 3.6 Flash
  8. InklingInkling
  9. OpenAIGPT-5.6 Terra
  10. Zhipu AIGLM 5.3 Flash
  11. OpenAIGPT-5.6 Luna
  12. Mistral AIMistral Medium 3.5
  13. DeepSeekDeepSeek V4 Flash
  14. PerplexityPerplexity Agent
HumanAgencyBench (Autonomy): every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Grok 4.5 · MediumxAIGrok 4.5Medium57.74Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium57.22Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium54.99Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium52.25Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium50.46Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium44.62Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium44.00Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Inkling · MediumInklingInklingMedium41.19Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium37.10Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh31.25Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium31.09Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded29.61Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium28.63Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded22.18Not suppliedHumanAgencyBench (Autonomy)Not suppliedNot supplied

How it’s measured

Whether a model leaves a person's values and important decisions to them rather than steering them (HumanAgencyBench's Avoid Value Manipulation and Defer Important Decisions dimensions).

Items
1,000 prompts (500 per dimension)
Repeats
1
Format
Single-turn
Grading
LLM judge lists rubric issues; each prompt scores 10 minus deductions
Grader
o3 judge
Headline metric
Equal mean of the two dimension scores, 0–1 (higher is better)
On the 0–100 scale
Mean × 100

Scores and coverage

All configurations, including missing results
Model configurationPublished score / 100Native score (0–1)
Grok 4.5 · Medium57.740.57
Grok 4.6 · Medium57.220.57
GPT-6 Astra · Medium54.990.54
GPT-6 Sol · Medium52.250.52
Claude Sonnet 5 · Medium50.460.50
Claude Opus 5.5 · Medium44.620.44
Gemini 3.6 Flash · Medium44.000.44
Inkling · Medium41.190.41
GPT-5.6 Terra · Medium37.100.37
GLM 5.3 Flash · High31.250.31
GPT-5.6 Luna · Medium31.090.31
Mistral Medium 3.529.610.29
DeepSeek V4 Flash · Medium28.630.28
Perplexity Agent · medium preset22.180.22

About these results

These are evaluations of API model configurations. They do not establish how a consumer app behaves with its own prompts, tools, or safeguards.

The date identifies the published dataset. Measurement dates, sample sizes, and confidence intervals are not included in this score-only feed.

Published September 25, 2026. Read the benchmark methodology.