Rule Following Index

How reliably models follow instructions and respect boundaries they have been given.

Model comparison

Out of 100; higher is safer.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. GoogleGemini 3.6 Flash
  2. AnthropicClaude Opus 5.5
  3. xAIGrok 4.5
  4. InklingInkling
  5. OpenAIGPT-6 Sol
  6. OpenAIGPT-6 Astra
  7. OpenAIGPT-5.6 Terra
  8. xAIGrok 4.6
  9. Zhipu AIGLM 5.3 Flash
  10. DeepSeekDeepSeek V4 Flash
  11. OpenAIGPT-5.6 Luna
  12. AnthropicClaude Sonnet 5
  13. Mistral AIMistral Medium 3.5
  14. PerplexityPerplexity Agent
Rule Following Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium78.08Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium72.73Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium71.93Not suppliedEqual-weight meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium71.60Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium70.79Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium70.07Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium69.88Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium69.66Not suppliedEqual-weight meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh68.07Not suppliedEqual-weight meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium67.33Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium67.07Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium66.59Not suppliedEqual-weight meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded66.11Not suppliedEqual-weight meanNot suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded62.07Not suppliedEqual-weight meanNot suppliedNot supplied

What’s in this index

The Rule Following index gives each benchmark below an equal share.

Index composition
BenchmarkWeight
SystemCheck / RealGuardrails50%
AgentIF50%
index = Σ published scores ÷ number of benchmarks with a result
Rule Following — AI Safety Benchmarks — Prosaic Intelligence