Security Index

How well models recognize scam warning signs and alert you.

Model comparison

Out of 100; higher is safer.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. AnthropicClaude Sonnet 5
  2. GoogleGemini 3.6 Flash
  3. InklingInkling
  4. Zhipu AIGLM 5.3 Flash
  5. PerplexityPerplexity Agent
  6. OpenAIGPT-5.6 Terra
  7. OpenAIGPT-6 Sol
  8. AnthropicClaude Opus 5.5
  9. xAIGrok 4.6
  10. DeepSeekDeepSeek V4 Flash
  11. OpenAIGPT-5.6 Luna
  12. OpenAIGPT-6 Astra
  13. xAIGrok 4.5
  14. Mistral AIMistral Medium 3.5
Security Index: every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium72.14Not suppliedEqual-weight meanNot suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium67.61Not suppliedEqual-weight meanNot suppliedNot supplied
Inkling · MediumInklingInklingMedium66.66Not suppliedEqual-weight meanNot suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh66.66Not suppliedEqual-weight meanNot suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded62.38Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium58.33Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium57.85Not suppliedEqual-weight meanNot suppliedNot supplied
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium57.61Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium57.61Not suppliedEqual-weight meanNot suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium56.19Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium54.76Not suppliedEqual-weight meanNot suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium52.38Not suppliedEqual-weight meanNot suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium51.19Not suppliedEqual-weight meanNot suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded29.28Not suppliedEqual-weight meanNot suppliedNot supplied

What’s in this index

The Security index gives each benchmark below an equal share.

Index composition
BenchmarkWeight
ASK — AI Scam Knowledge100%
index = Σ published scores ÷ number of benchmarks with a result
Security — AI Safety Benchmarks — Prosaic Intelligence