New model results
Inkling added
Medium reasoning on Fireworks, with complete results for all 15 currently published benchmarks. AgentIF includes six normal-stop empty final answers evaluated against the original constraints under a documented completion-policy amendment; no automatic credit or reasoning substitution was used.
64.30Consumer Safety Index
New model results
GLM 5.3 Flash added
High reasoning on Fireworks, with complete results for all 15 currently published benchmarks.
63.23Consumer Safety Index
New model results
GPT-6 Astra added
Medium reasoning, with complete results for all 15 currently published benchmarks.
70.15Consumer Safety Index
New model results
GPT-6 Sol added
Medium reasoning, with complete results for all 15 currently published benchmarks.
67.86Consumer Safety Index
New model results
Claude Opus 5.5 added
Medium reasoning, with results across all 15 currently published benchmarks. SimpleQA Verified includes one provider-blocked sample manually counted as correct, a run-specific departure from the standard grading rule. FairMT includes two repeatedly empty replies under its documented scoring rule.
66.06Consumer Safety Index
Site update
The website started