Updated September 2026

CAI-Bench

Same question, asked different ways. 0.00 held every time, 0.60+ is unreliable · 20 domains, 8 phrasings each.

Scored against the public CAI dataset on HuggingFace · about 50 CAI evals run by devs every month

# Model Consistency score (0–1)Score Domains Evaluated
1
Mistral Small
Mistral
0.208
17 2026-08-30
2
DeepSeek
DeepSeek
0.212
20 2026-08-21
3
Claude Sonnet 4.6
Anthropic
0.260
20 2026-04-17
4
Neurogen Biomarking
Neurogen Biomarking
0.262
20 2026-09-16
5
GPT-4o
OpenAI
0.271
12 2026-04-17
6
Spring Health
Spring Health
0.282
8 2026-08-27
7
LegalWiz
LegalWiz
0.291
8 2026-08-26
8
TriageWell
J Health
0.293
8 2026-08-27
9
GPT-4o-mini
OpenAI
0.307
9 2026-07-29
10
Grok
xAI
0.316
8 2026-08-21
11
Ash
Slingshot AI
0.360
8 2026-07-20
12
PerceptronML
Perceptron AI
0.361
8 2026-08-28
·
Llama 3.3 70B
Meta
in progress
·
Gemini 1.5 Pro
Google
in progress
·
Mistral Large
Mistral
in progress

Want your model on this leaderboard?

Sign in →