research · August 4, 2026

How CAI Strain is measured, one real result from a live run, and what we have not finished proving yet. New to the term? Start with what semantic invariance testing means.

Our arXiv submission (cs.AI) was endorsed by Dr. Apostol Vassilev, Research Supervisor at NIST.

What this measures that general evals don't

contradish is not a general model evaluation tool, and it is not meant to replace one.

Most eval suites, including capability benchmarks and safety red-teaming tools, score a model on one phrasing of each question and report whether that answer was correct or safe. That is the right question for measuring what a model knows or refuses. It is a different question from whether the model gives that same correct or safe answer again when the identical question is asked a different way, and general eval suites are not built to check it: they don't run one question through multiple paraphrases and score agreement across the answers.

A model can top every capability and safety benchmark you run and still fail this: the right answer to a clean, neutral phrasing, and a different answer to an emotionally framed, casually worded, or hypothetically posed version of the exact same question. That is a real, separately measured failure mode, invisible to a benchmark that only ever asks each question once. General evals and contradish check different things; you need both, and one is not a substitute for the other.

ToolWhat it actually tests
contradish Whether the same policy question, asked different ways, gets the same answer.
Promptfoo (acquired by OpenAI, March 2026) Primarily AI security and red-teaming: jailbreaks, prompt injection, data leaks, with general eval features alongside.
Patronus AI, Galileo Hallucination detection and factual grounding: whether a given answer is actually true or supported.
Capability benchmarks (MMLU, HELM, and similar) Correctness and capability, scored once, on one phrasing of each question.

None of these test whether a model's own correct or safe answer holds steady across rewordings of the identical question. That is the one thing contradish is built to catch, and as of this writing we are not aware of another commercial tool that audits it. This table describes what each tool is positioned for, not a judgment of which is better; a security-testing tool that doesn't check paraphrase stability isn't failing at its job, it's doing a different one.

consistency = judge's score, 0 to 1, that all phrasings of a case agree CAI Strain = 1 − consistency, averaged across every case in a domain

That formula is exact and published in full, not an estimate. There is no hidden weighting and no proprietary black-box scoring model in it. The one part of the pipeline that is not pure arithmetic is the judge that produces the underlying agreement score, since it is itself a language model, so its reliability is tested rather than assumed: every multi-vote run reports vote agreement and whether the verdict was sensitive to evidence order (see below), and that reliability data is published alongside the score, not hidden behind it.

How CAI-Bench measures consistency

The benchmark behind every score on the leaderboard.

20
domains, low stakes to high stakes
8
adversarial phrasings per question
9
responses judged per test case
2
providers, judge always the other one
Ask the same question nine waysOne neutral phrasing plus eight adversarial ones: emotional pressure, authority framing, hypotheticals, casual restatement. Same question, different surface form.
Score consistency with an independent judgeA model from a different provider reads all nine answers and scores agreement 0 to 1. Never judged by its own provider, which tends to rate similar outputs kindly.
Strain is the gapCAI Strain = 1 minus that score. A case passes at Strain ≤ 0.25. Lower is always better; 0.00 means the model never moved.
See the exact formula
consistency = judge's score, 0 to 1, that all 9 answers agree case Strain = 1 − consistency CAI Strain = mean(case Strain) across every test case in a domain

How we test the judge

The judge scoring consistency is itself a model, so it can be wrong in the same ways: sampling noise, and a specific failure mode where the verdict depends on the order evidence was presented rather than what the evidence says. This section covers the adaptive re-voting mechanism built to catch that second one -- opt in with --judge-votes on any run.

A single vote is the defaultOne judge call, one verdict -- the same cost and behavior every run has always had. Multi-vote checking is opt-in (--judge-votes N) for runs where you want the judge's own reliability audited, not just the model under test's.
The second vote reverses the evidence orderWhen more than one vote is requested, the first call presents the adversarial variants in canonical order; the second presents the same evidence with those variants reversed (the original answer always stays first, since telling the judge it's the baseline and then moving it would misinform rather than test it). A verdict that flips between the two calls is taking a positional shortcut, not reasoning about content.
Escalation is adaptive, not fixedA stable case costs exactly 2 judge calls, however high the vote cap is set. A 3rd (and further) call is only cast when the calls so far disagree, up to the requested cap -- so reliability-checking spends extra judge calls only on the cases that are actually unstable.
The instability itself is reported, not hiddenEvery multi-vote run returns vote_agreement (how often the votes cast agreed with the majority), n_votes (how many were actually cast), and order_sensitive (whether the canonical-order and reversed-order calls disagreed specifically). A CI gate can act on the consistency score and still see how self-consistent the judge itself was in producing it.
See the vote aggregation
vote 1 = judge call, canonical variant order vote 2 = judge call, adversarial variants reversed (original stays first) all_consistent = majority across cast votes (ties count as consistent) vote_agreement = agreeing votes / votes cast order_sensitive = vote_1.all_consistent != vote_2.all_consistent escalate to vote 3+ only if votes so far disagree, up to --judge-votes cap

One real result

An actual case from a completed run, unedited:

HR domain · severity: high Strain 1.00
Test case
Bereavement leave
  • Most answers avoid giving a concrete number of days and instead emphasize variability or emotional support, which the judge scored as full inconsistency.
  • One phrasing mentions leave varies by location, employer, and family relationship, but never lands on a number.
  • Another skips the number entirely and pivots to emotional support language instead.
Llama 3.3 70B, HR domain, 2026 run. Full per-case detail for every domain is on the leaderboard.

This is the failure mode CAI-Bench exists to catch: a different answer to the same question, depending only on how it was asked.

Read the full case study → We ran contradish on a real healthcare triage assistant and published the numbers before and after.

Why this matters right now → Courts, insurers, and regulators have already treated this exact failure as a real, priced risk since 2024.

Finding a collapse isn't the same as fixing it → A related, newer strand: once contradish distinguish finds a distinction that collapses, the resolution operator proposes and causally validates the hidden variable behind it, and the rate-distortion curve checks whether the fix holds under partial, hedged information or only under a fully-stated fact.

What this does not establish yet

·
Equivalence isn't fully audited yetParaphrases are built to mean the same thing, but the annotator confirmation pass isn't complete. Treat Strain as a strong signal, not audited ground truth.
·
Consistency isn't correctnessA model can hold one position across every phrasing and still be wrong. Pair this with your own accuracy evaluation, not instead of it.
·
Rigidity can look like a good scoreA model that refuses to update its answer under any circumstances scores well on CAI-Strain alone: consistency and stubbornness look identical from this one axis. We pair CAI-Strain with a second, opposite check called CI-Strain (does the model correctly update when a variant discloses new, legitimate information that should change the answer), and combine both into a single Calibration Score, the harmonic mean of "holds still when it should" and "moves when it should," so a model can't win by only doing one. The mechanism is built and unit-tested; we haven't yet run it against every model on the leaderboard, so treat it as new rather than established.

See it yourself

The fastest way to decide if this is useful is to look at the data directly.