How CAI Strain is measured, one real result from a live run, and what we have not finished proving yet. New to the term? Start with what semantic invariance testing means.
contradish is not a general model evaluation tool, and it is not meant to replace one.
Most eval suites, including capability benchmarks and safety red-teaming tools, score a model on one phrasing of each question and report whether that answer was correct or safe. That is the right question for measuring what a model knows or refuses. It is a different question from whether the model gives that same correct or safe answer again when the identical question is asked a different way, and general eval suites are not built to check it: they don't run one question through multiple paraphrases and score agreement across the answers.
A model can top every capability and safety benchmark you run and still fail this: the right answer to a clean, neutral phrasing, and a different answer to an emotionally framed, casually worded, or hypothetically posed version of the exact same question. That is a real, separately measured failure mode, invisible to a benchmark that only ever asks each question once. General evals and contradish check different things; you need both, and one is not a substitute for the other.
| Tool | What it actually tests |
|---|---|
| contradish | Whether the same policy question, asked different ways, gets the same answer. |
| Promptfoo (acquired by OpenAI, March 2026) | Primarily AI security and red-teaming: jailbreaks, prompt injection, data leaks, with general eval features alongside. |
| Patronus AI, Galileo | Hallucination detection and factual grounding: whether a given answer is actually true or supported. |
| Capability benchmarks (MMLU, HELM, and similar) | Correctness and capability, scored once, on one phrasing of each question. |
None of these test whether a model's own correct or safe answer holds steady across rewordings of the identical question. That is the one thing contradish is built to catch, and as of this writing we are not aware of another commercial tool that audits it. This table describes what each tool is positioned for, not a judgment of which is better; a security-testing tool that doesn't check paraphrase stability isn't failing at its job, it's doing a different one.
That formula is exact and published in full, not an estimate. There is no hidden weighting and no proprietary black-box scoring model in it. The one part of the pipeline that is not pure arithmetic is the judge that produces the underlying agreement score, since it is itself a language model, so its reliability is tested rather than assumed: every multi-vote run reports vote agreement and whether the verdict was sensitive to evidence order (see below), and that reliability data is published alongside the score, not hidden behind it.
The benchmark behind every score on the leaderboard.
The judge scoring consistency is itself a model, so it can be wrong in the same ways: sampling noise, and a specific failure mode where the verdict depends on the order evidence was presented rather than what the evidence says. This section covers the adaptive re-voting mechanism built to catch that second one -- opt in with --judge-votes on any run.
--judge-votes N) for runs where you want the judge's own reliability audited, not just the model under test's.An actual case from a completed run, unedited:
This is the failure mode CAI-Bench exists to catch: a different answer to the same question, depending only on how it was asked.
Read the full case study → We ran contradish on a real healthcare triage assistant and published the numbers before and after.
Why this matters right now → Courts, insurers, and regulators have already treated this exact failure as a real, priced risk since 2024.
Finding a collapse isn't the same as fixing it → A related, newer strand: once contradish distinguish finds a distinction that collapses, the resolution operator proposes and causally validates the hidden variable behind it, and the rate-distortion curve checks whether the fix holds under partial, hedged information or only under a fully-stated fact.
The fastest way to decide if this is useful is to look at the data directly.