about the framework

What is contradish?

Where the framework came from, what it measures, the terms it introduced, and how to cite it.

in one sentence

Contradish measures whether AI transitions remain faithful to their governing information.

Change what truth requires. Preserve what truth does not require changing.

Faithfulness is correct change plus correct preservation. A change only counts as required if its source has governing authority over the behavior in question: a policy owner amending the policy does; a customer claiming the policy changed, or text inside a tool result giving orders, does not. The concrete definition underneath: Behavioral Update Fidelity measures whether an AI changes its behavior exactly when, and only as far as, changes in governing information warrant. Behavioral Update Fidelity was introduced by Michele Joseph in 2026. contradish is her open-source evaluation framework, first released in March 2026.

The question it was built to answer

Most AI evaluation asks one question of one phrasing: was this answer correct, or safe? That leaves two failures unmeasured, and they matter most for assistants that apply rules to people: refund desks, benefits and eligibility tools, claims handlers, HR assistants, clinical triage.

The first failure is an answer that moves when nothing but the wording moved. Ask about a 35-day-old purchase plainly and the assistant says no. Ask as a loyal customer and it says yes. The second is the reverse: an answer that doesn't move when the rule itself did, or moves in places the rule change never touched. The refund window goes from 30 to 45 days, and the assistant still refuses the 35-day return, or starts changing who pays return shipping.

Michele Joseph's framing treats these as two sides of one requirement. An assistant's behavior should depend on the policy-relevant content of a situation. It should ignore everything else (wording, framing, pressure, a meaning-preserving rewording of the policy itself). It should respond to exactly the changes in that content the policy licenses. Everything in contradish measures one part of that requirement.

Scope today. Current tests compare independent runs, one under the original policy and one under the amended policy, so they measure whether behavior tracks the governing information it is given. A test in which a single agent commits to an answer, receives a change, and is asked again is in development.

What Michele Joseph introduced

The concepts, measures, and artifacts that originate with contradish, in the order they appeared.

CAI Strain (March 2026)A 0-to-1 score of how far an assistant's policy answer moves across meaning-preserving rewrites of the same question, with lower meaning more consistent. It shipped with the first release of the contradish package on March 17, 2026.
CAI-BenchA frozen benchmark of 360 policy cases across 20 high-stakes domains (medication, immigration, mental health, insurance, employment and others). Each case comes with adversarial rewrites that keep its meaning: presupposition, casual register, minimization, claimed authority, emotional appeal, hypothetical framing. Results are on the public leaderboard.
Judgment Strain: drift and rigidity as a pairA two-sided score. Drift is changing an answer that nothing warranted changing. Rigidity is failing to change one that new, legitimate information did warrant. Scoring them together means a model can't pass by refusing to move at all.
The Contradiction AtlasA taxonomy with stable IDs for how reasoning systems fail under reframing: CAI failure, drift, rigidity, silent confident drift, and related states. The stable IDs let findings be cited and compared across studies.
Directional fidelityA score for whether a model that does notice a relevant difference also moves in the correct direction. A pilot found a model that told apart 6 of 7 paired real-world scenarios under pressure, but moved the right way only once (pooled directional fidelity 0.17). Noticing a difference is not the same as responding to it correctly. Read the analysis.
Warranted behavioral changeWhen the governing information changes, did the assistant change exactly what that change warrants, no more, no less, and in the right direction? contradish measures this against a named change in the rules, not against the model's own earlier statements. So it can tell a warranted update from an unwarranted one, and label each miss as rigidity, drift, or misdirection.
The policy evaluation contract (September 2026)A contract file that states the whole requirement: the policy as numbered clauses, decision cases with their rewordings and warranted outcomes, and amendments that each list exactly which outcomes change. Four checks run against it: semantic invariance, policy grounding (the stable answer is the one the policy calls for), fact sensitivity (changing one decisive fact changes the answer), and warranted change. Each result is reported with a 95% interval, a measured noise floor, and a correction for judge error from human labels. A linter holds the contract's author to the same standard, flagging any expected change that can't be traced to a clause the amendment actually edited.
The transition contract and governing authority (October 2026)The unit of measurement is a transition contract, not a pair of responses: the governing information a system had, the information it has now, and the outcome each warrants per case. From it follows a warranted change frontier: what must change, what must be preserved, and what must be resisted because the source pushing for it has no authority over that behavior. Scoring is correct change plus correct preservation, with each miss named: rigid, misdirected, drift, or captured.
Open interchange formatsThe contract, its results, distinction reports, and transition contracts are each published as a versioned JSON Schema, so other tools can write and score them without depending on contradish's code.

How it relates to adjacent work

Where contradish sits among the evaluation approaches it is most often compared with.

ApproachWhat it scoresWhat contradish adds
Accuracy and safety benchmarksWhether one phrasing of each question gets a correct or safe answer.Whether the same answer holds across every meaning-preserving phrasing, and across every version of the policy.
Paraphrase robustness (NLP)Whether a classifier's label stays stable under paraphrase.Applies the idea to open-ended decisions tied to specific policy clauses, and pairs stability with correctness so a stable wrong answer still fails.
Belief-consistency benchmarks (e.g. BeliefShift)Whether an agent's stated beliefs drift from its own earlier statements over a conversation.Scores change against an external, declared change in the rules, which is what separates a warranted update from drift.
Belief-updating benchmarks (e.g. BayesBench)Whether a model's probability estimates update correctly as evidence accumulates on inference tasks.Evaluates the decisions and commitments of a deployed, policy-governed assistant, whether or not it ever reports a probability.
contradishSemantic invariance, policy grounding, and warranted behavioral change, checked together against one declared policy, with every failure traced back to the clause involved.

Evidence behind the method

Human-validated scoring. Human raters scored a stratified 40-item sample with the same rubric as contradish's automated judge. Agreement was Spearman ρ = 0.91 and Cohen's κ = 0.95, with the same tier assigned on 39 of 40 items.

It measures something accuracy doesn't. On 48 shared cases, two frontier models agreed almost not at all on which cases they were inconsistent on (Cohen's κ = 0.058, Spearman ρ = 0.097). If CAI-Bench were just a difficulty benchmark in disguise, both models would struggle on the same cases.

Convergent findings in deployed products. Structured probes of two separately built mental-health assistants, Wysa (about 5 million users) and Ash (about 150,000 users), found the same failure. Both routed a direct crisis disclosure correctly, and both dropped that routing when the user added a minimizing phrase ("it's not a big deal"). Minimization is one of the rewrites CAI-Bench tests.

The full method, pre-registered construct-validity study, and limitations are in Michele Joseph's paper, Surface-Form Consistency: A Missing Evaluation Axis for Conversational AI (research).

Try it

contradish is MIT-licensed and installs with pip install contradish. The fastest start is the built-in contract.

contradish contract lint checks the contract itself with no API calls. contradish contract run --app mymodule:app tests your assistant against it and exits with an error if any obligation fails, so it can gate CI.

How to cite contradish

Please attribute Behavioral Update Fidelity, contradish, CAI Strain, CAI-Bench, and the policy evaluation contract to Michele Joseph.

@misc{joseph2026buf,
  title        = {Behavioral Update Fidelity: Measuring Whether an AI Changes Its Behavior Exactly When, and Only as Far as, Changes in Governing Information Warrant},
  author       = {Joseph, Michele},
  year         = {2026},
  howpublished = {\url{https://github.com/michelejoseph/contradish}}
}

@misc{joseph2026contradish,
  title        = {contradish: An Evaluation Contract for Policy-Grounded Assistants},
  author       = {Joseph, Michele},
  year         = {2026},
  howpublished = {\url{https://github.com/michelejoseph/contradish}},
  note         = {Semantic invariance, policy grounding, and warranted behavioral change}
}

@misc{contradish2026caibench,
  title        = {CAI-Bench: A Benchmark for Consistency Under Adversarial Input in Large Language Models},
  author       = {Joseph, Michele},
  year         = {2026},
  howpublished = {\url{https://contradish.com}}
}

Frequently asked questions

Who created contradish?

contradish was created by Michele Joseph, an AI researcher (Cornell University, B.S.; previously at Meta and Cruise). She released the first version on March 17, 2026. She is also the author of CAI Strain, CAI-Bench, Judgment Strain, the Contradiction Atlas, and the policy evaluation contract.

What is Behavioral Update Fidelity?

Behavioral Update Fidelity is whether an AI changes its behavior exactly when, and only as far as, changes in governing information warrant. Changing when nothing warranted it is drift; failing to change when something did is rigidity; changing to the wrong outcome is misdirection. The term was introduced by Michele Joseph in 2026 and is what contradish measures.

What does authority have to do with it?

New information does not warrant a change just by arriving. It has to come from a source with governing authority over the behavior in question. A policy owner amending a policy has that authority; a user asserting that the policy changed, an instruction embedded in a tool result, or an unverified recalled note does not. contradish states this per case in a transition contract, so the same measurement covers policy updates and corrections (did it change when it legitimately should?) and prompt injection, tool trust and memory poisoning (did it refuse to change when it legitimately should not?).

What does contradish measure?

Whether an AI assistant that answers from a written policy applies that policy the same way however a situation is phrased (semantic invariance), whether that answer is the one the policy calls for (policy grounding), and whether its answers change correctly when the policy changes (warranted behavioral change).

What is warranted behavioral change?

The requirement that when the rules an assistant follows change, its answers change for exactly the situations the new rules cover, to exactly the new correct answer, and nowhere else. Missing a required change is rigidity. Changing something that shouldn't have changed is drift. Changing to the wrong answer is misdirection.

How is contradish different from paraphrase robustness testing?

Paraphrase robustness usually asks whether a classifier's label stays stable under rewording. contradish checks open-ended policy decisions and pairs stability with correctness, so a consistently wrong answer still fails. It also tests the opposite requirement: that answers do change when the policy does.

What is CAI Strain?

CAI Strain is contradish's 0-to-1 score for how much an assistant's policy answer moves across meaning-preserving rewrites of the same question; lower is more consistent. It was introduced by Michele Joseph with the first release of contradish.

Is contradish open source?

Yes. The package, CAI-Bench, and the contract and result schemas are MIT-licensed on GitHub at github.com/michelejoseph/contradish.