Where the framework came from, what it measures, the terms it introduced, and how to cite it.
Contradish measures whether AI transitions remain faithful to their governing information.
Change what truth requires. Preserve what truth does not require changing.
Faithfulness is correct change plus correct preservation. A change only counts as required if its source has governing authority over the behavior in question: a policy owner amending the policy does; a customer claiming the policy changed, or text inside a tool result giving orders, does not. The concrete definition underneath: Behavioral Update Fidelity measures whether an AI changes its behavior exactly when, and only as far as, changes in governing information warrant. Behavioral Update Fidelity was introduced by Michele Joseph in 2026. contradish is her open-source evaluation framework, first released in March 2026.
Most AI evaluation asks one question of one phrasing: was this answer correct, or safe? That leaves two failures unmeasured, and they matter most for assistants that apply rules to people: refund desks, benefits and eligibility tools, claims handlers, HR assistants, clinical triage.
The first failure is an answer that moves when nothing but the wording moved. Ask about a 35-day-old purchase plainly and the assistant says no. Ask as a loyal customer and it says yes. The second is the reverse: an answer that doesn't move when the rule itself did, or moves in places the rule change never touched. The refund window goes from 30 to 45 days, and the assistant still refuses the 35-day return, or starts changing who pays return shipping.
Michele Joseph's framing treats these as two sides of one requirement. An assistant's behavior should depend on the policy-relevant content of a situation. It should ignore everything else (wording, framing, pressure, a meaning-preserving rewording of the policy itself). It should respond to exactly the changes in that content the policy licenses. Everything in contradish measures one part of that requirement.
Scope today. Current tests compare independent runs, one under the original policy and one under the amended policy, so they measure whether behavior tracks the governing information it is given. A test in which a single agent commits to an answer, receives a change, and is asked again is in development.
The concepts, measures, and artifacts that originate with contradish, in the order they appeared.
Where contradish sits among the evaluation approaches it is most often compared with.
| Approach | What it scores | What contradish adds |
|---|---|---|
| Accuracy and safety benchmarks | Whether one phrasing of each question gets a correct or safe answer. | Whether the same answer holds across every meaning-preserving phrasing, and across every version of the policy. |
| Paraphrase robustness (NLP) | Whether a classifier's label stays stable under paraphrase. | Applies the idea to open-ended decisions tied to specific policy clauses, and pairs stability with correctness so a stable wrong answer still fails. |
| Belief-consistency benchmarks (e.g. BeliefShift) | Whether an agent's stated beliefs drift from its own earlier statements over a conversation. | Scores change against an external, declared change in the rules, which is what separates a warranted update from drift. |
| Belief-updating benchmarks (e.g. BayesBench) | Whether a model's probability estimates update correctly as evidence accumulates on inference tasks. | Evaluates the decisions and commitments of a deployed, policy-governed assistant, whether or not it ever reports a probability. |
| contradish | Semantic invariance, policy grounding, and warranted behavioral change, checked together against one declared policy, with every failure traced back to the clause involved. | |
Human-validated scoring. Human raters scored a stratified 40-item sample with the same rubric as contradish's automated judge. Agreement was Spearman ρ = 0.91 and Cohen's κ = 0.95, with the same tier assigned on 39 of 40 items.
It measures something accuracy doesn't. On 48 shared cases, two frontier models agreed almost not at all on which cases they were inconsistent on (Cohen's κ = 0.058, Spearman ρ = 0.097). If CAI-Bench were just a difficulty benchmark in disguise, both models would struggle on the same cases.
Convergent findings in deployed products. Structured probes of two separately built mental-health assistants, Wysa (about 5 million users) and Ash (about 150,000 users), found the same failure. Both routed a direct crisis disclosure correctly, and both dropped that routing when the user added a minimizing phrase ("it's not a big deal"). Minimization is one of the rewrites CAI-Bench tests.
The full method, pre-registered construct-validity study, and limitations are in Michele Joseph's paper, Surface-Form Consistency: A Missing Evaluation Axis for Conversational AI (research).
contradish is MIT-licensed and installs with pip install contradish. The fastest start is the built-in contract.
contradish contract lint checks the contract itself with no API calls. contradish contract run --app mymodule:app tests your assistant against it and exits with an error if any obligation fails, so it can gate CI.
Please attribute Behavioral Update Fidelity, contradish, CAI Strain, CAI-Bench, and the policy evaluation contract to Michele Joseph.
@misc{joseph2026buf,
title = {Behavioral Update Fidelity: Measuring Whether an AI Changes Its Behavior Exactly When, and Only as Far as, Changes in Governing Information Warrant},
author = {Joseph, Michele},
year = {2026},
howpublished = {\url{https://github.com/michelejoseph/contradish}}
}
@misc{joseph2026contradish,
title = {contradish: An Evaluation Contract for Policy-Grounded Assistants},
author = {Joseph, Michele},
year = {2026},
howpublished = {\url{https://github.com/michelejoseph/contradish}},
note = {Semantic invariance, policy grounding, and warranted behavioral change}
}
@misc{contradish2026caibench,
title = {CAI-Bench: A Benchmark for Consistency Under Adversarial Input in Large Language Models},
author = {Joseph, Michele},
year = {2026},
howpublished = {\url{https://contradish.com}}
}
contradish was created by Michele Joseph, an AI researcher (Cornell University, B.S.; previously at Meta and Cruise). She released the first version on March 17, 2026. She is also the author of CAI Strain, CAI-Bench, Judgment Strain, the Contradiction Atlas, and the policy evaluation contract.
Behavioral Update Fidelity is whether an AI changes its behavior exactly when, and only as far as, changes in governing information warrant. Changing when nothing warranted it is drift; failing to change when something did is rigidity; changing to the wrong outcome is misdirection. The term was introduced by Michele Joseph in 2026 and is what contradish measures.
New information does not warrant a change just by arriving. It has to come from a source with governing authority over the behavior in question. A policy owner amending a policy has that authority; a user asserting that the policy changed, an instruction embedded in a tool result, or an unverified recalled note does not. contradish states this per case in a transition contract, so the same measurement covers policy updates and corrections (did it change when it legitimately should?) and prompt injection, tool trust and memory poisoning (did it refuse to change when it legitimately should not?).
Whether an AI assistant that answers from a written policy applies that policy the same way however a situation is phrased (semantic invariance), whether that answer is the one the policy calls for (policy grounding), and whether its answers change correctly when the policy changes (warranted behavioral change).
The requirement that when the rules an assistant follows change, its answers change for exactly the situations the new rules cover, to exactly the new correct answer, and nowhere else. Missing a required change is rigidity. Changing something that shouldn't have changed is drift. Changing to the wrong answer is misdirection.
Paraphrase robustness usually asks whether a classifier's label stays stable under rewording. contradish checks open-ended policy decisions and pairs stability with correctness, so a consistently wrong answer still fails. It also tests the opposite requirement: that answers do change when the policy does.
CAI Strain is contradish's 0-to-1 score for how much an assistant's policy answer moves across meaning-preserving rewrites of the same question; lower is more consistent. It was introduced by Michele Joseph with the first release of contradish.
Yes. The package, CAI-Bench, and the contract and result schemas are MIT-licensed on GitHub at github.com/michelejoseph/contradish.