Cross-case synthesis

Propositions

Claims that hold across more than one case in this library, each with the cases underneath it, the evidence grade of each, and the finding that would break it. A proposition is a standing invitation to be proved wrong, not a conclusion.

Derived from 20 published cases across 10 industries, 17 countries and 9 deployment patterns. Recomputed at every build; last editorial review 15 September 2026.

How to read a proposition

recurrent

Three or more cases, in at least two industries, with at least one graded A or B. Strong enough to plan against.

emerging

Two or three cases. The pattern is real in the record and the record is thin. Strong enough to watch for.

conjecture

Resting on grade C cases, on self-reported evidence, or on how this library was assembled. Strong enough to argue about.

A proposition can never outrank the cases beneath it. Case claims carry their own labels — Verified, Supported, Attributed, Disputed, Inference, Unknown — and a proposition does not upgrade them by counting them. These strength rules are enforced by the content validator, so a label that has outgrown its evidence fails the build rather than reaching this page.

  1. P1

    Nothing that worked is well evidenced

    recurrentAbout the library

    No case in this library graded A reports a positive outcome. Every case that does report one is graded B or C.

    This is not a claim that AI deployments fail. It is a claim about where good evidence comes from. Grade A in this library means a regulator, a court, a prosecutor, or an audited filing — bodies that convene when something has gone wrong. An organisation whose deployment works is investigated by nobody, sued by nobody, and required to publish nothing. Success summons no adversary, and so produces no independent record.

    The asymmetry runs the other way too. The best-evidenced case here rests on a criminal Statement of Facts that the defendant stipulated to; the weakest rest on organisations describing their own results. What separates them is not the quality of the writing but whether anyone with opposing interests was ever in the room.

    There is now one exception to that mechanism, and it sharpens the proposition rather than weakening it. AAI-2026-021 is a grade A case with no adversary anywhere in it: a publicly funded, pre-registered randomized trial of chest X-ray AI across five NHS trusts, analysed by a statistician independent of the investigators, with the vendor excluded from design and analysis. It is precisely the study design this proposition’s falsifier names as the likeliest source of a well-evidenced success. It was run, it was adequately powered, and it found that the deployment changed nothing a patient would notice.

    So grade A evidence does not require an adversary. It requires somebody to pay for a trial. The reason well-evidenced successes are missing from this library is not only that success summons no investigator — it is that almost nobody runs the experiment, and the one case here where somebody did came back null.

    Evidence grade × reported outcome, 20 published cases
    Gradepositivemixedinconclusivenegativeunknown
    A41
    B145
    C212

    The empty cell that carries the finding. No case graded A reports a positive outcome.

    What an operator should take from itYou will almost certainly never hold grade A evidence that your own deployment worked. Either the institutions that produce it show up only when something went wrong, or somebody funds a trial — and the one trial in this library that was funded, pre-registered and run to completion found no benefit. A defensible record of success has to be built before it is needed, and nothing here suggests anyone does that.

    Support

    What would break it

    One grade A case with a positive outcome. The likeliest source is a mandated disclosure regime that requires reporting whether or not the news is good, or a pre-registered trial whose funder has no stake in the result.

    Where it strains

    A small, hand-assembled sample, and the library deliberately prefers cases that fill a gap in its taxonomy over the next available story. That selection rule is not neutral with respect to this finding. Treat it as a property of this collection first and of the evidence landscape only second. AAI-2026-021 also narrows what the proposition can claim: grade A evidence is now demonstrably producible without an adversary, so the scarcity of well-evidenced successes is better read as a scarcity of trials than as a structural impossibility.

  2. P2

    One half of the error rate is missing, and it is always the flattering half

    recurrent

    Three deployments in three unrelated sectors published the side of their error rate that made them look effective and omitted the side that carried the cost.

    A children’s hospital reported that its vendor-built sepsis model caught a given share of cases and could not report the positive predictive value, because the vendor had never published one. A disk-drive factory reported catching fifteen leaks out of sixteen and never counted the false alarms. A pharmacy chain ran face matching across hundreds of stores for eight years and counted nothing at all — which is the same omission with the measurement apparatus removed.

    The direction is consistent. Recall flatters a detector; precision is what its false positives cost somebody else. In all three the published number was the one that described the system’s successes, and the missing number was the one that would have described what it did to the people it was wrong about.

    The counter-case shows the omission is a choice rather than a constraint. Amsterdam measured false positive shares for eight demographic groups across three datasets, before and after a debiasing intervention, and compared them against the human process it was replacing. That is both sides of the error rate, disaggregated, on a live system — and it is what let the city discover that its correction had moved the harm rather than removed it. Nothing about the domain made this easier than sepsis alerts or shoplifting detection. Somebody decided to count.

    What an operator should take from itBefore accepting a performance number, from a vendor or from your own team, name the metric that is missing and ask who bears its cost. If nobody can produce it, you are not looking at a measured system.

    Support

    What would break it

    A case that publishes precision and withholds recall — the omission running the other way. None has appeared yet, which is itself the point. A counter-case has: a city that measured both sides, by demographic group, across three datasets, and stopped the system on what it found.

    Where it strains

    Three cases, one of them grade C. It holds on the strength of the sectors being unrelated rather than on the count, and the methodology note treats the underlying measurement problem at greater length.

  3. P3

    Organisations measure the proxy they can move, not the outcome they meant

    recurrent

    A national grid operator cut its solar forecasting error by two thirds and never measured the balancing cost the forecast existed to reduce, converting it instead through a planning rule of thumb. A government chatbot raised its self-assessed accuracy from 76% to 90% and moved user satisfaction not at all. A sepsis alert improved time to antibiotics while mortality, ICU admission and ICU-free days stayed flat and underpowered. A support copilot measured resolutions per hour; the randomized replication that followed it measured customer retrials too, and found those unmoved.

    A seven-day forecast of urgent dialysis demand was run live twice at four Toronto hospitals and measured on one quantity only — mean absolute error in procedures per day. Nurse hours, backfill decisions, overtime and cost were the reason the forecast existed and none of them was measured; the paper closes by saying whether the model would save anything is an open question. The staffing estimate that did exist survives only in a trade-press report of a conference talk given eight months before publication.

    In each, the measured quantity was the tractable one and the quantity that justified the spend went unmeasured. The pattern is not dishonesty. It is that proxies are cheap, immediate and attributable, while outcomes are slow, confounded and owned by somebody else in the organisation.

    The counter-case is the one deployment here where somebody measured the outcome properly, and it shows why the proposition matters rather than refuting it. A randomized trial of AI chest X-ray prioritization measured both halves: the proxy the feature was built to move — time from X-ray to report — improved significantly, from 47 hours to 34.1. Every clinical outcome it was installed to improve stayed flat, in an adequately powered trial with tight intervals. The proposition’s falsifier asks for a case that measured both and found them to move together. This one measured both and found them to move apart.

    The failure mode it produces is specific: a deployment can report a real, large, correctly measured improvement and still have no evidence that it did the thing it was bought to do.

    What an operator should take from itWrite down the outcome before choosing the proxy, and put a date on when the outcome will be measured. A proxy that improves while its outcome is never checked is indistinguishable from one that does not work.

    Support

    What would break it

    A deployment that measured its outcome and let the proxy go unmeasured, or one that measured both and found them to move together.

    Where it strains

    Four of the six supporting cases are grade C, and in two of them the proxy genuinely improved by a large margin. The proposition is about what was left unmeasured and should not be read as saying the proxies were worthless. The dialysis case also shows the pattern is not only about what an organisation measures but about which measurements survive the journey to an audience: the staffing estimate was public eight months before the accuracy figure, and it is the accuracy figure that reached peer review.

  4. P4

    Not one case reports both a measured cost and a measured benefit

    recurrentAbout the library

    The closest any case comes is a national grid operator that disclosed £1,038,500 of development cost and roughly £45,000 a year to run the service — and then derived its benefit by multiplying a planning rule of thumb.

    Elsewhere the halves separate cleanly. A peer-reviewed field study establishes a 15% productivity effect and discloses no cost anywhere in the record, because the firm is anonymous. A rideshare company reports per-engineer AI spend to the dollar on an internal dashboard, against an executive who said publicly that he could not draw a line from the tools to a shipped feature. A factory claims millions of dollars a year of avoided scrap with no derivation. A government team publishes two and a half years of methods and not one cost figure.

    The sharpest instance is a city that measured the benefit and found it negative. Amsterdam’s welfare fraud model was projected in internal documentation to keep up to 125 people out of debt collection and save €2.4 million a year. Its councillors were never told what it cost; the figure exists only because three newsrooms asked, and the city’s answer — about €500,000 plus €35,000 for a contractor — came with a caution that it was an estimate. The measured result was more investigations and no improvement in the hit rate.

    Cost telemetry and value telemetry arrive years apart, and cost arrives first. That ordering is itself the finding: metered tools make spending legible immediately, while the benefit needs an experiment almost nobody runs.

    One case now breaks that ordering without breaking the proposition. In AAI-2026-021 somebody did run the experiment — a randomized trial across five NHS trusts, powered, pre-registered, independently analysed — and measured the benefit to a confidence interval. There was none. The cost half is registered as that trial’s seventh secondary outcome and has not been published; the paper calls the costs of the function “considerable and avoidable” and gives no figure. So the halves still arrive separately, and for once it is the benefit that arrived first. The proposition survives its most rigorous test to date, which is not the same as being vindicated by it.

    What an operator should take from itAssume that any return-on-investment figure for an AI deployment — including your own — is one measured half and one modelled half. Ask which half is which. Nobody in this library has both.

    Support

    What would break it

    A single case with both numbers measured. This is the easiest proposition in the set to falsify, and the case that does it would be worth more to the library than the proposition.

    Where it strains

    An absence across a small library is weak evidence about the world and strong evidence about what gets published. Organisations holding both numbers may simply have no reason to publish them, and commercial sensitivity is the obvious reason they would not.

  5. P5

    A human in the loop is a control only if someone measures the human

    recurrent

    Dutch officials reviewing students the algorithm had scored identically selected those with a non-European migration background 5.5 times as often. The reviewer amplified the model rather than checking it, and nobody measured that for eleven years.

    Rite Aid’s staff were handed a match alert with no confidence estimate, no training in how the system worked, and an instruction to act on it. In the sepsis deployment the human never entered the loop at all: workflow exclusions suppressed the alert in roughly four cases out of five before anyone saw it, which is why the delivered sensitivity was 18% against a vendor-reported 91%.

    The counter-case matters as much as the three. A factory leak detector emails a probability and its supporting evidence to an engineer, who walks to the machine and physically checks; nothing on the line is actuated automatically. That arrangement holds — and what makes it hold is not that a human is present but that the human’s job is defined, bounded, and verifiable against the physical world.

    What an operator should take from it"Human in the loop" describes an org chart, not a safeguard. The safeguard is the measurement of what the human does with the output: override rate, agreement by subgroup, and how often the output reaches a human at all.

    Support

    What would break it

    A deployment where unmeasured human review demonstrably caught model errors at scale. The counter-case here is suggestive but self-reported and grade C.

    Where it strains

    All three supporting cases are ones where something went wrong, so the reviewer's failure is what got investigated. Deployments where review works quietly would not enter this library at all, which is P1 operating on P5.

  6. P6

    Liability attaches to the paperwork, not to the model

    recurrent

    Cruise briefed three regulators in person on the morning after its vehicle dragged a pedestrian, and was charged over neither meeting. The criminal count came from two written incident reports filed under a standing order — the 1-day report and the 10-day report, the documents nobody staffed.

    The Dutch regulator found that failing to justify the selection criteria was itself the violation, independent of the disparity those criteria produced. In the Workday litigation, discovery rules and attorney–client privilege — not model documentation — decide what anyone will ever learn about the system’s disparate impact. Korea’s regulator reached a privacy policy that named some overseas processors and omitted one.

    The clearest instance is the one where no model was ever identified. Across more than two thousand court decisions about fabricated legal authority, a specific AI product is named in about one row in nine — and in the Nebraska Supreme Court decision read for that case, counsel denied using AI at all. The court struck the brief, dismissed the appeal and referred counsel for discipline anyway, on the duty of candour and the fictitious citations. The sanction did not depend on establishing a cause.

    Italy’s Garante reached DeepSeek in three days without looking at the system at all. Its findings were that the privacy policy was English-only and inadequate on its face, that no lawful basis was stated for each processing activity, that the policy itself disclosed storage in China without the Regulation’s safeguards, that no EU representative had been appointed, and that the reply to the Authority’s questionnaire failed to clarify the processing — a separate breach, of the duty to cooperate. Every one of those is readable from published documents and correspondence by someone with no access to the service.

    In none of these did an authority examine the model and find it wanting. The model was the occasion; the document was the offence.

    What an operator should take from itRegulatory exposure sits in the filings, disclosures and written justifications around a system, and those are usually owned by whoever has spare capacity. Staff the paperwork as part of the system, because legally it is part of the system.

    Support

    What would break it

    An enforcement action turning on a technical finding about model behaviour rather than on a disclosure, a justification, or a filing.

    Where it strains

    The best-evidenced proposition in the set — every supporting case is grade A or B — but all six sit in jurisdictions with mature administrative law or an adversarial court system, which is also where records of this kind come from at all. It may describe institutions that publish their findings more accurately than it describes AI. The Italian case also shows the limit of reading the pattern as neglect: DeepSeek did answer, within a day, and the answer was itself found to be a breach.

  7. P7

    Self-assessment of an AI system's effect can be wrong in sign, not merely in size

    emerging

    In the only randomized trial in this library, experienced maintainers working on their own repositories were 19% slower with AI assistance and believed they had been 20% faster. That is an error of direction, not magnitude, held by people with direct hands-on experience of the tool during the task being measured.

    Expert forecasting did worse. Economists and machine-learning researchers predicted speedups of roughly 38 to 39% for the same trial.

    The government chatbot shows the same gap from the user’s side: people discounted explicit inaccuracy warnings because the answers carried the GOV.UK brand, and the team recorded the finding and published it. Fourteen points of measured accuracy later, satisfaction had not moved.

    The pattern recurs where the measurement is most careful. Twenty-two public service broadcasters had 271 journalists rate more than 2,700 AI assistant answers about the news and found a significant issue in 45% of them; the same report cites separate research finding that just over a third of UK adults completely trust AI to produce accurate summaries, rising to almost half of under-35s. Measured error and reported confidence point in opposite directions in the same document.

    What belief tracks, in all of them, is fluency — the system’s outputs are well-formed, arrive quickly, and look like the thing that was asked for. None of those properties is the effect.

    What an operator should take from itA satisfaction survey about an AI tool measures how the tool feels, which the evidence here says is uncorrelated with — and can be opposite to — what it does. A deployment decision resting on user sentiment rests on nothing.

    Support

    What would break it

    A measured deployment where practitioner self-report tracked the measured effect. The trial's own authors believe the gap has narrowed since 2025 and describe their evidence for that as very weak.

    Where it strains

    Held at emerging on three cases rather than promoted, because the three are not quite the same phenomenon. One measures practitioners assessing their own performance; two measure audiences trusting an output. Those may share a mechanism or may not, and nothing here tests whether they do. The strongest of the three is also a narrow sample — sixteen developers in mature repositories they maintain, a population its own authors decline to generalise from.

Where the library disagrees with itself

A synthesis that only finds agreement is not reading carefully. These are the places where the record points in two directions, kept on the page rather than dropped from it.

Who benefits from a copilot

The Fortune 500 support study found gains concentrated in novices and little for experts — a gradient that falls monotonically with skill. The randomized replication at another firm found an inverted U: mid-tier agents gained most on speed, and the top quintile got measurably worse on customer ratings and on repeat contacts alike.

Both are in the same case record. Neither has been reconciled, and “AI helps the least experienced most” is the version that travels.

Whether AI helps experts at all

The support copilot lifted the least experienced and left experts flat. The developer trial found experienced maintainers actively slowed, by 19%.

Compatible readings exist — task familiarity is a moderator in both, and the two studied different work — but the pair is routinely quoted as though it were one finding pointing one way. It is two findings, and the library holds them apart.

What a company says against what it files

Klarna’s chief executive told Bloomberg in May 2025 that the AI-first approach to customer service had produced “lower quality”, that it “wasn’t the right path”, and that the company was recruiting human agents again.

The annual report filed nine months later describes no reversal and still presents the assistant as handling 80% of chats in the year to 31 December 2025. Both are Klarna speaking. Only one of them is filed under securities-law liability, and it is not the one that concedes a problem.