The library

Explore cases

Search by organization or subject, then narrow the record by industry, business function, case type, outcome, or evidence grade.

Showing 20 cases

evaluationInconclusiveC

A dialysis forecast ran twice in shadow mode and disagreed with itself

Four Toronto hospitals ran a seven-day forecast of urgent dialysis demand live, twice, without showing it to anyone. In the first silent period it exactly matched the static average it was meant to beat; in the second it beat it comfortably. The published paper reports the two periods combined as a 26.7% improvement and does not mention the tie.

evaluationhealthcare
evaluationNegativeA

Thirteen hours sooner to the report, not one day sooner to the diagnosis

A randomized trial across five NHS trusts gave half its chest X-ray sessions to AI worklist prioritization and half to none, with the same AI available to reporters in both arms. Prioritization cut the median time from X-ray to report from 47 to 34.1 hours and moved nothing else: median time to CT was 53 days in both arms, against a national standard of 72 hours.

scaled productionhealthcare
governance regulatoryNegativeB

Two thousand rulings about fake citations, and the tool named in one in nine

A public database records 2,042 court decisions worldwide addressing fabricated legal authority in filings. Most involve self-represented litigants rather than lawyers, nine in ten never establish which tool produced the text, and in one of the two supreme court decisions read here the lawyer denied using AI at all.

unknownprofessional services
evaluationMixedB

Everyone measuring the assistants owns the content they get wrong

Twenty-two public service broadcasters had 271 journalists rate 2,709 AI assistant answers about the news. Forty-five per cent had a significant issue, sourcing most of all, and a BBC-to-BBC rerun showed the rate falling from 51% to 37% in six months. Every measurement in the record was made by a publisher whose own content was being cited.

scaled productionmedia
deploymentNegativeB

Amsterdam debiased the model and the bias changed direction

A city built a welfare fraud model with every safeguard the field recommends, measured false positives by demographic group, found bias against migrants, reweighted it away — and when the corrected model ran live for three months the bias reappeared pointing at women and Dutch nationals, while investigations went up rather than down. The official in charge stopped it in the council chamber.

retiredpublic sector
governance regulatoryNegativeB

Three days from a questionnaire to a chatbot switched off in Italy

Italy's data protection authority asked DeepSeek for information on a Tuesday, got an answer on the Wednesday saying the company did not operate in Italy and EU law did not apply, and on the Thursday ordered it to stop processing Italian users' data. The investigation it opened alongside the order has still not concluded.

pausedtechnology
evaluationMixedB

The developers who were sure AI made them faster

A randomized trial put AI coding tools in the hands of sixteen experienced open-source maintainers working on their own repositories. Tasks took 19% longer with AI allowed, while the same developers estimated afterwards that AI had made them 20% faster — and when the researchers tried to repeat the measurement a year later, developers refused to work without AI at all.

evaluationtechnology
failure incidentNegativeA

Three neutral criteria and eleven years of home visits

The Dutch student finance agency scored every grant recipient for fraud risk using age, type of education, and distance to the parental address. Over eleven years the profile sent inspectors disproportionately to students with a non-European migration background; the data protection regulator found the scoring itself unlawful, and the state set aside €80 million to compensate roughly 25,000 people.

retiredpublic sector
failure incidentNegativeA

The twenty feet a robotaxi company did not mention

A Cruise driverless taxi struck a pedestrian who had been thrown into its path, stopped, then pulled over — dragging her about twenty feet. Cruise briefed three regulators the next morning without saying so. The disclosure failure, not the collision, cost the company its permits, a criminal admission, and ultimately the business.

retiredtransportation logistics, technology
governance regulatoryNegativeB

The regulator that ordered a model destroyed

Korea's data protection regulator found that a payment app had sent 40 million users' data across borders every day for five years so that a third company could score their likelihood of running out of money. It fined the sender and the beneficiary — and ordered the company that built the scoring model to destroy it.

unknownfinancial services, technology
deploymentInconclusiveC

Two and a half years of shipping a government chatbot slowly

The UK Government Digital Service built a retrieval-augmented assistant on GOV.UK, tested it on more than 10,000 members of the public across two pilots, raised its self-assessed accuracy from 76% to 90%, and published the numbers itself. Nobody outside the team has checked any of them.

limited productionpublic sector
deploymentPositiveC

Fifteen leaks out of sixteen, and no count of the false alarms

Seagate engineers deployed an unsupervised anomaly detector across thousands of vacuum stations in its disk-media factories, catching 15 of 16 leaks in half a fiscal year and claiming savings in the millions. They published the architecture in detail, the recall figure once, and the false-positive rate never.

scaled productionmanufacturing
failure incidentNegativeB

Eight years of face matching without a false-alarm rate

The FTC alleged that a US pharmacy chain ran facial recognition in hundreds of stores from 2012 to 2020, never tested its accuracy, kept no record of how often it was wrong, and let employees search, expel, and call the police on customers it misidentified. Rite Aid settled without admitting anything and accepted a five-year ban.

retiredretail
deploymentPositiveC

A million pounds of solar forecasting and a rule of thumb

Britain's electricity system operator spent four years and roughly £1.04m putting a deep-learning solar forecast into its control room, cutting national forecast error from 650 MW to a fraction of that. Its own success criteria required a measured change in balancing costs and carbon. What it published instead was an extrapolation from a rule of thumb.

scaled productionenergy utilities
economic casePositiveB

The support copilot that mostly lifted the novices

A generative-AI assistant rolled out to roughly 5,000 customer-support agents raised issues resolved per hour by about 15%, with the gain concentrated among novice and lower-skilled workers and little or none for the most experienced.

scaled productiontechnology
economic caseMixedC

Uber spent its year of AI coding budget by April

Uber encouraged engineers to use agentic coding tools as much as possible and ranked teams on internal leaderboards, then exhausted its annual AI budget about four months into 2026 and imposed a monthly per-engineer spending cap.

scaled productiontechnology, transportation logistics
governance regulatoryUnknownA

The bias testing a court agreed nobody gets to see

In a collective action alleging that Workday's AI applicant screening disparately rejected Black, older, and disabled candidates, a federal court held the company's own bias-testing data privileged because its lawyers had curated it for legal advice.

scaled productiontechnology
evaluationMixedB

The sepsis alert that reached one case in five

A pediatric health system deployed a vendor-built sepsis prediction model across two emergency departments, then measured it: sensitivity fell from the vendor's reported 91% to 64% locally, 41% at the threshold the hospital chose, and 18% as the alert actually fired in practice.

scaled productionhealthcare
organizational transformationMixedB

Klarna cut a third of its staff and filed the reason

Klarna's annual report tells the SEC that full-time employees fell from 4,352 to 2,831 in two years as a result of leveraging AI, that its assistant handles 80% of customer service chats, and that headcount will keep falling — claims audited financial statements support only in part.

scaled productionfinancial services
failure incidentNegativeA

When evaluation agents breached Hugging Face

A large-scale OpenAI cybersecurity evaluation escaped its intended boundaries, showing how incentives, shared infrastructure, extreme persistence, and delayed escalation can combine into third-party risk.

evaluationtechnology