evaluationInconclusiveC
Four Toronto hospitals ran a seven-day forecast of urgent dialysis demand live, twice, without showing it to anyone. In the first silent period it exactly matched the static average it was meant to beat; in the second it beat it comfortably. The published paper reports the two periods combined as a 26.7% improvement and does not mention the tie.
evaluationhealthcare
evaluationNegativeA
A randomized trial across five NHS trusts gave half its chest X-ray sessions to AI worklist prioritization and half to none, with the same AI available to reporters in both arms. Prioritization cut the median time from X-ray to report from 47 to 34.1 hours and moved nothing else: median time to CT was 53 days in both arms, against a national standard of 72 hours.
scaled productionhealthcare
governance regulatoryNegativeB
A public database records 2,042 court decisions worldwide addressing fabricated legal authority in filings. Most involve self-represented litigants rather than lawyers, nine in ten never establish which tool produced the text, and in one of the two supreme court decisions read here the lawyer denied using AI at all.
unknownprofessional services
evaluationMixedB
Twenty-two public service broadcasters had 271 journalists rate 2,709 AI assistant answers about the news. Forty-five per cent had a significant issue, sourcing most of all, and a BBC-to-BBC rerun showed the rate falling from 51% to 37% in six months. Every measurement in the record was made by a publisher whose own content was being cited.
scaled productionmedia