AAI-2026-006 · evaluation

The sepsis alert that reached one case in five

A pediatric health system deployed a vendor-built sepsis prediction model across two emergency departments, then measured it: sensitivity fell from the vendor's reported 91% to 64% locally, 41% at the threshold the hospital chose, and 18% as the alert actually fired in practice.

Bevidencemixed outcome

Children's Healthcare of Atlanta · Epic SystemsVerified Sep 12, 2026

Case in one sentence

A pediatric health system implemented the sepsis prediction model built into its electronic health record, measured it against its own records, and found that the vendor’s reported 91% sensitivity became 18% by the time the alert was actually reaching clinicians at the bedside.

Executive summary

Children’s Healthcare of Atlanta deployed a vendor-developed pediatric sepsis prediction model across two emergency departments, wired to a nurse-facing interruptive alert that prompts staff to call a sepsis huddle. It then ran an independent evaluation across 599,163 ED visits spanning 2021 to 2024.1

The performance degraded in three identifiable steps. At the vendor’s recommended threshold of 6, the model caught 64% of sepsis cases locally, against the 91.25% the vendor reported from its development site. The hospital’s sepsis committee had raised the threshold to 8 to keep the positive predictive value above 10%, which cut sensitivity to 41%. And as the alert actually fired in clinical practice — after the workflow exclusions that suppress it — sensitivity was 18%, or 207 of 1,114 sepsis cases.1

Care processes still improved. Time to first fluid bolus fell by 16.7 minutes and time to first antibiotic fell from 112 to 102 minutes. Thirty-day mortality, ICU admission rate, and ICU-free days all moved in the favourable direction and none of those changes reached statistical significance.1

The authors’ own summary of the lesson is that local workflows, documentation patterns, and patient populations make published model performance hard to generalise. The specific gap this case adds to the library is that each step of the degradation is attributable and none of it would be visible to a buyer reading the vendor’s number.

Research question

When a health system deploys a prediction model that came with its electronic health record, how much of the vendor’s reported performance survives local configuration and the clinical workflow it fires into?

Organization and operating context

Children’s Healthcare of Atlanta runs the two emergency departments studied, one of them an academic quaternary care centre with 45 ED beds and over 70,000 visits a year, on an enterprise-wide Epic installation.1 Pediatric sepsis is a low-incidence, high-stakes target: 0.34% of ED visits in this population met the IPSO sepsis definition, and pediatric sepsis carries roughly 4% in-hospital mortality nationally with an estimated $7 billion in annual US hospitalisation costs.1

The model is Epic’s, a scaled linear model fitted with a non-linear reduced gradient method and developed on 116,419 emergency encounters from Nationwide Children’s Hospital. Its reported performance was 91.25% sensitivity at 95% specificity, at a threshold of 6.1

The situation before AI

Sepsis recognition in a paediatric ED is a timing problem. The condition is rare per visit, deteriorates fast, and the clinical signs overlap with far more common presentations. The existing answer is human vigilance supported by screening protocols, and the metric that matters operationally is how fast treatment starts once someone suspects it.

That baseline matters for reading this case: the model was not replacing a measurement, it was added on top of clinicians who already recognise most sepsis. As the paper’s own alert-suppression logic shows, many alerts never fire precisely because the care team has already acted.

The AI intervention or event

The model scored patients continuously, filing a score to the database every 15 minutes as new data arrived. When a score crossed the threshold, a nurse-facing interruptive alert fired, prompting a sepsis huddle.1 Alerts were suppressed for patients over 18, those with a sepsis order set already started, those with a documented huddle outcome, those in a trauma room, and those already discharged.1

Before going live, each site ran the model silently for months while stakeholders watched its behaviour. After roughly ten months of background operation, the sepsis committee reviewed the data and chose a threshold of 8 rather than the vendor’s 6, to hold the positive predictive value at 10% or better.1 Hospital A went live on 2022-01-13; Hospital B, after its own background period, on 2023-03-13.1

That threshold decision is the case in miniature. The committee was trading sensitivity for precision because at the vendor’s threshold the alert would have been wrong roughly 93 times in 100 — a positive predictive value of 7% locally.1

Outcomes and economics

Measured against 1,114 sepsis cases in the post-implementation cohort:1

Measure Vendor reported, threshold 6 Local, threshold 6 Local, threshold 8 As implemented in practice
Sensitivity 91.25% 64% (717/1,114) 41% (467/1,114) 18% (207/1,114)
Specificity 95% 97% 99% 99%
Positive predictive value Not published 7% (717/9,395) 24% (467/1,994) 11% (207/1,820)

Area under the curve post-implementation was 0.9, falling to 0.88 when only alerts firing before sepsis onset counted as true positives.1

On care processes, time to first fluid bolus fell 16.7 minutes and time to first antibiotic fell from 112 to 102 minutes.1 On patient outcomes, 30-day mortality moved from 6% to 4%, ED-to-ICU admission from 87% to 84%, and ICU-free days from 6 to 5 — none statistically significant.1

No cost figures are reported: not the licence, not the configuration and validation effort across roughly a year of background running per site, not the clinical time spent on huddles triggered by the nine-in-ten alerts that were not sepsis at the implemented threshold.

Causal assessment and competing explanations

This case is labelled associational, and that is a deliberate downgrade from how the finding is often described. The design is a retrospective pre/post comparison with no concurrent control: everything that changed in paediatric sepsis care between 2021 and 2024 at these hospitals is bundled into the “after” period alongside the model. The two sites also went live fourteen months apart, so the post-intervention cohort mixes one site with the model and one without for part of its span.

Sepsis care was a national quality-improvement target throughout this period, and the IPSO collaborative the sepsis definition comes from is itself an improvement programme. A secular trend toward faster fluids and antibiotics is a live competing explanation for a 16.7-minute change that the study cannot exclude.

On the performance gap, the authors attribute the shortfall to local workflows, documentation patterns, and patient population, and they flag that the vendor’s validation population had a sepsis incidence of 0.12% against 0.34% locally.1 That base-rate difference cuts in an awkward direction: a higher local incidence should, other things equal, raise the positive predictive value, so the 7% local PPV at the vendor’s threshold is not explained by rarity.

Failures, limitations, and governance

  • The number a buyer sees is not the number a patient gets. The vendor’s single reported figure, 91.25% sensitivity, described neither the threshold the hospital would choose nor the workflow the alert would fire into.
  • Unpublished metrics: the vendor reported no positive predictive value and no AUC, so the two measures that determine alert burden and discrimination were unavailable for comparison at purchase.1
  • The precision-sensitivity trade was made locally and quietly. A sepsis committee, reading ten months of silent-mode data, moved the threshold and halved sensitivity to make the alert tolerable. That is good governance and it is invisible outside the institution.
  • Alert suppression is not a bug: most of the drop from 41% to 18% is the alert correctly staying silent when clinicians had already acted. A sensitivity figure that ignores this flatters the model; one that includes it understates the model and describes the system.
  • No outcome signal: the changes in mortality and ICU use were not significant, and the study was not designed to detect them.
  • Single system, no replication: the same vendor’s adult model has produced opposite findings at different health systems, one reporting alert fatigue without benefit and another reduced mortality.1

What this case demonstrates

  1. Vendor-reported model performance is an upper bound measured under the vendor’s conditions, and the distance to delivered performance can be a factor of five.
  2. Degradation happens in identifiable stages — population shift, local threshold choice, workflow suppression — and each stage can be measured separately by the deploying institution, which is the only party that can see all three.
  3. A buyer cannot evaluate alert burden without a positive predictive value, and a vendor is not obliged to publish one.
  4. Silent-mode running before go-live is what made the threshold decision evidence-based here; ten months of it produced the local data the committee needed.
  5. Process measures move before outcome measures do, and a deployment can be genuinely useful on the first while showing nothing on the second.
  6. The honest denominator for a clinical alert includes the cases where it stayed quiet because humans got there first.

What this case does not demonstrate

  • It does not show the model caused the faster fluids and antibiotics; there is no concurrent control and sepsis care was improving nationally.
  • It does not show the model changed mortality, ICU admission, or ICU-free days, in either direction.
  • It does not establish that the vendor’s reported figure was wrong in its own setting.
  • It does not quantify the deployment’s cost, or the clinical time consumed by false alerts.
  • It does not generalise to other paediatric systems or other vendors. It is now flanked by evidence on the same vendor’s adult model, but that is a different model on a different population and corroborates the pattern, not the numbers.
  • It does not show that the pediatric model would behave as the adult model does. The vendor’s adult model was rebuilt between versions, from a penalised logistic regression to a gradient-boosted tree with local fine-tuning; no source read here says which generation the pediatric model belongs to.

Evidence assessment

Grade B — a rigorous evaluation whose numbers still have no corroboration. The primary source is peer-reviewed, published open access, read in full for this case, and unusually candid: it reports its model’s performance four ways, including the least flattering one, and states plainly that the vendor’s PPV and AUC were unavailable for comparison.1 The sepsis case definition was validated by two clinicians against the automated query. The institution had every incentive to report the version of sensitivity that made its deployment look best and did not.

What holds it at B is that the figures themselves still cannot be checked by anyone else. These are measurements of one health system’s own deployment against its own records; no outside party can reproduce them, and no second institution has published a comparable staged breakdown for this model. The conflict is structural rather than suspicious: the implementer is also the evaluator, which is what makes the alert-suppression analysis possible and what makes independent verification impossible.

A second chain now exists, and it is important to be exact about what it does and does not reach. A multicentre prospective validation published in February 2026 measured the same vendor’s adult sepsis model across 227,091 encounters at four US health systems, and found an AUROC between 0.82 and 0.92 with, in its own summary, “high institutional variability, low positive predictive value, and high alert burden” — concluding that institutions implementing the model should validate it locally.2 That is independent of this deployment, of its authors, and of the vendor, and it corroborates the general proposition this case is about: that this vendor’s sepsis models do not carry their reported performance across institutions, and that the deploying institution has to measure. It corroborates none of the pediatric figures. Different model, different population, different care setting. The case is therefore filed with two chains and stays at grade B, which is an editor’s judgment: a mechanical reading of the chain count would allow A, and this case declines it because a second chain that cannot touch the headline numbers is not corroboration of them.

One thing the second source does supply directly is the answer to a question this case leaves open. It reports the vendor’s own internal validation AUROCs for the adult model — 0.83 to 0.86 across three sites — and attributes them to a Microsoft Teams meeting with two named Epic employees on 25 February 2025.2 The vendor’s validation figures reach the peer-reviewed literature as a personal communication. That is not an accusation; it is the disclosure regime this case’s buyer was operating in, stated by someone other than the buyer.

Two internal inconsistencies in the paper are worth recording rather than smoothing over. The narrative text gives the positive predictive value at the implemented threshold as 22% (395/1,820) while Table 2 gives 24% (467/1,994); this case cites the table, which is labelled by column. And the reported confidence interval for the as-implemented AUC, “0.88 (95% CI, 0.86-0.87)”, does not contain its own point estimate, so the third decimal place of those AUCs should not be relied on.

Finally, the design bounds what the outcome figures can mean. A retrospective pre/post across a period when sepsis care was a national quality target cannot separate the intervention from the trend, which is why this case carries associational rather than the causal language the abstract uses.

Material claims

Each claim carries a controlled label, the evidence behind it, and what would change the label.

Claim Label Evidence What would change this
As the alert actually fired in practice, it reached 18% of sepsis cases. Verified Table 2: 207 of 1,114 cases1 A correction to the paper, or re-analysis under a different sepsis definition
At the vendor’s recommended threshold, local sensitivity was 64% against the 91.25% reported. Verified Table 2, with the vendor figure from the model’s development site1 The vendor publishing validation details that change the comparison
The vendor published no positive predictive value or AUC for comparison. Verified Stated in the results1 Publication of those metrics by the vendor
The hospital raised the threshold from 6 to 8, halving sensitivity, to hold PPV at 10% or better. Verified Implementation section and Table 21 A correction, or release of the underlying threshold analysis showing otherwise
Time to first fluid bolus fell by 16.7 minutes after implementation. Supported Pre/post comparison, P = .03, CI −31.8 to −1.5, with no concurrent control1 An interrupted time series or controlled design separating the model from secular trends
Time to first antibiotic fell from 112 to 102 minutes. Supported Pre/post comparison, P = .05, CI −19.1 to 0.1 — borderline by the authors’ own account1 A controlled design, or a larger sample resolving the borderline result
The model changed 30-day mortality, ICU admission, or ICU-free days. Unknown All three moved favourably; none reached statistical significance and the study was not powered for them1 A powered outcome study, or pooled data across institutions
The performance gap reflects population and workflow differences rather than an error in the vendor’s own figure. Inference The authors attribute it to local workflows, documentation, and population; the base-rate difference they flag does not by itself explain the local PPV Vendor disclosure of its validation population and method, or an independent audit of the original figure
The deployment was worth its cost. Unknown No licence, configuration, or clinical-time cost is reported, and no outcome benefit was established Cost disclosure alongside a powered outcome measure
The same vendor’s adult sepsis model varies widely in performance across institutions and needs local validation before deployment. Verified Prospective validation at four US health systems on 227,091 encounters: AUROC 0.82 to 0.92, the score threshold matching 60% sensitivity ranging from 14 to 37, and PPV from 0.13 to 0.262 A larger multicentre study finding consistent cross-site performance
The vendor’s own validation figures for the adult model reached the peer-reviewed literature as a personal communication rather than a publication. Verified The validation paper cites AUROC 0.83–0.86 across three internal sites to a Microsoft Teams meeting with two named Epic employees, 25 February 20252 The vendor publishing its validation study
The pediatric model’s local degradation would be reproduced at another pediatric site. Unknown No second institution has published a comparable staged breakdown for this model; the adult evidence shows wide cross-site variation, which cuts both ways A second pediatric health system publishing the same four measures

Direct quotations

“When implementing an externally developed model, local workflows, documentation patterns, and patient populations make it challenging to generalize published or reported model performance metrics to real world performance.”

— Kandaswamy and colleagues1 · locator: abstract, Discussion

“Vendor reported PPV and AUC are not available for comparison.”

— Kandaswamy and colleagues1 · locator: Results, Predictive performance

“In internal validation, Epic Systems has reported improved area under the receiver operating characteristic curve (AUROC) values for the ESM v2 between 0.83 and 0.86 across 3 internal validation sites (Tyler Sundberg and Elliot First, Epic Systems, Microsoft Teams meeting, February 25, 2025).”

— Wong and colleagues, citing the vendor’s own validation of its adult model to a personal communication2 · locator: Introduction, final paragraph

“High variability was noted in model performance across study sites, highlighting the need for local validation prior to deployment.”

— Wong and colleagues, on the same vendor’s adult model at four health systems2 · locator: Discussion, second paragraph

Revision notes

  • 2026-09-13 — Source review. Added a second chain and kept the grade. A multicentre prospective validation of the same vendor’s adult sepsis model — 227,091 encounters at four US health systems, published in JAMA Network Open in February 2026 — corroborates the general proposition this case rests on, that this vendor’s sepsis models do not carry their reported performance across institutions and have to be validated locally. It corroborates none of the pediatric figures, which remain measurements nobody outside Children’s Healthcare of Atlanta can reproduce, so evidence_grade stays B. That is an editor’s judgment and is flagged as one: with two source families a mechanical reading of the chain count would permit A, and this case declines it. The single_chain_rationale is removed, because its claim that “no second chain is reasonably available” turned out to be wrong — one was available in the adjacent adult literature and had simply not been looked for. Also newly recorded, and directly answering a gap this case had named: the vendor’s own internal validation AUROCs reach the peer-reviewed literature as a citation to a Microsoft Teams meeting with two named Epic employees. Three claims added, one “does not demonstrate” bullet split in two. Retrieved through the NCBI E-utilities full-text XML, because the PMC web page returns a reCAPTCHA challenge to this sandbox.
  • 2026-09-12 — Initial publication at Grade B. Single peer-reviewed source, read in full and cited with single_chain_rationale because no second chain is reasonably available for measurements of one institution’s own deployment. Labelled associational rather than causal: the design is a retrospective pre/post with no concurrent control, during a period when paediatric sepsis care was a national quality-improvement target. Two internal inconsistencies in the source — a PPV that differs between narrative and table, and a confidence interval that excludes its own point estimate — are recorded in the evidence assessment. Added clinical-care to the controlled business functions, which had no value for patient care.

Footnotes

  1. Kandaswamy, Orenstein, Muthu, McCarter, Braykov, Beus, and colleagues, “Early clinical evaluation of a vendor developed pediatric artificial intelligence sepsis model in the emergency department”, Journal of the American Medical Informatics Association 32(10):1542-1551, 2025-07-22, doi:10.1093/jamia/ocaf105. 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28

  2. Wong, Currey, Schwinne, Park-Egan, Meyer, and colleagues, “Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model”, JAMA Network Open 9(2):e260181, 2026-02-27, doi:10.1001/jamanetworkopen.2026.0181. Open access under CC-BY. 2 3 4 5 6

Evidence ledger

BWell supported

Credible original evidence supports the central account, but access, corroboration, measurement, completeness, or reproducibility carries a material limitation.

Material sources
2
Independent chains
2
Archived
0 of 2

The 2 sources below trace back to 2 separate origins. Whether those origins corroborate one another on this case's central claims is assessed in the case, not implied by the count.

What would strengthen this case

Grade A still needs corroboration of this deployment, which does not exist: another health system publishing a comparable staged breakdown for the same pediatric model, or a design with a concurrent control — an interrupted time series or a stepped-wedge rollout — that could separate the care-process improvements from secular trends. Release of the supplementary threshold analysis underlying the choice of 8 would let a reader check the trade-off the sepsis committee actually made. A multicentre prospective validation of the same vendor's adult model now corroborates the general finding — that this vendor's sepsis models vary widely by institution and need local validation — without touching the pediatric figures, and it records that the vendor's own validation AUROCs reached the literature through a personal communication rather than a publication.

  1. Chain 1 of 2Origin: esm-adult-external-validation

    • Multicenter Prospective Validation of an Updated Proprietary Sepsis Prediction Model

      Primary investigationIndependent reporting

      Access
      Prospective post-implementation data on 227,091 encounters at four US health systems — University of Michigan, Oregon Health & Science University, Emory Healthcare, and MetroHealth — with scores from both versions of the vendor's model computed every fifteen minutes on every emergency department and inpatient encounter. No patient data were shared between sites; each analysed its own.
      Method
      Prospective external validation reported against TRIPOD+AI, with sepsis defined by Sepsis-3 clinical criteria rather than by billing codes. Discrimination by AUROC; specificity, PPV, NPV and number needed to evaluate all reported at the threshold matching 60% sensitivity at each site, at encounter level and at prediction level over three time horizons. Includes a comparison against the timing of clinician recognition and a fairness audit by age, sex, race, and ethnicity.
      Conflicts
      None to the vendor. No author is employed by or funded by Epic Systems, the funders are named as having no role, and the disclosures are unrelated to sepsis prediction. The lead author led the 2021 external validation that found the vendor's first model wanting, which is a prior position rather than an interest. Every site is an implementer reporting on a model it has already bought and deployed — the same structural conflict as the pediatric study, four times over.
      Corroboration
      Its own design is the corroboration: four health systems, measured separately, reaching consistent conclusions about variability. Its AUROC range of 0.82 to 0.92 overlaps the 0.83 to 0.86 the vendor reports internally, so on discrimination the vendor's figure holds up where its first model's did not.
      Accountability
      Peer-reviewed, open access under CC-BY, institutional review board approval at each site, named data custodians per site, itemised author contributions, declared funding, and reported against a published reporting guideline.
      Notes
      Read in full from the PMC full-text XML through the NCBI E-utilities, because the PMC web page returns a reCAPTCHA challenge to this sandbox. Concerns the vendor's *adult* model — versions 1 and 2, patients aged 18 and over, inpatient and emergency department — and not the pediatric model this case is about. It is used here for the general finding and for what it records about vendor disclosure, never as corroboration of the pediatric figures.

      Accessed Sep 13, 2026No archive snapshot

  2. Chain 2 of 2Origin: chla-jamia-evaluation

    • Early clinical evaluation of a vendor developed pediatric artificial intelligence sepsis model in the emergency department

      Primary investigationDirect evidenceAnalysis

      Access
      The health system's own electronic health record data for 599,163 emergency department visits across two of its sites between 2021-01-01 and 2024-04-01, including model scores filed every 15 minutes, alert firings, and clinical outcomes.
      Method
      Retrospective cross-sectional pre/post comparison. Sepsis cases identified by IPSO criteria through an automated query validated by two clinicians at 95% sensitivity and 100% positive predictive value on manual review. Performance measured at the encounter level, reported separately at the vendor's recommended threshold, the locally implemented threshold, and as the alert actually fired in practice.
      Conflicts
      Written by the implementing institution about its own deployment, which is an interest in the result, and by informatics staff who configured the model. It is also the reason the paper can report what an outside evaluator could not observe: alert suppression by clinical workflow.
      Corroboration
      None available for this deployment. The paper notes that the same vendor's adult model produced widely different results across health systems, one reporting alert fatigue without outcome improvement and another reporting reduced mortality, which is context rather than confirmation.
      Accountability
      Peer-reviewed, published open access with a DOI, institutional review board approval recorded with a study number, and author contributions listed individually.
      Notes
      Read in full for this case, including Table 2. The vendor's reported figures come from the model's development at Nationwide Children's Hospital on 116,419 encounters, in a population with a sepsis incidence of 0.12% against 0.34% locally — a difference the authors flag as material to the comparison.

      Accessed Sep 12, 2026No archive snapshot

Cited in the synthesis