Back to the news portal
Health & Life SciencesResearch paperResearchSource analysisNorth AmericaUnited States

Did machine learning beat logistic regression on surgical sepsis?

A 328,292-case US study found no significant predictive advantage for LASSO, XGBoost or random forest. Sepsis was rare, the best risk group still had low absolute incidence, and no model has external validation.

By The Impact of AI Research DeskReleased 2 October 2026 at 04:05 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesClinical AISurgical riskSepsisModel comparisonBreast cancer surgery

Research topic

Whether machine-learning models improve prediction of rare 30-day sepsis after breast cancer surgery compared with conventional logistic regression

The Impact of AI research cover asking whether machine learning beat logistic regression on surgical sepsis, with conceptual clinical risk pathways and four model cards.
AI-generated editorial illustration. The clinical pathway and model cards are conceptual; they do not depict a patient, provider record, hospital system or measured chart.

At a glance

  • 1The retrospective ACS-NSQIP cohort included 328,292 women who underwent breast cancer surgery from 2008 to 2022; 838 had coded sepsis within 30 days, an incidence of 0.26%.
  • 2In a 64,819-patient held-out test set, AUCs ranged from 0.727 for logistic regression to 0.742 for random forest, with no statistically significant machine-learning advantage.
  • 3The highest predicted-risk decile contained 35.7% of test-set events, but its observed sepsis incidence was only 0.86%; low positive predictive value and absent external validation rule out deployment as a screening or triage tool.

Living evidence record

Impact record IAI-0D7CU8Z

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

2 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

A large dataset posed a difficult rare-event question

The study asked two related but different questions: which recorded characteristics were associated with sepsis after breast cancer surgery, and whether more complex machine-learning methods could predict that uncommon outcome better than conventional logistic regression. That separation matters. An association is not proof that changing a characteristic will prevent sepsis, while a predictive model can rank risk without explaining what caused it.

Researchers retrospectively analysed American College of Surgeons National Surgical Quality Improvement Program data from 2008 to 2022. The cohort included 328,292 adult women having partial, simple or radical mastectomy for breast cancer under general or monitored anaesthesia at participating US hospitals. Urgent and emergency cases and patients in the most severe ASA physical-status category were excluded, so the result concerns selected elective surgery rather than every person treated for breast cancer.[1][2]

Sepsis was rare and defined by the registry

The primary outcome was ACS-NSQIP-coded sepsis within 30 days. It occurred in 838 patients, or 0.26% of the full cohort. Incidence was 0.08% after partial mastectomy, 0.40% after simple mastectomy and 0.46% after radical mastectomy. Septic shock was recorded separately in 100 patients, or 0.03%, and was not combined with the sepsis endpoint used for inference and prediction.

The registry definition is important. It is anchored in older systemic inflammatory response criteria and is not interchangeable with the modern Sepsis-3 definition based on organ dysfunction. The database lacks the SOFA components needed to reclassify past cases. Trained reviewers used uniform audited criteria, but the reported incidence and the model's target remain a registry-specific construct, and ascertainment rules may also have evolved over the 15-year period.[2]

The comparison used a held-out test set

For risk-factor analysis, 4,201 people with a missing included predictor were removed, leaving 324,091 patients and 828 sepsis events. The predictive comparison randomly split this complete-case cohort into 259,272 training cases with 671 events and 64,819 held-out test cases with 157 events. The test set was not used to choose model settings, which provides a cleaner internal comparison than reporting training performance alone.

The four approaches were ordinary logistic regression, LASSO-penalised regression, XGBoost and random forest. LASSO used ten-fold cross-validation. XGBoost used class weighting and five-fold tuning based on precision–recall area, while random forest used balanced down-sampling and out-of-bag tuning. Those different imbalance treatments were reasonable but make clear that performance depended on implementation choices rather than an abstract contest between algorithm names.[2]

No machine-learning model achieved a significant advantage

On the held-out test set, logistic regression produced a receiver-operating-characteristic AUC of 0.727, with a 95% confidence interval from 0.687 to 0.766. LASSO reached 0.730, XGBoost 0.735 and random forest 0.742. None of those differences from logistic regression passed the study's statistical threshold: the reported DeLong p-values were 0.077, 0.409 and 0.089 respectively. The result does not prove machine learning is generally inferior; it shows that added complexity brought no measurable predictive advantage in this dataset.

Calibration strengthened that caution. Logistic regression and LASSO recorded Brier scores of 0.0024 and reasonable calibration. Before Platt recalibration, XGBoost and random forest substantially overpredicted risk, with Brier scores around 0.18. Recalibration improved the tree models, but the need for it underlines why discrimination alone is not enough. A model can order people from lower to higher risk while still assigning probabilities that are badly wrong.[2]

Risk concentration did not create a useful screening test

The internally validated logistic model showed an apparent AUC of 0.752 and an optimism-corrected AUC of 0.746 after 200 bootstrap samples. In the held-out comparison, the highest predicted-risk decile contained 56 of 157 sepsis events, or 35.7%, and the top two deciles contained 47.8%. That concentration may help frame follow-up research, but the absolute event rate in the highest decile was still only 0.86%.

At the threshold that maximised the Youden index, sensitivity was 77.7%, specificity 55.8% and positive predictive value 0.4%. Put plainly, about one in 250 people classified as higher risk developed the recorded outcome. The authors therefore say that no threshold from this analysis should be used to withhold reconstruction and that low positive predictive value plus absent external validation preclude screening or triage use.[2]

The associations are useful signals, not treatment effects

Prior sepsis had the largest adjusted association, with an odds ratio of 4.51 and a 95% confidence interval from 2.62 to 7.73. Functional dependence and dialysis were also strongly associated. Compared with no reconstruction, implant-based reconstruction had an adjusted odds ratio of 1.59 and autologous reconstruction 2.15. These numbers describe conditional associations in observational records; they do not show that reconstruction caused sepsis or that avoiding it would improve outcomes.

The database lacks tumour stage and biology, systemic treatment history, radiation exposure and the timing of neoadjuvant therapy. Those factors can influence immune function, wound healing, patient selection and the type of operation offered. Residual confounding is therefore a particularly serious alternative explanation for the reconstruction findings. Recorded race is also an administrative category and should not be interpreted as biology.[2]

Poor outcomes after sepsis show why validation matters

Patients with coded sepsis had much higher observed rates of reoperation, 46.5% versus 4.0%, and unplanned readmission, 73.7% versus 2.1%. Organ-space surgical-site infection was also more frequent, 34.0% versus 0.5%. These are unadjusted outcome comparisons around the same postoperative episode, not proof that an algorithm would prevent those events. They do explain why clinicians would value an accurate, calibrated warning if one can be validated prospectively.

For patients, the immediate lesson is not that an AI score should determine consent or procedure choice. It is that rare-event prediction can look respectable by AUC while producing many false alarms. An interpretable conventional model performed as well as the more complex alternatives here, which may be preferable when clinicians must explain risk, audit errors and decide whether any intervention tied to a score is proportionate.[2]

Generalisability remains untested

ACS-NSQIP is multi-institutional, but participating hospitals may not represent smaller, non-academic or international settings. The study used a random internal split rather than a later time period or a separate health system. It captures outcomes only within 30 days and lacks detailed timing, source, microbiology and clinical course. Coding can misclassify procedures, while very low prevalence leaves only 157 events in the test set and widens uncertainty around differences between models.

The licensed participant-use data are available through the American College of Surgeons subject to approval and are not public. The authors declared no competing interests and acknowledged ACS-NSQIP and participating hospitals as the data source; the manuscript does not report a separate funding statement. The American College of Surgeons did not verify the statistical analysis or conclusions.[2]

What would change the assessment

Confidence would rise with external validation in an independently assembled hospital network and a chronological test that reflects changes in surgery, cancer therapy and sepsis definitions. Results should include calibration, precision–recall performance and decision-curve analysis, with audits across age, disability, recorded race, hospital type and procedure. Repeating the study under contemporary Sepsis-3-compatible data would clarify whether the target remains clinically relevant.

A prospective study would then need to specify what action follows a high-risk score, compare that pathway with usual clinical assessment and measure patient outcomes, unnecessary interventions, workload and fairness. Until those steps are complete, the paper supports a simpler conclusion: on this rare outcome, more complex algorithms did not beat well-specified regression, and none is ready to decide care.[2]

What this means for people

  • The models are not ready to determine who receives reconstruction, screening or treatment.
  • Rare-event scores can produce many false alarms even when their AUC appears acceptable.
  • An interpretable model may be easier for patients and clinicians to question, explain and audit when performance is comparable.

Global context

The study draws on participating US hospitals and a US surgical registry. Breast-cancer pathways, coding, access to reconstruction and sepsis surveillance differ across countries, so the result should not be transferred internationally without new validation. Its wider relevance is methodological: in rare clinical outcomes, complex machine learning must beat a strong, calibrated regression baseline and demonstrate useful decisions rather than merely a slightly higher point estimate.

What the evidence does not yet show

  • The study is retrospective and observational, so adjusted associations do not establish causes or effects of changing surgical choices.
  • The registry lacks tumour biology, cancer stage, systemic therapy, radiation and neoadjuvant-treatment timing, leaving important residual confounding.
  • Sepsis used an older SIRS-based registry definition, outcomes stop at 30 days, and only 157 sepsis events appeared in the held-out test set.
  • Validation was internal; no independent hospital system, chronological cohort or international setting has tested the models.

What to watch next

  • External and chronological validation with contemporary sepsis definitions.
  • Calibration and false-alert burden across hospitals and patient groups.
  • Prospective evidence that a specified response to a risk score improves outcomes without unnecessary care.
  • Whether a simpler regression model remains competitive when richer oncology and treatment data are available.

Evidence trail

Sources used for this report

Links checked 2 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Health & Life Sciences

Can an AI digital twin shorten an ICU stay before it reaches a clinical trial?

ARPA-H has awarded the University of Vermont up to $38 million to model individual patients' immune responses. The five-year, milestone-based project targets a 25% reduction in ICU stays, but it has not yet demonstrated that result in patients and publishes no planned trial denominator.

6 min · 3 sources

Health & Life Sciences

Can an ICU glucose digital twin be trusted after testing on ten patients?

A peer-reviewed study reports that a pretrained transformer can produce 15- and 30-minute glucose forecasts for ten septic ICU patients within seconds. The result is technically useful, but it is a retrospective forecast test—not evidence that the system improves insulin decisions or patient outcomes.

10 min · 1 source

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.