Sepsis diagnosis is four serial waits — for growth, for purity, for identification, for susceptibility — and in up to half of cases it ends with nothing. Why correlation-based prediction has underdelivered at the bedside, why mechanistic models have a regulatory path that learned ones do not, and why a spacecraft is the same problem with the excuses removed.
The Waiting Problem
Sepsis diagnosis, mechanistic patient models, and clinical reasoning under delay
Nicolas Waern — WINNIIO AB, Gothenburg, Sweden ORCID 0000-0001-7970-2707 · Published August 2026 · CC BY 4.0
Abstract
Sepsis causes an estimated 11 million deaths each year, and the interval between presentation and a pathogen-specific diagnosis remains measured in days rather than hours.[1] This paper argues that the interval is structural rather than operational, and that the widely cited remedy — earlier antibiotic administration guided by predictive scoring — is not supported by randomised evidence in the form in which it is usually stated. Blood culture requires sequential waits for growth, for isolation of pure colonies, for organism identification, and for susceptibility testing; identification methods in routine use depend on pure cultures, which makes the sequence irreducible rather than merely slow. In a substantial proportion of presentations the sequence yields no organism at all. Machine-learning approaches to earlier recognition have shown that discrimination transfers between institutions while net clinical utility does not, and a randomised trial of large-language-model assistance found no improvement in physician diagnostic reasoning despite strong standalone model performance.[2] An alternative is to model the physiological mechanism rather than the observed correlation. Mechanistic patient models are the maturest form of clinical simulation, are already routine in antimicrobial dose individualisation, and — unlike standalone learned models — have an established regulatory credibility pathway under ASME V&V 40 and associated agency guidance. Spaceflight provides the limiting case: molecular diagnostic capability has been demonstrated in orbit, no component of a sepsis workup has, immune function is measurably impaired in the arm that controls bacterial infection, and communication latency precludes real-time ground support. The requirements that follow from that limiting case apply equally to any setting where expert consultation is delayed or unavailable.
Keywords: sepsis; digital twin; mechanistic modelling; clinical decision support; diagnostic delay; space medicine
1. Introduction
Sepsis is defined as life-threatening organ dysfunction caused by a dysregulated host response to infection.[3] Global Burden of Disease estimates place the annual burden at approximately 48.9 million cases and 11 million deaths, corresponding to roughly one in five deaths worldwide.[1] These figures are widely reproduced. The mechanism that produces them is less often examined.
This paper makes three claims. First, that the delay between presentation and pathogen-specific diagnosis is a property of the diagnostic method rather than of laboratory throughput, and therefore cannot be resolved by process improvement. Second, that the conventional justification for accelerated diagnosis — that earlier antibiotic administration proportionally improves survival — is not supported by randomised evidence, and that a defensible justification runs through diagnostic precision rather than through elapsed time. Third, that mechanistic patient models address the problem that correlation-based prediction has not, and possess a regulatory advantage that has been underused in argument.
A methodological note governs the clinical material. Sections 2 and 3 draw on working sessions recorded in 2025 with clinicians and laboratory scientists preparing a European research proposal that was subsequently not submitted. Participants are unattributed by agreement. The sessions were conducted partly in Swedish; quoted material has been translated into English by the author, and passages where the recording is indistinct are identified as such rather than reconstructed. Claims originating in those sessions are identified in the text and are not presented as published evidence.
2. The structure of diagnostic delay
The path from presentation to microbiological diagnosis proceeds in four sequential stages, and the sequence is dependent rather than parallel.
A patient presents with altered mental state, frequently with fever. Confusion is a strong presenting signal, and its diagnostic weight increases in populations that cannot provide a history — the very old and the newborn. Nursing staff obtain blood and urine samples in the emergency department, and the samples are physically transported to a laboratory.
Basic inflammatory markers return within hours. A participant in one recorded session described her own recent emergency presentation for persistent tachycardia, in which samples were drawn, she was returned to the waiting area, and the results took six hours. That interval represents the fastest component of the pathway.
Blood culture constitutes the slow component. The sample is inoculated into general growth medium and incubated. Rapidly growing organisms are detectable within approximately one day; slower-growing bacteria require up to three; fungal organisms may require up to two weeks. Detection of growth, however, does not constitute identification.
Identification methods in routine clinical use, principally matrix-assisted laser desorption/ionisation time-of-flight mass spectrometry, depend on reference spectra derived from pure isolates. A clinical microbiologist described the constraint directly: where two organisms are present in one sample and the spectrum is matched against a single-organism reference, approximately half the peaks correspond and the remainder do not, producing either misidentification or no identification at all. Pure colonies must therefore be isolated before identification can proceed, adding approximately one further day. Antimicrobial susceptibility testing follows identification, adding further days again.
The clinical consequence is that four dependent waits accumulate while the underlying pathophysiological process continues. One participant characterised the pathway as a waterfall in which each step waits for the preceding step to complete; the microbiologist's summary was that the process waits until something has happened.
2.1 Diagnostic yield
The sequence described above frequently terminates without a result. The microbiologist estimated that in as many as half of cases no organism is recovered, returning the clinician to the position occupied before the samples were drawn. In an earlier recorded session the same participant applied the figure to meningitis, pneumonia and sepsis collectively, adding that in the absence of a microbiological diagnosis empirical broad-spectrum therapy is initiated and maintained, with two consequences: selection pressure for antimicrobial resistance, and the possibility that the causative agent is viral and the therapy therefore inactive.
The neonatal case is constrained by sample volume. Recommended blood volumes for adult blood culture are substantially greater than the total circulating volume of a low-birth-weight neonate, and reduced inoculum volume is an established determinant of reduced culture sensitivity.[4] The diagnostic test that would resolve the clinical question cannot be performed.
A participant summarised the resulting clinical posture as treatment based substantially on belief rather than on knowledge — a judgement formed from a limited marker panel in which C-reactive protein is prominent, and which, as the microbiologist noted, does not discriminate bacterial infection from inflammation, viral infection or fungal infection. The marker establishes that a process is present. It does not establish which process.
3. Reconsidering the time-to-treatment argument
The intuitive response to diagnostic delay is to accelerate treatment rather than diagnosis. The supporting evidence is frequently reduced to a single figure: an approximately 8% increase in mortality for each hour of delay in effective antimicrobial administration in septic shock.[5] The randomised evidence does not sustain the claim in that form.
The PHANTASi trial randomised prehospital antimicrobial administration by ambulance crews against usual care. Antibiotics were administered substantially earlier in the intervention arm without improvement in survival.[6] A systematic review and meta-analysis found no significant mortality benefit at either of the commonly applied timing thresholds.[7] The Infectious Diseases Society of America declined to endorse the one-hour bundle. An analysis of the New York State sepsis mandate found that longer time to completion of the three-hour bundle and longer time to antibiotic administration were associated with higher risk-adjusted in-hospital mortality, but found no corresponding association for the rapid intravenous fluid bolus contained within the same bundle.[8] The components of the bundle do not behave uniformly, and the evidence supports timing of one component rather than of the bundle as a construct.
The clinicians in the recorded sessions did not argue for speed. Their argument was that vital-sign monitoring establishes whether a patient is unwell but not what is causing the illness, and that targeted therapy requires the second determination. The defensible causal chain is therefore not that faster diagnosis produces faster antibiotics and thereby improved survival, but that faster and more certain organism identification permits narrower therapy, reduces selection pressure, reduces inappropriate treatment, and produces a patient who can be stratified for further intervention. The 2026 revision of the Surviving Sepsis Campaign guidelines moves explicitly toward risk-stratified rather than uniform recommendations.[9]
4. Why correlation-based prediction has underperformed
If diagnosis arrives too late, an apparent solution is to predict deterioration from data already available in the electronic record. This approach has been pursued for over a decade, and its results are instructive in both directions.
External validation of a widely deployed proprietary sepsis prediction model across 27,697 patients found an area under the receiver operating characteristic curve of 0.63, against vendor-reported performance of 0.76 to 0.83. At the recommended alert threshold the model identified 183 of 2,552 sepsis cases not already recognised by clinicians, missed 67% of cases, and generated alerts on 18% of all hospitalisations.[10] Alert burden at that level is associated with documented desensitisation effects.[11]
Selective citation of that result would reproduce the error this paper argues against. Prospective multi-site deployment of a machine-learning early-warning system across 590,736 patients was associated with an 18.7% adjusted relative reduction in in-hospital mortality where a clinician confirmed the alert within three hours.[12] Multi-site external validation of a separately developed model across 205,005 encounters retained an AUROC between 0.906 and 0.960.[13]
The reconciling observation is that discrimination transfers between institutions while clinical utility does not. A methodological review of 91 sepsis real-time prediction models found median AUROC declining from 0.886 at six hours before onset under internal validation to 0.783 under full-window external validation, and median utility score declining from +0.381 internally to −0.164 externally — that is, to net harm at the operating threshold.[14]
The strongest evidence that model capability does not transfer through the clinical interface comes from a randomised trial of large-language-model assistance in diagnostic reasoning. Physicians using the model achieved a median score of 76% against 74% with conventional resources, a difference of two percentage points with a confidence interval spanning zero. The model operating alone scored 92%.[2] Two documented mechanisms act in opposition to produce this result: automation bias, in which clinicians defer to the system,[15] and sycophancy, in which systems trained on human preference data defer to the user.[16] The composite behaviour converges on the clinician's prior, which is the failure mode identified in section 2.1.
4.1 The label problem
A structural explanation underlies these performance ceilings. Sepsis prediction models are trained against a consensus clinical construct that was itself revised in 2016[3] and is operationalised differently across institutions and coding practices. The training target is a definition, not a disease state, and a model that learns the definition inherits its instability.
A multi-centre retrospective analysis across three harmonised intensive-care cohorts comprising 216,536 stays found that fine-tuning — the standard remedy for distribution shift — consistently underperformed retraining, fusion training and domain adaptation.[17] Adaptation cannot recover a representation that was fitted to the wrong quantity.
5. Mechanistic patient models
The alternative is to represent the physiology that generates the observations and to fit that representation to an individual patient.
This approach is established rather than speculative. Patient-specific cardiac models constructed from imaging have guided ablation for infarct-related ventricular tachycardia, with simulation-targeted lesions substantially smaller than those delivered conventionally, and the method has proceeded to trial under a regulatory Investigational Device Exemption.[18] Population-of-models methods generate large candidate parameter sets and filter them against experimental data to construct virtual patient cohorts.[19] Whole-body physiological simulation environments spanning cardiovascular, respiratory, renal, endocrine and metabolic systems have been maintained for over a decade with model structure represented in data rather than compiled into code.[20]
For sepsis specifically, mechanistic modelling predates the machine-learning literature. Agent-based models of the innate inflammatory response were used to reproduce in silico the failure of cytokine-directed clinical trials — a result concerning the response to intervention, which correlation cannot address.[21,22] Reduced ordinary-differential-equation models of acute inflammation reproduce the observed clinical trajectories from a small number of coupled equations.[23]
Mechanistic modelling has already entered routine sepsis care without being described as such. The DALI study, conducted across more than 500 patients in over 60 European intensive care units, found that conventional beta-lactam dosing fails to attain pharmacokinetic/pharmacodynamic targets in a majority of critically ill patients.[24] Model-informed dosing with therapeutic drug monitoring in critical illness is now the subject of a joint position paper endorsed by the relevant European professional societies.[25] The administered dose is calculated from a model of the individual patient.
5.1 The regulatory asymmetry
A mature framework exists for establishing the credibility of computational models in medical device submissions. ASME V&V 40, developed jointly by device manufacturers, simulation vendors and the United States Food and Drug Administration, specifies a risk-informed workflow keyed to a declared context of use and a model risk assessment across thirteen credibility factors.[26] The FDA operationalised the standard in guidance issued 16 November 2023.[27]
That guidance applies to physics-based and mechanistic models and explicitly excludes standalone machine-learning models. A mechanistic patient model therefore has a documented route to regulatory credibility that a correlation-based predictor does not. The European position is developing along comparable lines, with mechanistic models routed through the European Medicines Agency qualification of novel methodologies procedure, and with verification, validation and uncertainty quantification against a declared context of use established as the expected methodological framing for in-silico evidence.[28]
The timetable is material. Regulation (EU) 2024/1689 classifies artificial intelligence forming a safety component of a regulated medical device as high-risk under Article 6(1).[29] Regulation (EU) 2026/1744, published 24 July 2026, defers those obligations to 2 August 2028 for the medical-device route and to 2 December 2027 for standalone high-risk systems, while transparency obligations remain applicable from 2 August 2026.[30] Acknowledged notified-body capacity constraints are unaffected by the deferral.[31]
5.2 Limitations of the mechanistic argument
Three results constrain any general claim that biological modelling requires mechanism rather than statistics.
Decades of physics-based protein structure prediction were superseded by a learned model.[32] A protein language model trained on sequence alone, without alignment or physical priors, subsequently achieved comparable accuracy at substantially greater speed.[33] In numerical weather prediction — the canonical success of mechanistic simulation — a learned graph neural network outperformed the leading operational system on 90% of 1,380 verification targets.[34]
The defensible position is narrower. Statistical models reproduce the distribution on which they were trained, and fail characteristically on dynamics, alternative states and response to perturbation. Evaluated against 98 fold-switching proteins, structure prediction produced models corresponding to a single conformation for 81% of sequences.[35] The model reproduced the commonly observed state and did not represent the existence of the alternative.
Sepsis occupies that failure regime. The clinically operative question is not the patient's present classification but the predicted response to a specific intervention, which is a counterfactual. The most direct supporting evidence is a randomised trial stratifying patients by immune phenotype rather than treating sepsis as a single condition, which reported a difference on the organ-dysfunction endpoint of 35.1% against 17.9% (p = 0.002). The result requires qualification: twenty-eight-day mortality did not differ significantly, the endpoint is a surrogate, the trial is phase II, and serious adverse events occurred in 88.8% of participants.[36] It constitutes a first randomised signal that stratification by mechanism alters treatment response, not a demonstration of clinical benefit.
5.3 A terminological correction
The author has previously used the term Large Quantitative Model inconsistently across two published papers, defining it in one as a physics-based engine computing from equations rather than fitted correlation, and in the other as a model trained at cloud scale and contrasted with smaller models executing at the network edge.
These are not competing definitions of one property but two independent axes. The first is epistemological and concerns whether the model computes from mechanism or fits from correlation. The second is architectural and concerns where the model executes. A model may be mechanistic and computationally large, or mechanistic and small enough to execute on constrained hardware. The second combination is what the clinical case requires: mechanistic enough to evaluate a counterfactual, and small enough to execute where the patient is. The term Small Quantitative Model is defined here as that combination, and required both axes to be stated before it carried meaning.
6. The limiting case: clinical reasoning without consultation
Spaceflight alters immune function measurably, and in the direction that impairs control of bacterial infection. Natural killer cell cytotoxic function declines by approximately half by flight day 90 while cell numbers and receptor expression remain unchanged, a dissociation that a conventional differential count would not detect.[37] Neutrophil counts rise at landing while optimal chemotactic response declines approximately tenfold.[38] Alterations in adaptive immunity persist throughout six-month missions,[39] and eighteen cytokines and chemokines changed measurably following a three-day flight.[40]
Latent herpesvirus reactivation is well documented, with at least one virus shed by 53% of Space Shuttle and 61% of International Space Station crew members, and with reactivation increasing in frequency, duration and amplitude on longer missions.[41,42] Culturable infectious varicella zoster virus has been recovered from asymptomatic astronauts.[43]
The microbiological literature requires careful reading. Salmonella Typhimurium flown on Space Shuttle showed altered expression of 167 transcripts and increased virulence in a murine model.[44] Pseudomonas aeruginosa formed thicker biofilms in orbit with a column-and-canopy architecture not observed at normal gravity,[45] which is the plausible route to device-associated infection aboard a spacecraft. Escherichia coli cultured in space grew in gentamicin concentrations that are inhibitory at normal gravity.[46] However, under identical modelled microgravity with steam rather than antibiotic sterilisation, organisms failed to acquire resistance to any of twenty antibiotics tested; resistance emerged only under antibiotic selective pressure.[47] The most comprehensive station-wide survey concluded that the genomic and physiological features selected by International Space Station conditions do not appear directly relevant to human health.[48] Spaceflight measurably alters microbial behaviour; it has not been demonstrated to generate clinically resistant pathogens, and the four findings should be cited together.
6.1 The diagnostic asymmetry
Molecular diagnostic capability has been demonstrated in orbit. Nanopore sequencing aboard the International Space Station produced 276,882 reads across nine runs over six months without performance decay, and assembled an E. coli genome de novo into a single contig covering 99.9% of the reference.[49] Microbial identification has been performed entirely off Earth from sample to result.[50] Miniaturised polymerase chain reaction has been demonstrated in flight,[51] as has point-of-care blood chemistry.[52]
No component of a sepsis workup has flown. Lactate trending, blood culture, procalcitonin and functional differential counts are absent from the demonstrated in-flight capability set.
The resulting configuration for an exploration-class mission is therefore: demonstrated molecular tooling, no sepsis diagnostic pathway, immune function impaired in the relevant effector arm, a pharmacy whose stability data are limited and internally inconsistent — one study found 9 of 35 formulations meeting content criteria after up to 28 months in orbit independent of labelled expiry,[53] while a second found no unusual degradation across 9 medications after 550 days[54] — and a one-way communication latency to Earth precluding real-time consultation. Culture is unavailable, waiting is unavailable, consultation is unavailable, and pharmacokinetics are poorly characterised. A locally executing model of the patient capable of evaluating a counterfactual is the remaining candidate.
This is recognised in the operational literature. NASA exploration medical work frames the requirement as Earth-independent medical operations.[55] Observed in-flight clinical incidence has been characterised at 3.40 events per flight-year across 46 astronauts and 20.57 flight-years, with 46% of crew reporting a notable event.[56] An in-flight venous thrombosis has been diagnosed in orbit and managed with the available formulary through telemedicine.[57] Clinical decision-making without a physician present is not prospective.
6.2 Terrestrial transfer
Transfer from spaceflight medicine to terrestrial practice is documented rather than asserted. Rotating-wall vessel hardware developed to model microgravity became a mainstream platform for three-dimensional organotypic host–pathogen culture.[58] Remotely guided ultrasound performed by a minimally trained operator aboard the International Space Station, completing a focused assessment examination in approximately five and a half minutes, is directly ancestral to remote imaging protocols.[59] Microgravity crystallisation of a monoclonal antibody produced a uniform suspension where the ground control was bimodal, and the process was subsequently translated into terrestrial manufacturing.[60]
Each constraint described in section 6.1 is a sharpened form of a constraint present somewhere in terrestrial practice: a district hospital without on-site microbiology, an Antarctic overwintering station where immune activation escalates measurably after three to four months of isolation,[61] a submarine, a field hospital, or a neonatal unit unable to obtain sufficient sample volume.
7. Design requirements
Seven requirements follow, stated in the order in which they constrain design.
Models should represent physiological mechanism rather than the consensus label, because a model fitted to a syndrome definition inherits the instability of that definition and cannot be inspected for the reason it is wrong. The context of use and the associated model risk should be declared before the model is constructed, following ASME V&V 40 and the corresponding agency guidance, because that declaration is what renders a model falsifiable. Models should be small enough to execute where the patient is, since a model requiring data-centre resources is unavailable in precisely the circumstances that justify it. Every input should carry its declared staleness, because a patient in septic trajectory changes materially within hours and a simulation executed on twelve-hour-old inputs is not a delayed result but an incorrect one presented with the confidence of a current one. Degraded-mode behaviour should be specified as a primary design output rather than a fallback, because for exploration missions the degraded mode is the operating mode and for an under-resourced hospital at night it is also the operating mode. Systems should evaluate counterfactuals rather than produce classifications, since the operative clinical question concerns predicted response to a specific intervention and the evidence in section 4 indicates that classification systems have not altered outcomes at the point of care. Finally, attribution must be preserved through the full chain, because every preceding requirement depends on the ability to establish which model version, operating on which inputs, produced a given recommendation, and under whose authority it was accepted.
8. Conclusion
A concurrent engineering programme established at the Jet Propulsion Laboratory in 1995 reduced early space mission design from approximately nine months to approximately three weeks in elapsed time, with the engineering work itself conducted in approximately nine hours of facilitated sessions distributed across three days. The reported efficiency gain is less than one tenth of the previous elapsed duration at less than one third of the variable cost.[62] The mechanism was co-location, shared display of a common model, and networked models updating in real time, rather than automation of individual tasks.
The same analysis models the distinction quantitatively. For a task of forty weeks' duration, reducing every constituent subtask to a one-minute decision without addressing response latency yields a total duration of twenty weeks. Reducing coordination latency to one minute yields three hours. Automation halves the duration; removal of waiting collapses it.[62]
The author has previously cited this result incorrectly, compressing the elapsed and session figures into a single comparison and attaching an improvement of 99% to it. The correction is stated here because the error is instructive: the bottleneck in a complex process is rarely located where activity is visible, and is characteristically located in the intervals between activities, where the process appears idle and time is nevertheless consumed.
A patient in an emergency department is not waiting for a laboratory to operate more efficiently. The laboratory operates at the rate its method permits. The patient is waiting for a method that does not require the disease to declare itself first. That is the function of a mechanistic patient model: not a faster test, but a form of clinical reasoning that does not begin by waiting.
Conflicts of interest
The author is founder and chief executive of WINNIIO AB, which develops digital-twin and clinical decision-support systems, and therefore has a commercial interest in the class of approach advocated in section 5. The 2025 working sessions cited in sections 2 and 3 were conducted in preparation for a European research proposal in which the author would have participated. That proposal was not submitted.
Data availability
No primary data are reported. Recordings of the 2025 working sessions are not available for release, as participants consented to unattributed use only.
References
1. Rudd KE, Johnson SC, Agesa KM, et al. Global, regional, and national sepsis incidence and mortality, 1990–2017: analysis for the Global Burden of Disease Study. Lancet. 2020;395(10219):200–211. doi:10.1016/S0140-6736(19)32989-7 2. Goh E, Gallo R, Hom J, et al. Large language model influence on diagnostic reasoning: a randomized clinical trial. JAMA Netw Open. 2024;7(10):e2440969. doi:10.1001/jamanetworkopen.2024.40969 3. Singer M, Deutschman CS, Seymour CW, et al. The Third International Consensus Definitions for Sepsis and Septic Shock (Sepsis-3). JAMA. 2016;315(8):801–810. doi:10.1001/jama.2016.0287 4. Connell TG, Rele M, Cowley D, Buttery JP, Curtis N. How reliable is a negative blood culture result? Volume of blood submitted for culture in routine practice in a children's hospital. Pediatrics. 2007;119(5):891–896. doi:10.1542/peds.2006-0440 5. Kumar A, Roberts D, Wood KE, et al. Duration of hypotension before initiation of effective antimicrobial therapy is the critical determinant of survival in human septic shock. Crit Care Med. 2006;34(6):1589–1596. doi:10.1097/01.CCM.0000217961.75225.E9 6. Alam N, Oskam E, Stassen PM, et al. Prehospital antibiotics in the ambulance for sepsis: a multicentre, open label, randomised trial. Lancet Respir Med. 2018;6(1):40–50. doi:10.1016/S2213-2600(17)30469-1 7. Sterling SA, Miller WR, Pryor J, Puskarich MA, Jones AE. The impact of timing of antibiotics on outcomes in severe sepsis and septic shock: a systematic review and meta-analysis. Crit Care Med. 2015;43(9):1907–1915. doi:10.1097/CCM.0000000000001142 8. Seymour CW, Gesten F, Prescott HC, et al. Time to treatment and mortality during mandated emergency care for sepsis. N Engl J Med. 2017;376(23):2235–2244. doi:10.1056/NEJMoa1703058 9. Prescott HC, et al. Surviving Sepsis Campaign: international guidelines for management of sepsis and septic shock 2026. Crit Care Med. 2026;54(4):725–812. 10. Wong A, Otles E, Donnelly JP, et al. External validation of a widely implemented proprietary sepsis prediction model in hospitalized patients. JAMA Intern Med. 2021;181(8):1065–1070. doi:10.1001/jamainternmed.2021.2626 11. Ancker JS, Edwards A, Nosal S, Hauser D, Mauer E, Kaushal R. Effects of workload, work complexity, and repeated alerts on alert fatigue in a clinical decision support system. BMC Med Inform Decis Mak. 2017;17:36. doi:10.1186/s12911-017-0430-8 12. Adams R, Henry KE, Sridharan A, et al. Prospective, multi-site study of patient outcomes after implementation of the TREWS machine learning-based early warning system for sepsis. Nat Med. 2022;28(7):1455–1460. doi:10.1038/s41591-022-01894-0 13. Evaluating Sepsis Watch generalizability through multisite external validation of a sepsis machine learning model. npj Digit Med. 2025;8. doi:10.1038/s41746-025-01664-5 14. Wang Z, Wang W, Sun C, et al. A methodological systematic review of validation and performance of sepsis real-time prediction models. npj Digit Med. 2025;8. doi:10.1038/s41746-025-01587-1 15. Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. J Am Med Inform Assoc. 2012;19(1):121–127. doi:10.1136/amiajnl-2011-000089 16. Sharma M, Tong M, Korbak T, et al. Towards understanding sycophancy in language models. arXiv:2310.13548. Published as conference paper, ICLR 2024. 17. Evaluating deep learning sepsis prediction models in ICUs under distribution shift: a multi-centre retrospective cohort study. npj Digit Med. 2026. doi:10.1038/s41746-026-02364-4 18. Prakosa A, Arevalo HJ, Deng D, et al. Personalized virtual-heart technology for guiding the ablation of infarct-related ventricular tachycardia. Nat Biomed Eng. 2018;2:732–740. doi:10.1038/s41551-018-0282-2 19. Britton OJ, Bueno-Orovio A, Van Ammel K, et al. Experimentally calibrated population of models predicts and explains intersubject variability in cardiac cellular electrophysiology. Proc Natl Acad Sci USA. 2013;110(23):E2098–E2105. doi:10.1073/pnas.1304382110 20. Hester RL, Brown AJ, Husband L, et al. HumMod: a modeling environment for the simulation of integrative human physiology. Front Physiol. 2011;2:12. doi:10.3389/fphys.2011.00012 21. An G. In silico experiments of existing and hypothetical cytokine-directed clinical trials using agent-based modeling. Crit Care Med. 2004;32(10):2050–2060. doi:10.1097/01.ccm.0000139707.13729.7d 22. Vodovotz Y, An G. Agent-based models of inflammation in translational systems biology: a decade later. WIREs Syst Biol Med. 2019;11(6):e1460. doi:10.1002/wsbm.1460 23. Reynolds A, Rubin J, Clermont G, Day J, Vodovotz Y, Ermentrout GB. A reduced mathematical model of the acute inflammatory response: I. Derivation of model and analysis of anti-inflammation. J Theor Biol. 2006;242(1):220–236. doi:10.1016/j.jtbi.2006.02.016 24. Roberts JA, Paul SK, Akova M, et al. DALI: defining antibiotic levels in intensive care unit patients — are current β-lactam antibiotic doses sufficient for critically ill patients? Clin Infect Dis. 2014;58(8):1072–1083. doi:10.1093/cid/ciu027 25. Abdul-Aziz MH, Alffenaar JWC, Bassetti M, et al. Antimicrobial therapeutic drug monitoring in critically ill adult patients: a position paper. Intensive Care Med. 2020;46(6):1127–1153. doi:10.1007/s00134-020-06050-1 26. American Society of Mechanical Engineers. V&V 40-2018: Assessing Credibility of Computational Modeling through Verification and Validation — Application to Medical Devices. New York: ASME; 2018. 27. US Food and Drug Administration. Assessing the Credibility of Computational Modeling and Simulation in Medical Device Submissions: Guidance for Industry and Food and Drug Administration Staff. CDRH; issued 16 November 2023. Docket FDA-2021-D-0980. 28. Viceconti M, Pappalardo F, Rodriguez B, Horner M, Bischoff J, Musuamba Tshinanu F. In silico trials: verification, validation and uncertainty quantification of predictive models used in the regulatory evaluation of biomedical products. Methods. 2021;185:120–127. doi:10.1016/j.ymeth.2020.01.011 29. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. OJ L, 2024/1689, 12.7.2024. 30. Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 amending Regulation (EU) 2024/1689 as regards simplification of implementation. OJ L, 2026/1744, 24.7.2026. 31. Medical Device Coordination Group. MDCG 2022-14: Transition to the MDR and IVDR — Notified Body Capacity and Availability of Medical Devices and IVDs. European Commission; August 2022. 32. Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi:10.1038/s41586-021-03819-2 33. Lin Z, Akin H, Rao R, et al. Evolutionary-scale prediction of atomic-level protein structure with a language model. Science. 2023;379(6637):1123–1130. doi:10.1126/science.ade2574 34. Lam R, Sanchez-Gonzalez A, Willson M, et al. Learning skillful medium-range global weather forecasting. Science. 2023;382(6677):1416–1421. doi:10.1126/science.adi2336 35. Chakravarty D, Porter LL. AlphaFold2 fails to predict protein fold switching. Protein Sci. 2022;31(6):e4353. doi:10.1002/pro.4353 36. ImmunoSep investigators. Personalised immunotherapy in sepsis: a randomised clinical trial. JAMA. 2026. 37. Bigley AB, Agha NH, Baker FL, et al. NK cell function is impaired during long-duration spaceflight. J Appl Physiol. 2019;126(4):842–853. doi:10.1152/japplphysiol.00761.2018 38. Stowe RP, Sams CF, Mehta SK, et al. Leukocyte subsets and neutrophil function after short-term spaceflight. J Leukoc Biol. 1999;65(2):179–186. doi:10.1002/jlb.65.2.179 39. Crucian B, Stowe RP, Mehta S, et al. Alterations in adaptive immunity persist during long-duration spaceflight. npj Microgravity. 2015;1:15013. doi:10.1038/npjmgrav.2015.13 40. Kim J, Tierney BT, Overbey EG, et al. Single-cell multi-ome and immune profiles of the Inspiration4 crew reveal conserved, cell-type, and sex-specific responses to spaceflight. Nat Commun. 2024;15(1):4954. doi:10.1038/s41467-024-49211-2 41. Rooney BV, Crucian BE, Pierson DL, Laudenslager ML, Mehta SK. Herpes virus reactivation in astronauts during spaceflight and its application on Earth. Front Microbiol. 2019;10:16. doi:10.3389/fmicb.2019.00016 42. Mehta SK, Laudenslager ML, Stowe RP, et al. Latent virus reactivation in astronauts on the International Space Station. npj Microgravity. 2017;3:11. doi:10.1038/s41526-017-0015-y 43. Cohrs RJ, Mehta SK, Schmid DS, Gilden DH, Pierson DL. Asymptomatic reactivation and shed of infectious varicella zoster virus in astronauts. J Med Virol. 2008;80(6):1116–1122. doi:10.1002/jmv.21173 44. Wilson JW, Ott CM, Höner zu Bentrup K, et al. Space flight alters bacterial gene expression and virulence and reveals a role for global regulator Hfq. Proc Natl Acad Sci USA. 2007;104(41):16299–16304. doi:10.1073/pnas.0707155104 45. Kim W, Tengra FK, Young Z, et al. Spaceflight promotes biofilm formation by Pseudomonas aeruginosa. PLoS One. 2013;8(4):e62437. doi:10.1371/journal.pone.0062437 46. Zea L, Larsen M, Estante F, et al. Phenotypic changes exhibited by E. coli cultured in space. Front Microbiol. 2017;8:1598. doi:10.3389/fmicb.2017.01598 47. Tirumalai MR, Karouia F, Tran Q, et al. Evaluation of acquired antibiotic resistance in Escherichia coli exposed to long-term low-shear modeled microgravity and background antibiotic exposure. mBio. 2019;10(1):e02637-18. doi:10.1128/mBio.02637-18 48. Mora M, Wink L, Kögler I, et al. Space Station conditions are selective but do not alter microbial characteristics relevant to human health. Nat Commun. 2019;10(1):3990. doi:10.1038/s41467-019-11682-z 49. Castro-Wallace SL, Chiu CY, John KK, et al. Nanopore DNA sequencing and genome assembly on the International Space Station. Sci Rep. 2017;7(1):18022. doi:10.1038/s41598-017-18364-0 50. Burton AS, Stahl SE, John KK, et al. Off Earth identification of bacterial populations using 16S rDNA nanopore sequencing. Genes. 2020;11(1):76. doi:10.3390/genes11010076 51. Boguraev AS, Christensen HC, Bonneau AR, et al. Successful amplification of DNA aboard the International Space Station. npj Microgravity. 2017;3:26. doi:10.1038/s41526-017-0033-9 52. Smith SM, Davis-Street JE, Fontenot TB, Lane HW. Assessment of a portable clinical blood analyzer during space flight. Clin Chem. 1997;43(6 Pt 1):1056–1065. 53. Du B, Daniels VR, Vaksman Z, Boyd JL, Crady C, Putcha L. Evaluation of physical and chemical changes in pharmaceuticals flown on space missions. AAPS J. 2011;13(2):299–308. doi:10.1208/s12248-011-9270-0 54. Wotring VE. Chemical potency and degradation products of medications stored over 550 Earth days at the International Space Station. AAPS J. 2016;18(1):210–216. doi:10.1208/s12248-015-9834-5 55. Russell BK, et al. The value of a spaceflight clinical decision support system for earth-independent medical operations. npj Microgravity. 2023;9:46. doi:10.1038/s41526-023-00284-1 56. Crucian B, Babiak-Vazquez A, Johnston S, Pierson DL, Ott CM, Sams C. Incidence of clinical symptoms during long-duration orbital spaceflight. Int J Gen Med. 2016;9:383–391. doi:10.2147/IJGM.S114188 57. Auñón-Chancellor SM, Pattarini JM, Moll S, Sargsyan A. Venous thrombosis during spaceflight. N Engl J Med. 2020;382(1):89–90. doi:10.1056/NEJMc1905875 58. Barrila J, Radtke AL, Crabbé A, et al. Organotypic 3D cell culture models: using the rotating wall vessel to study host–pathogen interactions. Nat Rev Microbiol. 2010;8(11):791–801. doi:10.1038/nrmicro2423 59. Sargsyan AE, Hamilton DR, Jones JA, et al. FAST at MACH 20: clinical ultrasound aboard the International Space Station. J Trauma. 2005;58(1):35–39. doi:10.1097/01.ta.0000145083.47032.78 60. Reichert P, Prosise W, Fischmann TO, et al. Pembrolizumab microgravity crystallization experimentation. npj Microgravity. 2019;5:28. doi:10.1038/s41526-019-0090-3 61. Feuerecker M, Crucian BE, Quintens R, et al. Immune sensitization during 1 year in the Antarctic high-altitude Concordia environment. Allergy. 2019;74(1):64–77. doi:10.1111/all.13545 62. Chachere J, Kunz J, Levitt R. The Role of Reduced Latency in Integrated Concurrent Engineering. CIFE Working Paper WP116. Stanford University; 2009. https://purl.stanford.edu/bd089dx8723
Methodology referenced: SMILE v6.4.3. doi:10.5281/zenodo.21757691
Where this goes next
Want this applied to your organisation?
One call is enough to know if we're a fit.