Nature Biomedical Engineering review proposes three levels of evidence for claims about the causal effects of medical AI
Nature Biomedical Engineering review proposes three levels of evidence for claims about the causal effects of medical AI
On 7 August, in a review of causal graph neural networks in healthcare, the authors distinguished among systems that identify stable associations in data, models with explicit causal assumptions, and findings supported by validated causal inference.
A hospital leaves traces of its own operations in its data, including scanner settings, the order of examinations, and the composition of its patient population. In a study of chest radiographs from three hospital systems, a model designed to detect signs of pneumonia was trained on 158 323 images. Its AUC was 0.931 in the pooled internal evaluation. AUC measures how well an algorithm distinguishes images with pneumonia from those without it. In the external dataset from the third hospital system, the AUC fell to 0.815. A separate network identified the hospital system from an image with almost no errors. When pneumonia prevalence differed among hospitals, this hospital-specific signal helped the model infer the diagnosis within the source dataset.
The problem can arise at an earlier stage. In an analysis of virtual cell benchmarks, models tested on their ability to predict the effects of an intervention sometimes performed worse than the mean profile. A standard metric barely registered the several dozen genes altered by the intervention against a background of roughly 20 thousand other genes. Such an evaluation does not establish that the model has captured even the specific effect of the intervention. The new review sets a higher standard: the model must show that the effect has a causal explanation and transfers across clinical contexts.
A graph neural network represents genes, proteins, drugs, and symptoms as nodes in a single network and accounts for the relationships among them. A causal model adds a testable hypothesis about which factors affect which outcomes. Given its assumptions, the model can simulate a change in the prescribed drug and estimate the expected consequences.
“The question now facing medical AI is not whether a system can achieve high accuracy in retrospective tests, but whether it can withstand clinical use across institutions while preserving effectiveness and fairness.”
The authors divide causal claims into three levels. At the first level, the architecture uses causal ideas, for example by seeking features that remain stable when the hospital changes, but its output is still a prediction. At the second level, the model estimates the effect of an intervention using an explicitly defined causal structure. At the third level, a new causal finding must be consistent with a biological mechanism, reproduce in at least two independent patient groups, undergo experimental or prospective validation on new data, and remain robust to hidden factors that could reverse the conclusion.
The authors use the term digital twin for a future computational model of an individual patient. It would need to combine the patient’s medical history, imaging, molecular data, and known biological relationships to compare the expected outcomes of different interventions. For such a model, the three-level framework sets the requirements: a causal hypothesis about the patient, independent data, and validation through an intervention.