HealthFlux: Stanford and Harvard built a “world model” of human health that predicts the risk of 195 diseases more accurately than earlier models, outperforms clinical cardiovascular and mortality risk scores, and maintains predictive performance for ten years without new tests
HealthFlux: Stanford and Harvard built a “world model” of human health that predicts the risk of 195 diseases more accurately than earlier models, outperforms clinical cardiovascular and mortality risk scores, and maintains predictive performance for ten years without new tests
On September 21, Stanford researcher Kejun Ying, neurologist Tony Wyss-Coray, and colleagues from Harvard published a preprint describing HealthFlux. Six days later, Ying shared the work on X. The system combines medical records, laboratory tests, genetic data, plasma proteins, and brain and body scans from 502 166 UK Biobank participants into a single, continually updated representation of health, which it uses to estimate the risk of 195 diseases and death years in advance.
Health cannot be measured directly. A blood test, diagnosis, or MRI captures only a single moment, with no observations between visits. HealthFlux maintains a latent, 384-dimensional health state that evolves continuously with age between visits, much as GPS tracks a car between satellite signals, and is updated abruptly when a new test result arrives. This hybrid approach accommodates measurements taken at different intervals: blood tests every few months, MRI scans every few years, and diagnoses whenever a condition is clinically recognized. The model was trained to predict which event would occur next and when, much as a language model predicts the next word, rather than to predict diagnoses directly. A separate component was then trained to estimate disease risk from the resulting health state.
The dataset comprised 502 166 UK Biobank participants. The model received 5 647 variables from 11 sources, including clinical records, blood tests, physical measurements, lifestyle data, 2 923 plasma proteins, 249 metabolites, brain and body MRI scans, and genetic data.
HealthFlux predicts 195 diseases and death over five years with an AUROC of 0,816 (0,5 indicates chance performance; 1,0 indicates perfect prediction). This exceeds the previous best model, Delphi-2M (0,715, Nature, 2025), and established clinical risk scores. For mortality, the Charlson score achieved 0,754 compared with 0,846 for HealthFlux. For cardiovascular risk, QRISK3, SCORE2, and Framingham achieved 0,707 to 0,743 compared with 0,776. When one disease was excluded entirely from training, the system still predicted its risk with an AUROC of 0,769. Its ability to predict a disease it had never been trained on indicates that it learned general patterns of health over time rather than simply memorizing diagnoses. Health states simulated 5 to 10 years into the future without new test results continued to predict disease, and the performance advantage over Delphi-2M grew over time.
Combining data sources captures cases missed by models that use a single type of measurement. Among people aged 65 to 75 with hypertension, HealthFlux identified 27,9 to 55,5% of cases missed by models using only one data type, whether medical history, genetic profiles, or another individual source.
Ying, for whom HealthFlux was his first paper as senior author, described the ability to generalize to previously unseen diseases as the finding he was most proud of:
“A general representation of health allows us to predict conditions the model was never trained to predict.”
Ying wrote this on X on September 27. In the same post, he proposed a more ambitious use: simulate two futures, one with a drug and one without it, and compare disease risk. His reasoning was that a real trial takes years, while a simulation takes minutes. This could allow researchers to explore new uses for existing drugs and preventive interventions before undertaking an expensive human trial.
Ten days before Ying’s post, the scientific journal Cell published a review that, for the first time, identified “world models” as a direction for biomedical AI. The competing model RisQ appeared 2,5 months earlier and trails HealthFlux on every metric. The coauthors include Stanford neurologist Tony Wyss-Coray and Harvard’s Vadim Gladyshev. Six months earlier, in a roadmap for tissue and cell replacement, they wrote that a measure representing the whole body was missing. HealthFlux’s 384-dimensional health state now provides that representation.