A GPT-4o AI agent independently scored the severity of all 10,211 human genetic disease phenotypes from the medical literature, and its list of severe genes matched 95% of the genes in Australia's national carrier screening program Mackenzie's Mission
A GPT-4o AI agent independently scored the severity of all 10,211 human genetic disease phenotypes from the medical literature, and its list of severe genes matched 95% of the genes in Australia's national carrier screening program Mackenzie's Mission
On September 17, a team at the University of New South Wales in Sydney (UNSW) posted a preprint describing an autonomous AI agent that searches PubMed, the database of biomedical scientific publications, and scores the severity of genetic phenotypes according to the criteria of the American College of Medical Genetics and Genomics (ACMG). The agent classified all 10,211 terms in the Human Phenotype Ontology, the international clinical phenotype vocabulary, and aggregated scores up to the level of 8738 gene–disease pairs.
For a gene to be included in a carrier screening panel for prospective parents, ACMG requires two conditions to be met: the gene must be common enough among healthy carriers of a single copy (more than one in 200 people, a frequency that existing genomic databases already allow to be calculated), and a child who inherits a damaged copy from each of two carrier parents must develop a severe disease. The second condition has been assessed manually by committees of clinical geneticists for decades. Commercial screening panels have varied in scope from 41 to 1792 conditions, and when 12 geneticists independently rated the severity of 176 genes by hand in 2020, they agreed with one another in only 60.8% of cases.
A raw language model answer with no source references cannot be verified in a clinical setting, so the UNSW team built their agent on the ReAct (reasoning and acting) framework: the agent searches PubMed on its own and must cite a source for every conclusion, with no recourse to the model's internal knowledge. After benchmarking six models, including GPT-5 and Grok-4, the authors chose GPT-4o for its 92% accuracy and the best balance of predictability and reasoning flexibility. A separate verifier agent checks each statement against its cited source and assigns an evidence weight; because of low confidence, the agent itself flagged 1161 out of 10,211 assessments for human review.
On a sample of 941 phenotypes labeled independently by two clinical geneticists, the agent agreed with their ratings in 93.55% of cases. At the gene level, the agent classified 79.5% of the 8738 gene–disease pairs as severe or profound. Its list of severe genes was then compared with the genes in Australia's Mackenzie's Mission program, which has already been tested on thousands of prospective parent couples and was described in the New England Journal of Medicine in 2024: the overlap was 95.2%, and with the ACMG's own carrier screening panel the overlap was 99.3%. Processing a single phenotype costs roughly ten cents; the entire database of 10,211 terms cost approximately one thousand dollars, a task that has taken expert committees years to accomplish.
Of 29 genes with independently published severity ratings, the agent reproduced the published rating in 20 cases, assigned a stricter category than the source in eight, and underestimated only one gene, BCKDHB, rating its severity as "moderate" instead of "profound": the reference databases simply did not capture intellectual disability as a manifestation of this gene. The agent could not assign a severity tier to roughly a third of all terms; these were mostly phenotypes such as headache that fall outside the inherited conditions scale to begin with.
The technology is owned by UNSW Sydney and partially licensed to 23Strands: the company's director is a co-author of the paper, and the first author receives a scholarship top-up from the same company. This kind of severity assessment, inexpensive and already validated against an operational screening program, removes the bottleneck that has constrained carrier screening panels for prospective parents.