Kristina Emilie Sørensen Proposes Assessing AI Biological Risk Across the Full Workflow of a Trained Biologist
Kristina Emilie Sørensen Proposes Assessing AI Biological Risk Across the Full Workflow of a Trained Biologist
On July 25, Kristina Emilie Sørensen published an essay on the biological risks of language models. She proposes assessing a model together with the actions of a trained biologist, including DNA design, synthesis, laboratory work, and organizational review.
Biological risk assessments of language models often focus on a conversation: a researcher submits a dangerous request, and the developer counts the model's refusals. This type of test compares how different model versions behave in dialogue. Sørensen proposes also tracing the subsequent actions that turn the information obtained into a biological result.
In a June report from the Danish Centre for Biosecurity and Biopreparedness, AI forms part of a long sequence of actions. A trained specialist uses software to predict the effect of a genetic modification, design a DNA sequence, and evaluate its properties before synthesis. After synthesis, the specialist needs materials, equipment, and laboratory procedures. Screening of synthetic DNA orders already checks both the sequence and the customer. Earlier filters were less effective at recognizing proteins with novel sequences that retained their original function.
Sørensen argues that the biologist's objective gives meaning to a hundred ordinary questions. These questions may concern a protein, an expression system, or sample stability. The objective has already been chosen before the dialogue begins, so the filter sees only the wording of each question, while the workflow emerges from the person's subsequent decisions.
Sørensen proposes measuring this sequence. In her model, a language model helps a trained specialist explore options more quickly. The next stages involve ordering DNA, gaining access to equipment, working with laboratory staff, and following internal rules. These actions have different points of control: the supplier screens the order, the laboratory controls access to instruments, and the organization reviews an unusual project. The assessment connects the model's response to the actions and resources that make the next step possible.
In a 2023 MIT exercise, students without relevant training used chatbots for an hour and obtained information that the authors considered dangerous. In a 2024 RAND study, teams developed plans for a biological attack with and without a language model. RAND found no statistically significant difference in the viability of the plans. These studies measured different tasks and outcomes. A single score for a model's “danger” combines the conversation with the model, the planning process, and the material work that follows.
Research laboratories need software for designing genetic interventions and proteins, including laboratories seeking new therapies. In Sørensen's model, risk assessment begins with a question: what next step can a person take with the model's help, and where does that step leave a trace that can be reviewed?