In a cell-based experiment, alternative versions of the same gene proposed by language models were more likely to produce a strong signal
In a cell-based experiment, alternative versions of the same gene proposed by language models were more likely to produce a strong signal
On 16 August, researchers at the pharmaceutical company Roche posted a paper that has not yet undergone peer review on the bioRxiv preprint server. Language models proposed alternative ways to encode the same gene without changing its protein product. The authors then tested these sequences in cells to determine which ones produced a stronger signal.
A codon is a three-letter sequence in the genetic code that specifies one amino acid in a protein. Most amino acids can be specified by several codons. Replacing one such codon with another does not change the protein, but it can alter how the RNA folds, how stable it is in the cell, or how the ribosome reads it. The ribosome is the molecular machine that assembles proteins. Different nucleotide sequences encoding the same protein can therefore produce different amounts of that protein.
Conventional optimization programs select codons according to how frequently they occur in an organism's genes. The authors compared these programs with three types of language models, including one model tested in two modes. These models learn to predict a masked codon from nearby and distant regions of the sequence. For the experiment, the researchers created 68 versions of the SEAP gene, which encodes an enzyme whose activity can be measured easily in the medium surrounding the cells. In the study of Pichia-CLM, a model developed for yeast, codon sequences were selected to increase the production of six proteins. The new preprint applies this task to human cells and compares a library of sequences generated by several models with randomly recoded sequences.
After 24 hours in HEK293A human cells, 83.8% of the model-generated sequences produced higher SEAP activity than the original gene sequence. Among the 17 randomly recoded sequences, 5.9% exceeded the original sequence. The authors then integrated one copy of each sequence into the cellular genome so that differences in DNA quantity would not distort the comparison. In this second stage, half of the model-generated sequences outperformed the original sequence, and the model-generated group showed greater activity than the randomly recoded group.
Sequences in which approximately 5–12% of the codons had been replaced produced a stronger SEAP signal. The mode that rewrote almost half of the 507 SEAP codons produced the weakest signal after genomic integration. The authors suggest that preserving most of the original codon context may be beneficial for some genes.
The authors propose using a model to generate a set of synonymous gene sequences and then testing them in cells to identify those that produce a stronger signal.