The Codebook project identified the short DNA sequences preferred by 177 poorly studied proteins that regulate gene activity
The Codebook project identified the short DNA sequences preferred by 177 poorly studied proteins that regulate gene activity
On August 5, the Codebook project published a paper in Nature. The authors studied 332 putative human transcription factors, which are proteins involved in regulating gene activity, and identified motifs for 177 of them. A motif is a short pattern of preferences among DNA bases that indicates which sequences a protein is more likely to bind. For 130 factors, data from cells showed where in the genome these preferences occur.
In a 2018 catalog, researchers listed 1 639 putative human transcription factors. The motif was unknown for more than a quarter of them. Once a motif is known, researchers can test whether changing a single DNA base alters the binding of a specific protein and then investigate how that change affects gene activity.
To distinguish a protein's binding preference in a laboratory assay from its behavior in a cell, the Codebook team compared several types of experiments. Across 4 804 experiments, the team studied 393 proteins: 332 candidates and 61 previously characterized factors used as controls. In some experiments, the protein selected the DNA sequences to which it bound. In others, it was presented with fragments of the human genome. ChIP-seq showed which DNA fragments the protein bound in cultured human HEK293 cells. The researchers considered a motif reliable when at least two methods produced similar results and other experiments confirmed its predictions.
A short sequence can occur by chance at many locations in the genome. The authors therefore looked for regions supported by three lines of evidence: the region contained the motif, the protein bound it in experiments using genomic fragments, and the protein was detected at the same location in cells. At least one such region was found for 85 of the 101 factors with both types of data. Within these regions, the motif was better conserved across mammals than the neighboring DNA. In total, the authors identified 113 577 such conserved regions: 82 760 for Codebook factors and 30 817 for control factors.
This map makes it possible to frame a testable question for a DNA variant: which protein can distinguish between the two versions of the sequence, and where in the genome should researchers look for effects on gene activity? Among 2 260 variants that substantially changed the similarity to a motif, the prediction agreed with measurements of which of the two variants the protein bound more often in 1 682 cases. For each such variant, the catalog provides a specific hypothesis about which protein and which genomic region should be tested.