Skip to main content

World Reporter

UC San Diego Researchers Use AI to Decode Genetic “On Switch” Found in 60% of Human Genes

UC San Diego Researchers Use AI to Decode Genetic On Switch Found in 60% of Human Genes
Photo Credit: Unsplash.com

A research team at the University of California San Diego has used machine learning to decode the DNA signature of the initiator, a molecular switch that marks the precise point where gene reading begins in approximately 60% of focused human gene promoters. The study, published July 31, 2026, in the journal Genes & Development, gives scientists a quantitative tool to predict — for the first time — whether specific DNA mutations at the initiator site will disrupt gene activation, a capability with direct implications for cancer research, genetic diagnostics, and the broader effort to map the complete human gene expression code.

Key Takeaways

  • Researchers in Professor James T. Kadonaga’s laboratory at UC San Diego used high-throughput DNA sequencing to measure gene expression activity across approximately 500,000 different versions of the initiator region, then trained a machine learning model to identify the initiator’s characteristic DNA pattern.
  • The AI model identified the initiator sequence in roughly 60% of focused human gene promoters, providing the first reliable method for predicting its presence or absence across the genome.
  • The study also uncovered a previously unknown TATA-specific initiator variant, detected the overlapping but functionally distinct TCT motif, and mapped the relationships between multiple core promoter elements that govern how genes are activated.
  • The decoded initiator signature allows researchers to predict whether mutations at that site will disrupt gene expression — a capability relevant to understanding how certain mutations may drive cancer and other disorders.
  • The research was led by graduate student Torrey E. Rhyne-Carrigg and co-authored by Long Vo ngoc, Claudia Medrano, Kassidy E. Gillespie, and James T. Kadonaga, with computational support from the San Diego Supercomputer Center’s Expanse CPU.
  • Approximately 30% of focused human promoters still have no identified regulatory sequence explaining how they direct gene transcription, marking a significant frontier for future research.

The Initiator Controls Where Gene Reading Begins

Every cell in the human body contains roughly six billion DNA base pairs. Within that vast sequence, specialized segments of DNA known as promoters are responsible for orchestrating when, where, and to what extent each gene gets activated. The initiator is a specific element within those promoters that marks the exact starting point where the information encoded in a gene begins to be converted — or expressed — into the enzymes, hormones, proteins, and other functional products that sustain cellular life.

When the initiator functions correctly, it serves as a precise molecular signal that tells the cell’s transcription machinery where to begin reading a gene. When mutations disrupt the initiator, the gene may fail to activate properly, potentially causing the cell to malfunction. That breakdown in gene expression is one of the mechanisms that can lead to cancer and other disorders linked to abnormal cell behavior. Despite the initiator’s importance to gene activation, researchers had never fully decoded its DNA sequence pattern in human genes — until the UC San Diego team applied machine learning to the problem.

The challenge was one of scale. The initiator region is short, and the sequence differences between a functional initiator and a nonfunctional one can be subtle. Analyzing those differences across the entire human genome required the kind of pattern recognition that AI is particularly well suited to perform, especially when trained on a sufficiently large dataset of sequence variants with known expression outcomes.

500,000 DNA Variants Trained the AI Model

The research team, working in Kadonaga’s laboratory within UC San Diego’s Department of Molecular Biology, began by generating approximately 500,000 different DNA sequence variants of the initiator region. Using high-throughput DNA sequencing, the team measured the gene expression activity of each variant — effectively building a massive dataset that paired specific DNA sequences with their functional output.

That dataset became the training material for a machine learning system. The AI model learned to identify the characteristic pattern of DNA bases associated with a functioning initiator, distinguishing it from sequences that lacked initiator activity. Once the model had decoded the signature, the researchers deployed it across the human genome to search for the initiator’s telltale sequence in actual gene promoters. The result: approximately 60% of focused human promoters contain the initiator.

The computational demands of the work were substantial. The team used the Expanse CPU at the San Diego Supercomputer Center, supported through the National Science Foundation’s Advanced Cyberinfrastructure Coordination Ecosystem program. Additional funding came from the National Institutes of Health and the NSF Graduate Research Fellowship Program.

Kadonaga described the decoded initiator as a step forward in the combined use of laboratory experiments and AI to decipher the information embedded in human DNA sequences. The UC San Diego team’s approach — pairing high-throughput experimental data with machine learning — represents a methodology that could be applied to other unresolved elements of the gene expression code.

The Study Revealed New Relationships Between Promoter Elements

Beyond identifying the initiator itself, the machine learning analysis uncovered structural relationships between the initiator and other core promoter elements that had not been clearly defined. The study found a strict and synergistic interaction between the initiator and a downstream element known as the DPR (downstream promoter region), meaning the two elements work together in a tightly coordinated way to activate genes. The research also documented an inverse relationship between the TATA box — another well-studied promoter element — and the DPR, suggesting that promoters tend to use one regulatory strategy or the other, but rarely both at full strength.

The team identified a novel TATA-specific initiator variant, a previously unrecognized version of the initiator that appears to function specifically in promoters that also contain a TATA box. The study also detected the TCT motif, a sequence that overlaps with the initiator but serves a functionally distinct role. Notably, the TCT motif behaved differently in human genes than it does in fruit fly genes, a finding that highlights how core promoter elements have evolved distinct properties across species even when their DNA sequences appear similar.

These findings refine the scientific understanding of how genes are controlled at the most fundamental level. The promoter is not a single switch with one mechanism. It is a combinatorial system in which multiple elements interact — sometimes cooperatively, sometimes independently — to determine whether and how strongly a gene is expressed. The AI model’s ability to parse those interactions across hundreds of thousands of sequence variants is what made these structural discoveries possible.

Mutation Prediction Opens a Window Into Cancer Research

The practical significance of the decoded initiator lies in what it enables researchers to do next. With the initiator’s DNA signature identified, the AI model can now take a promoter sequence it has never analyzed before and predict how strongly it will drive gene expression. More critically, the model can predict how a specific mutation in the initiator region will change that expression — whether the mutation will reduce gene activation, eliminate it entirely, or leave it unchanged.

That predictive capability is relevant to cancer research because many cancer-associated mutations occur in non-coding regions of DNA — areas that do not directly encode proteins but regulate how genes are turned on and off. Mutations in the initiator region can silence a tumor-suppressing gene or alter the activation of a gene involved in cell growth. Before the UC San Diego study, researchers had limited ability to assess whether a given mutation at the initiator site would actually affect gene expression. The AI model now provides a quantitative framework for making those assessments.

The data and models produced by the study could also be used to design synthetic promoters with customized functions — engineered sequences that activate genes at specific levels, in specific cell types, or under specific conditions. That application has implications for gene therapy, where precise control over gene expression is a critical engineering challenge.

A Small Piece of a Larger Code Remains Unfinished

Kadonaga framed the initiator discovery as one component of a much larger scientific goal: building an AI model of the complete human gene expression code. That code, embedded across six billion DNA base pairs, determines which of the body’s roughly 20,000 genes are active in any given cell at any given time. A comprehensive AI model of the gene expression code would, in theory, allow researchers to predict the activity of every gene variant in every individual — a capability that would fundamentally change genetic medicine.

The current study makes progress toward that goal but also highlights how far the field has to go. Even after accounting for the initiator, the TATA box, and other known core promoter elements, approximately 30% of focused human promoters have no identified regulatory sequence that explains how they direct gene transcription. Those promoters clearly function — the genes they control are expressed in cells — but the DNA instructions governing their activation remain undecoded.

The UC San Diego research joins a broader wave of AI-driven genomics work unfolding in 2026. Google DeepMind’s AlphaGenome platform, also launched this year, can analyze up to one million DNA base pairs to predict regulatory and functional genomic elements. The convergence of experimental sequencing data, machine learning models, and high-performance computing infrastructure is accelerating the pace at which researchers can map the regulatory architecture of the human genome — a project that Kadonaga described with measured optimism, noting that expanding AI models of the human gene expression code is achievable in the not-too-distant future.

 

Disclaimer: This article reports on published peer-reviewed research and is intended for informational purposes only. It does not constitute medical advice, diagnosis, or treatment recommendations. Readers with questions about genetic testing, cancer risk, or personal health should consult a qualified healthcare professional.

 

FAQs

What is the initiator in human DNA?

The initiator is a specific DNA sequence element within a gene promoter that marks the exact point where gene transcription — the process of converting genetic information into functional products — begins. The UC San Diego study found that the initiator is present in approximately 60% of focused human gene promoters.

How did the researchers use AI in this study?

The team generated approximately 500,000 DNA sequence variants of the initiator region, measured the gene expression activity of each using high-throughput sequencing, and then trained a machine learning model on that data to identify the characteristic DNA pattern associated with a functioning initiator. The model can now predict whether specific mutations at the initiator site will affect gene expression.

What does this mean for cancer research?

Many cancer-associated mutations occur in non-coding regulatory regions of DNA, including gene promoters. The decoded initiator gives researchers a quantitative tool to predict whether mutations at the initiator site will disrupt gene activation, potentially helping identify mutations that silence tumor-suppressing genes or alter genes involved in cell growth. The research is at the foundational science stage and does not yet translate into clinical treatments or diagnostics.

Where was the study published?

The study, titled “Machine learning analysis of the human initiator region reveals key features of different types of core promoters,” was published on July 31, 2026, in Genes & Development, a peer-reviewed journal published by Cold Spring Harbor Laboratory Press. The research was funded by the National Institutes of Health and the National Science Foundation.

World Reporter

Bringing the World to Your Doorstep: World Reporter.