Illumina reports 17 percent more disease-associated variants with SpliceAI2, tested on Genomics England data

The successor to the 2019 SpliceAI model — cited in more than 3,400 publications and embedded in ClinGen recommendations — uses DNA sequence alone to predict full transcripts.

Macro illustration of a DNA-like thread with one small segment glowing faintly, a metaphor for detecting more disease-associated variants within genomic sequence.
Illustration
Gift article

Illumina reports 17 percent more disease-associated variants with SpliceAI2, tested on Genomics England data

The successor to the 2019 SpliceAI model — cited in more than 3,400 publications and embedded in ClinGen recommendations — uses DNA sequence alone to predict full transcripts. But all the numbers come from Illumina itself, and clinical status remains an open question.

On October 8, 2026, Illumina launched SpliceAI2, a genomic AI model from the company's BioInsight AI Lab that predicts how genetic variants alter RNA splicing (Unite.AI). In an ordinary product context, this would be a minor news item. But the original 2019 model has, according to Illumina, been cited in more than 3,400 publications and is incorporated into the recommendations for interpreting splice variants from ClinGen, a clinical research body that sets standards for clinical genomics. A tool in that position affects not just one company's product portfolio — it touches an established practice in rare disease genomics worldwide.

The question, then, is not only whether SpliceAI2 is technically better, but what updating such a widely used reference tool means for researchers and clinicians who have built their workflows around the previous version.

How the model works

The technical leap lies in what the model delivers from the same input. SpliceAI2 requires only a DNA sequence, according to Illumina, which the company says makes transcript-level analysis possible without RNA data from tissue that is difficult to biopsy.

According to a manuscript describing the work, the model processes 196,608 base pairs of genomic sequence with roughly 13 million trainable parameters. Splice sites and splice junctions are predicted and integrated into a so-called splice graph, from which complete transcripts and their usage frequencies are inferred. That is a substantial extension over merely predicting local splicing events: the model is meant to say something about how the entire transcript is assembled, and to what degree different isoforms are actually used.

The training: a dataset more than a hundred times larger

Illumina says SpliceAI2 was trained on a dataset more than a hundred times larger than its predecessor's. According to the manuscript, the model was trained end-to-end on 314,745 RNA sequencing samples from human and nine other mammalian species, together covering more than 46 million observed splice junctions after filtering. In addition came 330 long-read RNA sequencing samples from the public ENCODE project — data that provide full-length transcripts and thus train the ability to reconstruct entire isoforms, not just local splicing events.

The numbers — and who stands behind them

According to the manuscript, as relayed in secondary reporting, SpliceAI2 outperformed the original SpliceAI, Pangolin and AlphaGenome on several benchmarks:

  • GTEx, cryptic splice variants: auPRC 0.77 versus 0.66 for the next-best model.
  • Quantifying splice site usage: Spearman correlation 0.63 versus 0.47 — the largest relative improvement among the results.
  • sQTL classification (splicing quantitative trait loci): auROC 0.76 versus 0.73.
  • OpenSplice, a massively parallel reporter assay: Spearman 0.59 versus 0.55.

According to the reporting, the comparisons with AlphaGenome were run by collaborators at the University of Oxford.

The most consequential figure is nonetheless the one closest to clinical research: In an analysis of 7,504 probands from the Genomics England 100,000 Genomes Project, variants prioritized by SpliceAI2 are reported to be significantly enriched in phenotype-matched disease genes, with 17 percent more disease-associated variants than any other tested splicing model — at matched confidence thresholds. That the thresholds are matched matters: it means the gain cannot simply be explained by the model flagging more variants, but that it reportedly ranked them more accurately at the same level of stringency.

One caveat applies to all of these figures, however: they come from Illumina and the manuscript, relayed through secondary reporting, and have not been independently verified. The conclusion of "17 percent more variants" therefore applies to Illumina's own benchmarks and one specific dataset — not a documented gain in general clinical practice.

Where it may matter

The practical benefit lies primarily in rare disease. Tools like the original SpliceAI already have a fixed place in the prioritization of variants that affect splicing. A model that both captures more cryptic variants and reconstructs full transcripts could, in theory, move more variants from "unknown significance" to "actionable candidate" — without the laboratory needing tissue-specific RNA material.

SpliceAI2 also fits into a larger strategy. The model joins PromoterAI and PrimateAI-3D in Illumina's series of genomic AI models, which together cover splice, promoter and missense variants. Illumina says the three models collectively allow researchers to identify up to twice as many variants with predicted biological effect — again a company claim, not an independent measurement.

The open questions

Several key questions are unanswered in what has been reported so far:

  • Research or clinic? It is not stated whether SpliceAI2 is intended for clinical use or research only, and no regulatory status is mentioned.
  • Will ClinGen follow? The original is embedded in ClinGen recommendations for splice variant interpretation. It is unclear whether or when the recommendations will reference the successor — and until that happens, many laboratories will continue with the 2019 model.
  • Publication status. The available reporting does not indicate whether the manuscript is peer-reviewed or a preprint.
  • Independent verification. Since the benchmarks come from the developer itself, independent replications — ideally on cohorts other than Genomics England — will be the real test.

SpliceAI2 is thus, for now, a promising but self-evaluated update to one of genomics' most cited AI tools. How much it changes practice around splice variants depends less on the benchmark numbers than on the answers to the questions that remain.

AIMag.no
AIMag.no
The AIMag.no editorial team covers artificial intelligence, tools, research, and regulation.

Get the best of AI MAG in your inbox

News, analysis, and ideas at the intersection of AI and society.