Illumina SpliceAI2 adds deeper variant analysis with research limits

Illumina announced SpliceAI2 on October 8, expanding its tools for finding genetic changes that disrupt RNA splicing. The company reports stronger results in rare disease research and access through DRAGEN Annotation and Emedgene. For research teams considering the model, the important questions extend beyond the headline gain to how that result was measured, what the software requires and who may use it.
What changes in the model
Splicing assembles RNA transcripts before protein production. A genetic change can alter where that assembly happens, including at sites deep inside regions called introns. SpliceAI2 predicts splice sites, the use of junctions between them and complete transcript forms. Illumina says its training incorporated 330 samples with long RNA reads, helping the model learn which splice events belong to the same transcript.
The model also incorporates activity measurements from 147 splicing regulators to represent tissue context. That addition has limits. Illumina acknowledges that predicting how a particular variant changes its effect between tissues remains less successful. A more detailed splicing prediction therefore does not settle every question about where a variant matters in the body.
What the 17 percent result measures
The preprint makes the rare disease comparison more specific. Researchers selected 7,504 participants with particular inherited disorder characteristics from the Genomics England project. They tested whether predicted variants were enriched in genes linked to each participant’s symptoms, using shuffled symptom assignments to estimate the background expectation. At a fixed enrichment odds ratio of two, SpliceAI2 recovered 133 excess variants against 114 for the next best model. That difference underlies the roughly 17 percent improvement.
This was an analysis of candidate variant enrichment, rather than a trial showing 17 percent more patients receiving a diagnosis. The manuscript also compares the model with SpliceAI, Pangolin and AlphaGenome using separate splice variant benchmarks. Those comparisons excluded essential splice variants. The evidence supports a focused advance in the tested prediction tasks, while the authors identify tissue, cell type and developmental context as areas needing further work.
What a research team needs to run it
The public repository documents a CUDA capable GPU requirement. Variant scoring takes a reference genome, model weights and a table specifying chromosome, position, reference allele, alternate allele and strand. When strand is unknown, the instructions call for evaluating both strands separately and combining the predictions. Outputs include a summary score plus 260 columns of detailed predictions.
Existing SpliceAI users should also check their filtering rules. The repository recommends SpliceAI2 thresholds of 0.1 for high recall, 0.25 for balanced precision and recall, and 0.5 for high precision. Its corresponding legacy SpliceAI thresholds are 0.2, 0.5 and 0.8. Copying an old numerical cutoff unchanged would therefore miss the developer’s intended correspondence.
Reproducibility needs attention
Illumina’s model card says predictions are averaged across two models, so both model folders are required. It also warns that mixed precision computation and compilation can produce small score differences across GPU architectures and library versions. Changes in ordering can alter reported distances as well. For teams comparing runs, the practical implication is to record the software environment and hardware alongside the predictions rather than assume every installation will produce identical output.
Public access comes with restrictions
The repository’s license limits access to noncommercial use by qualifying noncommercial organizations. It explicitly excludes commercial affiliates, contract research organizations and research performed for the benefit of a commercial entity. The terms also restrict redistribution and use in commercial offerings. Teams should review the actual agreement before downloading or building a service around the model. Public code availability does not establish unrestricted commercial permission.
The same terms label the model for research use and prohibit diagnostic procedures. They warn that results may be inaccurate or unsuitable for particular populations, variants or applications. For a research group, SpliceAI2 offers a more detailed way to prioritize variants for investigation, with access and validation requirements that remain part of the adoption decision.






