Models in this group take nucleotide sequence, a VCF, or sequence plus phenotype terms. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
Evo 2 (1B / 7B / 40B, 1M-token context)
Vendor: Arc Institute with NVIDIA and Stanford
What it does: reads up to a megabase of sequence in one context and scores variants with no task-specific training — about 0.95 AUROC separating pathogenic from benign BRCA1 variants zero-shot, and strong where classical tools are weakest, outside coding regions.
About 0.95 AUROC classifying pathogenic against benign BRCA1 variants zero-shot; strong non-coding variant performance (developer-reported, Nature 2026).
Evo
Vendor: Arc Institute
What it does: is the first-generation genome model, for sequence likelihoods, embeddings and generation at a lower hardware cost than Evo 2.
Task-dependent across genomic modelling and generation tasks.
Vendor: InstaDeep with NVIDIA
What it does: matches or beats specialised models across 18 genomics tasks from one encoder, with sizes from 50M to 2.5B so the footprint can follow the budget.
Matches or beats specialised models on 18 genomics prediction tasks (independent).
DNABERT-2
Vendor: DNABERT-2 authors
What it does: encodes DNA efficiently for classification and regression, a practical base when your labelled set is small.
Task-dependent across genomic sequence classification and regression benchmarks.
HyenaDNA
Vendor: Hazy Research / Stanford
What it does: handles very long sequences at low cost, which is what long-range regulatory work needs.
Task-dependent across long-range genomics benchmarks.
Vendor: Google DeepMind
What it does: predicts expression and epigenomic tracks from about 200 kb of sequence, at r around 0.85 across tracks — the reference for sequence-to-expression work.
Substantial improvement over prior models for gene expression prediction from sequence, r about 0.85 across tracks (independent).
GET
Vendor: GET authors
What it does: predicts transcription and regulatory activity using cellular context alongside the sequence, rather than sequence alone.
Task-dependent across transcription and regulatory-genomics benchmarks.
SpliceAI
Vendor: Illumina
What it does: predicts splice gain and loss at 95 percent top-k accuracy and is already standard in clinical variant pipelines — the least speculative model on this page.
95 percent top-k accuracy for splice site prediction; standard tool in clinical variant pipelines (independent, Cell 2019).
AlphaMissense
Vendor: Google DeepMind
What it does: scores missense variants for pathogenicity from the protein sequence, for triage inside a variant pipeline.
Benchmark-dependent; high pathogenic against benign discrimination reported, no single universal accuracy.
Exomiser
Vendor: Monarch Initiative / Jackson Laboratory
What it does: ranks candidate causal variants against the patient’s own phenotype terms, putting the causative variant first in 70–97 percent of solved rare-disease cases. A rule-and-ML hybrid, fully on-premise, and the most defensible option in a clinical pipeline.
Causative variant in the top rank in about 70–97 percent of solved rare-disease cases depending on cohort (independent).