Skip to main content

Model reference — AI Medical · Genomics

AI models for Single-cell analysis

Leading performance on cell-type annotation, batch integration and perturbation prediction is reported for the best-established model in this group, and strong few-shot results on network biology and disease gene prioritisation for the other.

All of these are representation models: they produce embeddings, and the analysis you care about is a head trained on top. That is why reported accuracy is task-dependent throughout and why local evaluation is quick — the heads train in minutes.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Single-cell analysis service AI Medical services Pricing

Input type — Single-cell expression data

Models in this group take single-cell gene-expression matrices or profiles. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

scGPT

Vendor: University of Toronto (Bo Wang lab)

What it does: leads on cell-type annotation, batch integration and perturbation prediction — the broadest single base for single-cell work, and heads train on top of it in minutes.

Leading performance on cell-type annotation, batch integration and perturbation prediction (independent, Nature Methods 2024).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (cells/hour)≈ 30,000≈ 105,000≈ 270,000

Geneformer

Vendor: Broad Institute / MIT (Ellinor lab), Theodoris Lab

What it does: is strong few-shot on network biology and disease gene prioritisation, and supports in-silico perturbation before an experiment is run.

Strong few-shot performance on network biology and disease gene prioritisation (independent, Nature 2023).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (cells/hour)≈ 30,000≈ 105,000≈ 270,000

scFoundation

Vendor: BioMap

What it does: is a large single-cell foundation model for annotation and prediction across downstream benchmarks.

Task-dependent across single-cell downstream benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (cells/hour)≈ 10,000≈ 35,000≈ 90,000

scBERT

Vendor: scBERT authors

What it does: encodes expression profiles for cell-type annotation at a small footprint, cheap enough for a whole atlas.

Task-dependent across single-cell annotation tasks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (cells/hour)≈ 90,000≈ 315,000≈ 810,000

Universal Cell Embeddings (UCE)

Vendor: UCE authors

What it does: produces cell embeddings that transfer across datasets, which is what makes combining data from several sources workable.

Task-dependent across cell-type annotation and cross-dataset transfer.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (cells/hour)≈ 30,000≈ 105,000≈ 270,000

Choosing between them

Pick on the task mix and the size of your atlas. Where you need perturbation prediction, the models differ meaningfully; where you need annotation, several are close and the choice is practical. Sizing below is per model rather than per dataset, since memory scales with genes retained.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Single-cell analysis service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for single-cell analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.