Skip to main content

Model reference — AI Medical · Whole-slide images

AI models for Digital pathology analysis

Pathology foundation models are the development that changed what is practical in this field: one pass over a slide produces a representation that many downstream tasks reuse.

They differ in the scale of tissue they were built on and in whether they handle text. Vision-only encoders are the workhorses for tiles. Vision-language models pair regions with descriptions, which lets an archive be searched in words. Slide-level models take the whole image rather than tiles, and cell-level models do the separate job of finding and classifying individual nuclei.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Digital pathology analysis service AI Medical services Pricing

Input type — Whole-slide images

Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

UNI 2 / UNI2-h

Vendor: Mahmood Lab, Harvard Medical School

What it does: a tile encoder built on a very large slide collection, and the strongest general starting point for classification and retrieval across tissue types.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 25≈ 90≈ 250

Virchow 2 / Virchow2G

Vendor: Paige

What it does: a tile encoder trained on a very large clinical archive, notable for holding accuracy across scanners and staining variation.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 18≈ 65≈ 180

Prov-GigaPath

Vendor: Microsoft Research and Providence

What it does: works at whole-slide scale rather than tile scale, which suits slide-level predictions where the relevant signal is spread across the tissue.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 12≈ 45≈ 130

CONCH

Vendor: Mahmood Lab, Harvard Medical School

What it does: pairs slide regions with text, so an archive can be searched by description instead of by a fixed label set.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 200

Phikon-v2

Vendor: Owkin

What it does: a lighter tile encoder that runs comfortably on one mid-range card, for a first pipeline or a cost-sensitive archive pass.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 30≈ 110≈ 300

CellViT++

Vendor: University Hospital Essen

What it does: detects and classifies individual nuclei rather than regions, and adds cell-level classification on top of segmentation, so a slide yields counts and densities rather than a heat map.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 15≈ 55≈ 150

Tile encoders

Independent benchmarking — Nature Biomedical Engineering 2025 across 19 models and 31 tasks, and Nature Communications 2026 across 32 models — found no single winner: CONCH and Virchow2 lead on average at AUROC 0.71, with tissue-specific strengths, and ensembling two or three encoders beats any one. Several of the strongest models are released under non-commercial terms, which we check against your intended use before building. Tiers and rates read as above; use them for initial sizing only.

H-optimus-0

Vendor: Bioptimus, with Aignostics, Mayo and Charité collaborators

What it does: is a 1.1B tile encoder that led pan-cancer performance in one independent benchmark, and is Apache-licensed for the weights — an unusual combination in this category.

Best pan-cancer performance in one independent benchmark; AUROC 0.68 mean across 31 tasks in another.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

H-optimus-1

Vendor: Bioptimus

What it does: is the newer version of the same encoder, for classification, grading, retrieval and prognosis work.

Task-dependent across classification, grading, retrieval, prognosis and slide-level benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

UNI

Vendor: Harvard Medical School (Mahmood lab)

What it does: is the widely cited Harvard tile encoder, strong on brain, bladder, breast and pan-cancer tasks. Gated and non-commercial, which is the constraint to check before it becomes your base layer.

Strong on brain, bladder, breast and pan-cancer tasks; AUROC 0.68 mean across 31 independent tasks. CC-BY-NC-ND, gated.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Virchow

Vendor: Paige AI, with Microsoft

What it does: is the first-generation Paige encoder, trained on a very large clinical archive and joint-best overall in independent benchmarking.

Joint-best overall in independent benchmarking, AUROC 0.71 across 31 tasks; leads on colon and prostate.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Hibou-B

Vendor: HistAI

What it does: is the Apache-licensed member of the modern encoder set — nearly the accuracy of the gated models with none of the licence friction, which often decides the choice.

AUROC 0.67 mean across 31 independent tasks; Apache-licensed, which is rare in this category.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Hibou-L

Vendor: HistAI

What it does: is the larger Hibou encoder, at similar benchmark accuracy and a non-commercial licence.

AUROC 0.67 mean across 31 independent tasks (CC-BY-NC for the L variant).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Phikon

Vendor: Owkin

What it does: is the small, fast Owkin encoder: a few points behind the leaders and cheap enough to run across a whole archive.

AUROC 0.65 mean across 31 independent tasks; small and fast.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

EXAONEPath

Vendor: LG AI Research

What it does: is among the top encoders for brain tissue tasks specifically, which is the case for it over a general leader.

Among top models for brain tissue tasks (independent). Non-commercial.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Path Foundation

Vendor: Google Research

What it does: produces patch embeddings that linear probes turn into tumour detection or grading with small label budgets. Superseded by MedSigLIP, and listed here for continuity.

Efficient linear probes for tumour detection and grading with small label budgets (developer-reported). Legacy model — MedSigLIP is the current recommendation.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

REMEDIS-Pathology

Vendor: Google Research

What it does: is a Google pathology representation model for classification, grading and retrieval from patch embeddings.

Task-dependent across classification, grading, retrieval and prognosis benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Kaiko

Vendor: Kaiko

What it does: is an openly published tile encoder for pathology classification and retrieval work.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Lunit-BT

Vendor: Lunit

What it does: is one of four Lunit encoders trained with different self-supervised objectives, which makes the set useful for ensembling rather than picking one.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Lunit-DINO

Vendor: Lunit

What it does: is the DINO-trained member of the Lunit set, for patch embeddings and downstream prediction.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Lunit-MoCoV2

Vendor: Lunit

What it does: is the MoCo v2 member of the Lunit set.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Lunit-SwAV

Vendor: Lunit

What it does: is the SwAV member of the Lunit set.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

CTransPath

Vendor: Sichuan University / Tencent AI Lab

What it does: is a long-established pathology encoder, still a solid retrieval baseline and light to run.

Task-dependent across pathology classification and retrieval benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

RetCCL

Vendor: Sichuan University / Tencent AI Lab

What it does: is built for retrieval specifically — finding visually comparable regions across an archive rather than labelling them.

Task-dependent across pathology retrieval benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

PathoDuet

Vendor: Shanghai Jiao Tong University

What it does: is a pathology encoder for patch-level classification and downstream prediction.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

BEPH

Vendor: Shanghai Jiao Tong University

What it does: produces patch embeddings for classification, grading and prognosis work.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

PathOrchestra

Vendor: Shanghai AI Lab

What it does: covers a wide set of pathology tasks from one model, which reduces the number of encoders a pipeline has to maintain.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

GPFM

Vendor: Smart Lab / HKUST collaborators

What it does: is a general pathology foundation model covering patch and slide-level prediction.

Task-dependent across pathology benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

HIPT

Vendor: Mahmood Lab

What it does: works up a hierarchy from patch to region to slide, which is how it reaches slide-level prediction without a separate aggregator.

Task-dependent across classification, grading and prognosis benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Slide-level and vision-language models

These models work at the level of a whole slide, or pair slide regions with text so an archive can be searched and labelled in words. Tiers and rates read as above; use them for initial sizing only.

CONCH1.5 / CONCHv1.5

Vendor: Harvard Medical School (Mahmood lab)

What it does: pairs slide regions with text, so an archive can be labelled and searched in words with no training. Joint-best overall in independent benchmarking and best on lung and prognostic tasks — gated and non-commercial.

Task-dependent; CONCH is joint-best overall at AUROC 0.71 across 31 tasks and best on lung and prognostic tasks (independent). CC-BY-NC-ND, gated.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

TITAN

Vendor: Harvard Medical School (Mahmood lab)

What it does: predicts at whole-slide level and beats multiple-instance-learning baselines on rare disease and few-shot classification, which is where slide-level labels are scarcest.

Outperforms MIL baselines for rare-disease and few-shot slide classification (developer-reported). CC-BY-NC-ND.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

MUSK

Vendor: Stanford University (Li Lab)

What it does: combines pathology images with clinical text for prognosis and immunotherapy response prediction — the furthest into outcome prediction any model here goes.

Leading multimodal pathology results including prognosis and immunotherapy response prediction (independent, Nature 2025). CC-BY-NC.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

PLIP

Vendor: Stanford University

What it does: is the smallest vision-language option: weakest of the modern set on accuracy, and easy enough to run that it suits a first pipeline or a cheap archive pass.

AUROC 0.64 mean across 31 independent tasks — weakest of the modern set but very small and easy to run. Non-commercial research use.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 45≈ 157.5≈ 405

MI-Zero

Vendor: MI-Zero authors

What it does: labels whole slides zero-shot from text prompts, with no slide-level training set required.

Task-dependent across zero-shot slide classification benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

QuiltNet

Vendor: QuiltNet authors

What it does: pairs pathology images with text for zero-shot labelling and retrieval.

Task-dependent across pathology classification and retrieval benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

PathAsst

Vendor: PathAsst authors

What it does: answers questions about a pathology image and generates descriptive text, for review and teaching rather than reporting.

Task-dependent across pathology VQA and classification benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

CHIEF

Vendor: Yu Lab / Harvard Medical School

What it does: produces slide-level embeddings, disease classes and prognosis scores from tiles or whole slides.

Task-dependent across classification, grading, retrieval and prognosis benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

PRISM

Vendor: Paige AI, with Microsoft

What it does: aggregates tile features into a slide embedding and generates report text from it.

Task-dependent across slide-level benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

COBRA

Vendor: Kather Lab

What it does: produces whole-slide embeddings for downstream slide-level prediction.

Task-dependent across slide-level benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

MADELEINE

Vendor: Mahmood Lab

What it does: learns across several stains of the same tissue, which is what makes it useful where your lab routinely runs more than H&E.

Task-dependent across slide-level benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

GigaSSL

Vendor: CBIO / Mines Paris

What it does: produces whole-slide embeddings cheaply, trained self-supervised at slide rather than tile level.

Task-dependent across slide-level benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (whole slides/hour)≈ 9≈ 31.5≈ 81

Cell and nucleus models

These models find and classify individual nuclei rather than regions, which is what cell counting and density measurement require. Tiers and rates read as above; use them for initial sizing only.

CellViT

Vendor: University Hospital Essen (IKIM)

What it does: detects and classifies individual nuclei at the strongest published accuracy for the task — PanNuke panoptic quality about 0.50 and detection F1 about 0.82 — which is what cell counting and density measurement actually need.

PanNuke panoptic quality about 0.50, mean detection F1 about 0.82 — state of the art for nuclei instance segmentation (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

HoVer-Net

Vendor: University of Warwick (TIA Centre)

What it does: is the long-standing reference implementation for nucleus instance segmentation and classification, and still a dependable baseline.

PanNuke and CoNSeP detection F1 about 0.80; long-standing reference implementation (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

StarDist / Cellpose 3

Vendor: MPI-CBG and community / MouseLand

What it does: segments cells across varied microscopy without tuning per dataset — the generalist option when your inputs are not all H&E.

Generalist cell segmentation, average precision above 0.80 on varied microscopy (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 45≈ 157.5≈ 405

Pathology pipeline tooling

Platforms and frameworks rather than models: they read whole-slide formats, manage annotation and train weakly-supervised classifiers on slide-level labels. Tiers and rates read as above; use them for initial sizing only.

QuPath + StarDist / Cellpose extensions

Vendor: University of Edinburgh / QuPath community

What it does: is the platform rather than the model: it reads whole-slide formats, manages annotation and hosts the segmentation models above.

Pipeline platform; accuracy depends on the embedded model.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (whole slides/hour)≈ 90≈ 315≈ 810

CLAM / TRIDENT

Vendor: Harvard Medical School (Mahmood lab)

What it does: trains a slide-level classifier from slide-level labels alone, with attention heatmaps showing which regions drove the prediction — no region annotation project required.

Weakly-supervised MIL framework; typical slide-level AUROC 0.85–0.98 depending on task and encoder.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (whole slides/hour)≈ 20≈ 70≈ 180

Choosing between them

Tile encoders differ less in accuracy than in licence terms and hardware appetite, so the choice is usually a practical one. Whether you need a text-capable or slide-level model follows from the task, not the budget. We benchmark on your own scanner output, because stain and scanner variation moves these numbers more than model choice does.

At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.

Digital pathology analysis service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for digital pathology analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.