Skip to main content

Model reference — AI Medical · Medical images

AI models for Medical image segmentation

Segmentation models label every voxel in a study, which is what turns an image into a measurement: a volume, a diameter, a distance to a margin.

Three approaches are worth knowing before you choose. Prompt-driven models segment whatever a clinician points at, with no training and no wait — the fastest route to a working pipeline. Trained models cover a fixed list of structures automatically and are more accurate on them, which is what routine throughput needs. Whole-body models sit in between: no prompt, more than a hundred structures, one pass per study.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Medical image segmentation service AI Medical services Pricing

Input type — Medical images

Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

MedSAM2

Vendor: Bo Wang Lab, University of Toronto

What it does: segments any structure you point at in CT, MRI and ultrasound, including through a 3D volume, with one click or box per structure. No training needed, which makes it the fastest route to a working pipeline.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 40≈ 140≈ 380

nnU-Net v2

Vendor: German Cancer Research Center

What it does: configures and trains itself to your own annotated cases and remains the accuracy reference for organ and tumour segmentation. The right choice where you hold annotations and accuracy governs.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 30≈ 110≈ 300

TotalSegmentator v2

Vendor: University Hospital Basel

What it does: labels 117 anatomical structures in a whole-body CT in a single pass, with no prompt and no training. Used for automatic measurement across an entire archive.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 25≈ 90≈ 240

SAM-Med3D

Vendor: Shanghai AI Laboratory

What it does: works natively in three dimensions rather than slice by slice, so a structure stays consistent through the volume instead of drifting between slices.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 20≈ 70≈ 200

STU-Net

Vendor: Shanghai AI Laboratory

What it does: pretrained on large CT collections and fitted to a new structure from a small number of annotated cases, which suits structures no released model covers.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 18≈ 60≈ 170

MONAI Auto3DSeg

Vendor: NVIDIA and the MONAI consortium

What it does: selects, trains and ensembles segmentation models for your dataset automatically. Heavier to run, and the option when the target has no suitable model at all.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 12≈ 45≈ 130

Prompt-driven segmentation

These models segment whatever you point at — a click, a box or a phrase — with no task-specific training, which is the fastest route to a working pipeline. Tiers and rates read as above; use them for initial sizing only.

MedSAM

Vendor: University of Toronto (Bo Wang lab)

What it does: outlines any structure you box on a 2D medical image, across ten modalities, with no training at all — the quickest way to get usable masks out of a mixed archive.

Median Dice 0.85–0.92 across 10 imaging modalities with box prompts (independent, Nature Communications 2024).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 300≈ 1,100≈ 2,700

LiteMedSAM

Vendor: University of Toronto (Bo Wang lab)

What it does: does the same job as MedSAM at a fraction of the compute, which is what makes it viable on every image rather than on a sample.

Dataset-dependent; optimised for much faster inference than MedSAM.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 900≈ 3,200≈ 8,100

SAM-Med2D

Vendor: OpenGVLab

What it does: segments from a click or a box, trained on 4.6 million medical images and nearly 20 million masks — the largest prompted 2D medical training set released.

Dataset-dependent; trained on 4.6M medical images and 19.7M masks, evaluated with Dice and IoU.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 300≈ 1,100≈ 2,700

SegVol

Vendor: SegVol authors / BAAI

What it does: segments a 3D volume from a phrase, a box or a point, and needs far fewer prompts than slice-by-slice approaches because it works in three dimensions natively.

Competitive 3D Dice with far fewer prompts than slice-wise SAM (developer-reported).

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

MedLSAM

Vendor: OpenMedLab

What it does: finds the target in a 3D volume and then segments it, so an operator does not have to locate the structure first.

Dataset-dependent; reported with mean IoU for localisation and Dice for segmentation.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

SAT

Vendor: SAT authors

What it does: segments from a text prompt across a very large vocabulary of structures, which suits work where the target list is long and changes.

Task-dependent across large-vocabulary medical segmentation datasets.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

BiomedParse-v2

Vendor: Microsoft Research

What it does: segments a 3D volume from a phrase across CT, MRI, ultrasound, PET and microscopy, and also reports whether the object is present at all — first place in the CVPR 2025 challenge for this task.

Task-dependent; first place in the CVPR 2025 text-guided 3D biomedical segmentation challenge according to the official repository.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

BiomedParse-v1

Vendor: Microsoft Research

What it does: segments, detects and labels more than a hundred biomedical object types from text across nine modalities, on 2D images.

Task-dependent across 100+ biomedical object tasks and nine modalities.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 300≈ 1,100≈ 2,700

Automatic and whole-body segmentation

These models segment a fixed list of structures with no prompt, which is what a routine measurement pipeline needs. Tiers and rates read as above; use them for initial sizing only.

VISTA3D / VISTA-3D / NV-Segment-CT

Vendor: NVIDIA / Project MONAI

What it does: segments 127 CT classes automatically and lets a clinician correct the result interactively, which is usually what gets a mask signed off rather than re-drawn.

Dice about 0.87 across 127 CT classes; interactive correction supported (developer-reported).

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

VISTA-2D

Vendor: NVIDIA / Project MONAI

What it does: segments 2D medical and microscopy images, as the 2D counterpart of the same toolchain.

Task-dependent; reported per dataset.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 300≈ 1,100≈ 2,700

NV-Segment-CTMR

Vendor: NVIDIA / Project MONAI

What it does: covers more than 300 anatomical classes across CT and MRI with no prompt, which makes whole-archive measurement a scheduled job rather than a project.

Task-dependent across 300+ supported classes; no single universal Dice score.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

TotalSegmentator MRI

Vendor: University Hospital Basel

What it does: brings whole-body segmentation to MRI across 80 structures on any sequence, and beat other public tools by a clear margin in independent evaluation.

Dice 0.839 on the internal MRI test set across 80 structures; beat other public tools 0.862 against 0.759 on 40 structures (independent, Radiology 2025).

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

MONAI Model Zoo bundles

Vendor: NVIDIA / Project MONAI

What it does: supplies ready-trained bundles for specific targets — spleen, pancreas, lung nodule, multi-organ, whole-body CT — so a common task does not need a training run.

Bundle-specific, typically Dice 0.85–0.96 on the source challenge dataset (developer-reported per bundle).

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

Universal Model

Vendor: Universal Model authors

What it does: segments abdominal organs and tumours, optionally directed by a language target, from a single CT volume.

Task-dependent across abdominal organ and tumour classes.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

SegmentAnyBone

Vendor: SegmentAnyBone authors

What it does: segments bone in MRI across anatomical sites, which general soft-tissue models handle poorly.

Task-dependent across anatomical sites; evaluated with Dice.

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 90≈ 315≈ 810

Choosing between them

The decision comes down to three things: whether your targets are standard anatomy, how many annotated cases you already hold, and whether a clinician is available to place a prompt. Send us a sample of your studies and we will come back with a model, the fine-tuning it needs to reach your accuracy target, and the hardware to run it.

At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.

Medical image segmentation service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for medical image segmentation. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.