One model in this group reports zero-shot chest X-ray AUC 0.90 for cardiomegaly, 0.93 for lung opacity and 0.91 for pleural effusion, and beats the earlier single-modality foundation models on most tasks. It is also the producer’s own recommended replacement for several legacy models we still list for continuity.
Independent benchmarking is more sober: mean pathology AUROC 0.66 across a 31-task benchmark for one widely used biomedical image-text model. Both numbers are real; they measure different things.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take a medical image, optionally with text, and return embeddings rather than a report. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
MedSigLIP-448 (400M)
Vendor: Google Health
What it does: puts chest radiographs, CT and MRI slices, dermatology, ophthalmology and pathology images into one searchable image-text space — so a single index serves every modality, and a new finding can be scored without training a classifier for it.
Zero-shot CXR AUC 0.90 cardiomegaly, 0.93 lung opacity, 0.91 pleural effusion; EyePACS 5-class DR grading competitive with task-specific models; beats legacy CXR, Derm and Path Foundation models on most tasks (developer-reported).
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 4,000
≈ 14,000
≈ 36,000
BiomedCLIP (PMC-15M)
Vendor: Microsoft Research
What it does: indexes biomedical figures and images against their captions for retrieval and zero-shot labelling. Widely used and well understood; independent benchmarking puts its pathology performance modestly, which is worth knowing before it becomes your only index.
State of the art biomedical image-text retrieval and zero-shot classification at release; mean pathology AUROC 0.66 in an independent 31-task benchmark.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 12,000
≈ 42,000
≈ 108,000
MedImageInsight (0.36B)
Vendor: Microsoft Research
What it does: covers X-ray, CT, MRI, dermatology, OCT, ultrasound and pathology in a 0.36B model — a small footprint for that breadth, which makes indexing a large archive cheap.
State of the art or near it on 14 image-classification and retrieval tasks spanning X-ray, CT, MRI, dermatology, OCT, ultrasound and pathology (developer-reported).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 12,000
≈ 42,000
≈ 108,000
Choosing between them
Coverage decides it. If your archive spans several modalities, a multi-modality model gives you one index instead of four. If you work in one modality with a large volume, a specialist encoder may score better. Indexing is a one-off cost we size separately from query serving.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for medical image retrieval. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.