Chest radiography is the highest-volume examination in most hospitals, which is why automated reading is further along here than anywhere else in imaging.
Three classes of model, three cost profiles. Classifiers return a probability per finding, run on modest hardware and are enough to order a queue. Vision-language models draft the findings text and cost considerably more per study. Image encoders produce no report at all — they turn a radiograph into a vector for retrieval, prior-study comparison, or a classifier trained on your own labels.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
CheXagent-8B / CheXagent-2-3B
Vendor: Stanford AIMI
What it does: interprets a chest radiograph and drafts structured findings text, trained specifically for this examination rather than adapted from a general model.
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (studies/hour)
≈ 400
≈ 1,500
≈ 4,200
MAIRA-2
Vendor: Microsoft Research
What it does: drafts a report in which each finding is grounded to the region of the image that supports it, so a reader can verify every statement.
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (studies/hour)
≈ 250
≈ 900
≈ 2,600
LLaVA-Rad
Vendor: Microsoft Research
What it does: a smaller radiology vision-language model that drafts findings on a single mid-range GPU, which makes per-study cost practical at department volume.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (studies/hour)
≈ 600
≈ 2,200
≈ 6,000
TorchXRayVision
Vendor: Mila, Université de Montréal
What it does: returns probabilities for eighteen common findings. Light, well characterised across datasets and the dependable choice for triage ordering.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (studies/hour)
≈ 9,000
≈ 32,000
≈ 90,000
RAD-DINO
Vendor: Microsoft Research
What it does: produces an image representation rather than text, for similarity search, retrieval of comparable prior studies, and training a classifier on your own labels.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (studies/hour)
≈ 5,000
≈ 18,000
≈ 50,000
MedGemma 4B multimodal
Vendor: Google DeepMind
What it does: a general medical vision-language model that answers questions about a radiograph and drafts text, useful where the same service also handles other image types.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (studies/hour)
≈ 800
≈ 2,800
≈ 7,500
Choosing between them
If the goal is a better-ordered worklist, a classifier does the job on a single mid-range card. If the goal is draft text, a vision-language model is required and the sizing changes accordingly. We benchmark the shortlist against studies your own readers have already reported, so the accuracy figures you plan around come from your population rather than a public leaderboard.
At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for chest x-ray reporting support. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.