Skip to main content

Model reference — AI Medical · Medical images

AI models for Chest X-ray reporting support

Chest radiography is the highest-volume examination in most hospitals, which is why automated reading is further along here than anywhere else in imaging.

Three classes of model, three cost profiles. Classifiers return a probability per finding, run on modest hardware and are enough to order a queue. Vision-language models draft the findings text and cost considerably more per study. Image encoders produce no report at all — they turn a radiograph into a vector for retrieval, prior-study comparison, or a classifier trained on your own labels.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Chest X-ray reporting support service AI Medical services Pricing

Input type — Medical images

Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

CheXagent-8B / CheXagent-2-3B

Vendor: Stanford AIMI

What it does: interprets a chest radiograph and drafts structured findings text, trained specifically for this examination rather than adapted from a general model.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 400≈ 1,500≈ 4,200

MAIRA-2

Vendor: Microsoft Research

What it does: drafts a report in which each finding is grounded to the region of the image that supports it, so a reader can verify every statement.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (studies/hour)≈ 250≈ 900≈ 2,600

LLaVA-Rad

Vendor: Microsoft Research

What it does: a smaller radiology vision-language model that drafts findings on a single mid-range GPU, which makes per-study cost practical at department volume.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 600≈ 2,200≈ 6,000

TorchXRayVision

Vendor: Mila, Université de Montréal

What it does: returns probabilities for eighteen common findings. Light, well characterised across datasets and the dependable choice for triage ordering.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 9,000≈ 32,000≈ 90,000

RAD-DINO

Vendor: Microsoft Research

What it does: produces an image representation rather than text, for similarity search, retrieval of comparable prior studies, and training a classifier on your own labels.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 5,000≈ 18,000≈ 50,000

MedGemma 4B multimodal

Vendor: Google DeepMind

What it does: a general medical vision-language model that answers questions about a radiograph and drafts text, useful where the same service also handles other image types.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (studies/hour)≈ 800≈ 2,800≈ 7,500

Choosing between them

If the goal is a better-ordered worklist, a classifier does the job on a single mid-range card. If the goal is draft text, a vision-language model is required and the sizing changes accordingly. We benchmark the shortlist against studies your own readers have already reported, so the accuracy figures you plan around come from your population rather than a public leaderboard.

At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.

Chest X-ray reporting support service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for chest x-ray reporting support. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.