Skip to main content

Model reference — AI Medical · Radiology

AI models for Chest X-ray screening

Reported accuracy clusters around AUC 0.88–0.93 for the common findings, and one model in this group is explicitly cross-dataset validated at AUC 0.78–0.85 — a lower number that is more honest about what transfers.

Three families: probability classifiers, image encoders trained without text supervision, and image-text models that support zero-shot classification and retrieval. The knowledge-enhanced models add an external medical knowledge source to the image-text pairing.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Chest X-ray screening service AI Medical services Pricing

Input type — Chest radiographs

Models in this group take a chest radiograph, some with text prompts alongside. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

TorchXRayVision model zoo

Vendor: Mila / Machine Learning and Medicine Lab (Cohen et al.)

What it does: scores eighteen common findings on a chest radiograph, fast and on modest hardware. Its published accuracy is cross-dataset validated, which makes it lower than the competition and more likely to hold on your data.

AUC 0.78–0.85 across 18 findings, explicitly cross-dataset validated so the numbers are modest but honest (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (images/hour)≈ 30,000≈ 105,000≈ 270,000

RAD-DINO

Vendor: Microsoft Research

What it does: turns a chest radiograph into an image representation without using report text to learn, and is best in class for the classifiers and report models built on top of it.

Best-in-class CXR image encoder for downstream classification and report generation, trained without text supervision (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

CXR Foundation / ELIXR-B

Vendor: Google Health

What it does: scores findings zero-shot or trains a classifier from a small labelled set, at AUC around 0.89–0.93 on the common findings.

Zero-shot AUC 0.89 cardiomegaly, 0.93 pleural effusion, 0.88 pneumonia on CheXpert-style labels (developer-reported). Legacy model — MedSigLIP is the current recommendation.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

CheXzero

Vendor: Stanford / CheXzero authors

What it does: scores pathology from a text prompt alone, so a new finding needs a phrase rather than a labelled dataset.

Task-dependent across chest X-ray pathology labels; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

MedCLIP

Vendor: MedCLIP authors

What it does: indexes chest radiographs against text for zero-shot classification and retrieval, useful where your finding list changes faster than you can retrain.

Task-dependent across zero-shot X-ray classification and retrieval.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

CXR-CLIP

Vendor: CXR-CLIP authors

What it does: pairs radiographs with text for classification and retrieval, and is light enough to index a large archive quickly.

Task-dependent across chest X-ray classification and retrieval benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

MedKLIP

Vendor: MedKLIP authors

What it does: adds an external medical knowledge source to the image-text pairing, which improves diagnosis on the findings where plain text pairing is weakest.

Task-dependent across chest X-ray diagnostic datasets.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

KAD

Vendor: KAD authors

What it does: brings structured medical knowledge into chest X-ray diagnosis, returning disease probabilities with the knowledge trail behind them.

Task-dependent across chest X-ray diagnosis benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

AFLoc

Vendor: AFLoc authors

What it does: localises the pathology on the image as well as scoring it, so a reader can check the claim against the pixels in seconds.

Task-dependent across localisation and classification benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

Choosing between them

For triage, a light classifier is enough and runs at tens of thousands of images per hour. For a changing finding list, a zero-shot image-text model avoids a retraining cycle each time. For your own endpoints, an encoder plus a linear probe is the most label-efficient route. Several models here carry research-only licences.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Chest X-ray screening service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for chest x-ray screening. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.