Skip to main content

Model reference — AI Medical · Clinical language

AI models for Medical visual question answering

Accuracy here is reported as benchmark averages across VQA suites, and the spread between models is wide. One 32B model reports about 66.6 average on its cited suite against about 61.8 for its 7B sibling — a useful illustration that size still buys accuracy in this group.

The practical split is by size and by modality coverage. 4B models run on one card and handle common modalities. 7–14B models broaden coverage. 27–32B models lead on reasoning and cost several times as much per answer.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Medical visual question answering service AI Medical services Pricing

Input type — Medical images plus text

Models in this group take one or more medical images with a text prompt and return text. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

Lingshu-7B / Lingshu-32B

Vendor: Alibaba DAMO Academy and collaborators

What it does: answers questions and drafts findings across twelve imaging modalities, and led open multimodal medical results at release. The 32B version buys about five points of benchmark average over the 7B at several times the cost per answer.

32B reported average about 66.6 and 7B about 61.8 across the cited multi-benchmark medical VQA suite; top open multimodal medical results at release (independent).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

MedGemma-4B-IT / MedGemma-4B-PT / MedGemma 1.5 4B / MedGemma-27B-Multimodal

Vendor: Google DeepMind / Google Research

What it does: covers chest radiographs, CT, MRI, pathology, dermatology and fundus images from one family, at sizes from a single-card 4B to a 27B node. Notably, 81 percent of its generated chest X-ray reports were judged clinically equivalent to the radiologist original.

4B-IT: MedQA 64.4, MedMCQA 55.7, PubMedQA 73.4, MMLU-Med 70.0; 81 percent of generated CXR reports judged clinically equivalent by a board-certified radiologist. 27B multimodal: MedQA about 85.3 zero-shot. 4B-PT: MIMIC-CXR RadGraph F1 29.5 before tuning. 1.5 4B: MedQA 69.1, EHRQA 89.6 (developer-reported).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Hulu-Med-4B / Hulu-Med-7B / Hulu-Med-14B / Hulu-Med-32B

Vendor: Zhejiang University / ZJU-AI4H

What it does: handles both 2D images and 3D volumes with text, across four sizes — useful when the same pipeline has to cover radiographs and CT without two deployments.

Benchmark-dependent; strong reported results across OmniMedVQA, VQA-RAD, SLAKE, PathVQA and related suites.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

HuatuoGPT-Vision-7B / 34B

Vendor: FreedomIntelligence

What it does: answers questions about medical images and leads its size class on several VQA benchmarks, under a permissive licence.

Leading open medical VLM on multiple VQA benchmarks in its size class (developer-reported).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

LLaVA-Med-v1.5-7B

Vendor: Microsoft Research

What it does: is the open reference baseline for biomedical visual question answering: small, well documented and the sensible first benchmark before paying for a larger model.

Beats prior supervised state of the art on 3 biomedical VQA sets after a single day of training; one reported modern comparison gives average medical-VQA score about 37.8 (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

RadFM (14B)

Vendor: Shanghai Jiao Tong University

What it does: takes 2D radiographs and 3D CT or MRI interleaved with text, which is rarer than it sounds — most medical vision-language models handle slices only.

Handles 2D and 3D radiology inputs; competitive on RadBench report generation and VQA (developer-reported).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

RadSight-4B / RadSight-8B

Vendor: Alibaba DAMO Academy

What it does: covers X-ray, CT and MRI for detection, question answering and report drafting, at two sizes that both fit a single node.

Task-dependent across disease detection, VQA and report generation; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

BiomedGPT

Vendor: Stanford and academic consortium

What it does: works across nine biomedical modalities and led 16 of 25 tasks in independent evaluation — the broadest coverage per deployment in this group.

State of the art on 16 of 25 tasks spanning 9 biomedical modalities (independent, Nature Medicine 2024).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

BioMedGPT

Vendor: BioMedGPT authors

What it does: handles biomedical images and text together for answers, labels and embeddings from one model.

Task-dependent across biomedical vision-language and text tasks; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Med-Flamingo

Vendor: Med-Flamingo authors

What it does: learns a new task from a handful of examples in the prompt, which suits questions you cannot assemble a training set for.

Benchmark-dependent across medical VQA and few-shot tasks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

MedVInT

Vendor: MedVInT authors

What it does: answers questions about medical images following visual instructions, at a size that runs on one card.

Benchmark-dependent across PMC-VQA and related medical VQA tasks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

PubMedCLIP

Vendor: PubMedCLIP authors

What it does: pairs medical images with text for question answering features and retrieval, and is light enough to run in front of a heavier model as a filter.

Task-dependent across medical VQA benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 2,500≈ 8,800≈ 22,500

Choosing between them

Match modality coverage to your actual case mix first: a model strong on chest radiographs may be weak on pathology. Then pick the smallest size that holds your accuracy target, because per-answer cost scales steeply. Several models here are research-only, which we check before building.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Medical visual question answering service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for medical visual question answering. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.