Models in this group take one or more medical images with a text prompt and return text. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
Lingshu-7B / Lingshu-32B
Vendor: Alibaba DAMO Academy and collaborators
What it does: answers questions and drafts findings across twelve imaging modalities, and led open multimodal medical results at release. The 32B version buys about five points of benchmark average over the 7B at several times the cost per answer.
32B reported average about 66.6 and 7B about 61.8 across the cited multi-benchmark medical VQA suite; top open multimodal medical results at release (independent).
MedGemma-4B-IT / MedGemma-4B-PT / MedGemma 1.5 4B / MedGemma-27B-Multimodal
Vendor: Google DeepMind / Google Research
What it does: covers chest radiographs, CT, MRI, pathology, dermatology and fundus images from one family, at sizes from a single-card 4B to a 27B node. Notably, 81 percent of its generated chest X-ray reports were judged clinically equivalent to the radiologist original.
4B-IT: MedQA 64.4, MedMCQA 55.7, PubMedQA 73.4, MMLU-Med 70.0; 81 percent of generated CXR reports judged clinically equivalent by a board-certified radiologist. 27B multimodal: MedQA about 85.3 zero-shot. 4B-PT: MIMIC-CXR RadGraph F1 29.5 before tuning. 1.5 4B: MedQA 69.1, EHRQA 89.6 (developer-reported).
Hulu-Med-4B / Hulu-Med-7B / Hulu-Med-14B / Hulu-Med-32B
Vendor: Zhejiang University / ZJU-AI4H
What it does: handles both 2D images and 3D volumes with text, across four sizes — useful when the same pipeline has to cover radiographs and CT without two deployments.
Benchmark-dependent; strong reported results across OmniMedVQA, VQA-RAD, SLAKE, PathVQA and related suites.
HuatuoGPT-Vision-7B / 34B
Vendor: FreedomIntelligence
What it does: answers questions about medical images and leads its size class on several VQA benchmarks, under a permissive licence.
Leading open medical VLM on multiple VQA benchmarks in its size class (developer-reported).
LLaVA-Med-v1.5-7B
Vendor: Microsoft Research
What it does: is the open reference baseline for biomedical visual question answering: small, well documented and the sensible first benchmark before paying for a larger model.
Beats prior supervised state of the art on 3 biomedical VQA sets after a single day of training; one reported modern comparison gives average medical-VQA score about 37.8 (independent).
RadFM (14B)
Vendor: Shanghai Jiao Tong University
What it does: takes 2D radiographs and 3D CT or MRI interleaved with text, which is rarer than it sounds — most medical vision-language models handle slices only.
Handles 2D and 3D radiology inputs; competitive on RadBench report generation and VQA (developer-reported).
RadSight-4B / RadSight-8B
Vendor: Alibaba DAMO Academy
What it does: covers X-ray, CT and MRI for detection, question answering and report drafting, at two sizes that both fit a single node.
Task-dependent across disease detection, VQA and report generation; no single universal accuracy.
BiomedGPT
Vendor: Stanford and academic consortium
What it does: works across nine biomedical modalities and led 16 of 25 tasks in independent evaluation — the broadest coverage per deployment in this group.
State of the art on 16 of 25 tasks spanning 9 biomedical modalities (independent, Nature Medicine 2024).
BioMedGPT
Vendor: BioMedGPT authors
What it does: handles biomedical images and text together for answers, labels and embeddings from one model.
Task-dependent across biomedical vision-language and text tasks; no single universal accuracy.
Med-Flamingo
Vendor: Med-Flamingo authors
What it does: learns a new task from a handful of examples in the prompt, which suits questions you cannot assemble a training set for.
Benchmark-dependent across medical VQA and few-shot tasks.
MedVInT
Vendor: MedVInT authors
What it does: answers questions about medical images following visual instructions, at a size that runs on one card.
Benchmark-dependent across PMC-VQA and related medical VQA tasks.
PubMedCLIP
Vendor: PubMedCLIP authors
What it does: pairs medical images with text for question answering features and retrieval, and is light enough to run in front of a heavier model as a filter.
Task-dependent across medical VQA benchmarks.