Models in this group take clinical or biomedical text and return text. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
Baichuan-M2-32B
Vendor: Baichuan Intelligence
What it does: holds a multi-turn clinical conversation and currently posts the strongest physician-graded scores of any open model. At 32B it runs quantised on one or two 48 GB cards, which makes it the best quality-per-GPU option on this page.
HealthBench 60.1 overall, 34.7 Hard, 91.5 Consensus — highest among compared open models and above most closed models (developer-reported, arXiv 2509.02208; not independently reproduced).
gpt-oss-120b / gpt-oss-20b
Vendor: OpenAI (open-weight release)
What it does: answers clinical questions with a controllable reasoning budget — more thinking on hard cases, less on routine ones. The 120B version fits a single 80 GB GPU, so leading open-weight quality does not require a multi-GPU node.
Leading open-weight scores on the HealthBench public leaderboard (July 2026); 120b fits one 80 GB GPU (independent, llm-stats.com).
MedGemma 27B text-it
Vendor: Google DeepMind / Google Research
What it does: answers medical questions and drafts clinical text, and is the strongest exam-benchmark performer here by a wide margin. The obvious default where licence terms permit your use.
MedQA 87.7 percent zero-shot, 89.8 best-of-5; MedMCQA 74.2; PubMedQA 76.8; MMLU-Med 87.0; AfriMed-QA 84.0; exceeds physician performance on AgentClinic-MedQA (developer-reported).
HuatuoGPT-o1-8B / 70B / 72B
Vendor: FreedomIntelligence (CUHK-Shenzhen / SRIBD)
What it does: works through a differential step by step and shows its reasoning, which is what makes an answer reviewable rather than merely plausible. The 8B version is the best in its size class, so the reasoning does not have to cost 70B hardware.
8B: MedQA 72.6, MedMCQA 60.4, PubMedQA 79.2, average 63.9 across 7 benchmarks — best in its size class (independent, arXiv 2412.18925).
Meditron-7B / 70B, Meditron3-8B / 70B
Vendor: EPFL (LiGHT lab) / Yale
What it does: answers from a base trained on medical literature and guidelines, which shows in how it phrases clinical reasoning. Meditron3 follows instructions more closely, which matters when output must fit a fixed template.
Meditron-70B about 70 percent MedQA; Meditron3 adds a Llama-3.1 base and better instruction following (independent).
Med42-v2 8B / 70B
Vendor: M42 Health (Abu Dhabi)
What it does: is tuned for clinical instruction following, so it stays inside the format and scope you give it. Strong on examination-style questions at both sizes.
About 72 percent (8B) and about 80 percent (70B) MedQA; strong on USMLE-style sets (developer-reported).
OpenBioLLM-8B / 70B
Vendor: Saama AI Research
What it does: answers biomedical and clinical questions and holds up well on published QA suites. Note that independent reproduction came in below the original developer claims — we benchmark it against your own material.
70B: MedQA 76.1, MedMCQA 74.7, PubMedQA 79.2 (independent reproduction; developer claims were higher).
UltraMedical-8B / 70B
Vendor: Tsinghua University (UltraMedical project)
What it does: targets medical question answering specifically, and the 70B version is competitive with much larger general models on the standard sets.
70B: MedQA 82.2, MedMCQA 71.8; 8B: MedQA 71.1 (independent).
BioMistral-7B
Vendor: Avignon Université / Nantes Université
What it does: runs on a single mid-range card, which puts a medical-tuned model within reach of high-volume, narrow tasks where a 70B model would never pay for itself. Accuracy is correspondingly modest.
MedQA 45.0, MedMCQA 40.2, PubMedQA 66.9 (independent, HuatuoGPT-o1 comparison table).
Me-LLaMA
Vendor: Me-LLaMA authors
What it does: is a medical-domain language model for question answering, generation and embeddings from one deployment, useful where several text tasks share a pipeline.
Benchmark-dependent; evaluated across medical NLP and QA tasks.
MedAlpaca-7B / MedAlpaca-13B
Vendor: MedAlpaca authors
What it does: holds a conversational medical exchange, and is small and well understood enough to serve as a baseline before a larger model is justified.
Benchmark-dependent; conversational medical model, no single universal accuracy.
PMC-LLaMA-7B / PMC-LLaMA-13B
Vendor: PMC-LLaMA authors
What it does: answers from a base trained on biomedical literature, which suits question answering grounded in published evidence rather than in hospital notes.
Benchmark-dependent; evaluated on medical QA tasks.
Asclepius-Llama3-8B
Vendor: Stanford and community (Asclepius project)
What it does: summarises and answers questions over clinical notes, and is trained entirely on synthetic notes — so it is redistributable without the licence friction most clinical models carry.
Approaches GPT-3.5 on clinical note QA and summarisation; trained purely on synthetic notes so it is redistributable (developer-reported, ACL).
MedFound-7B / MedFound-176B
Vendor: Chinese Academy of Sciences and collaborators
What it does: reads a clinical narrative and returns a diagnosis list with its reasoning. Developer evaluation reports accuracy competitive with clinicians across specialties, on internal data.
Diagnostic accuracy competitive with clinicians on internal multi-specialty evaluation (developer-reported, Nature Medicine 2025).
Diabetica-7B / Diabetica-o1
Vendor: Shanghai Jiao Tong University
What it does: is built for one disease area and beats much larger general models inside it — the case for a specialist model when your scope is narrow.
Outperforms GPT-4 on diabetes-specific QA sets (developer-reported).
Qwen3 (8B / 32B / 235B-A22B)
Vendor: Alibaba Qwen
What it does: is a strong general model, not a medical one: paired with retrieval over your own guidelines it closes most of the gap, and its permissive licence removes a common blocker.
Strong general reasoning; medical performance below dedicated medical tunes unless paired with guideline retrieval (independent).
DeepSeek-R1 / R1-distill variants
Vendor: DeepSeek AI
What it does: produces long chain-of-thought reasoning, which is what complex differentials need. It is the heaviest option here, and the distilled variants exist for exactly that reason.
MedQA about 90 percent for the full model; heavy compute, strong chain-of-thought for complex differentials (independent).