Skip to main content

Model reference — AI Medical · Clinical language

AI models for Clinical reasoning assistant

None of these is a diagnostic device. Exam-style benchmarks such as MedQA are multiple-choice proxies that correlate poorly with consultation quality; rubric-graded benchmarks written by physicians are a better signal, and both are gameable. We report what producers and independent evaluations state and treat all of it as shortlisting evidence.

Three tiers matter commercially. Small models (7–8B) run on one card and suit high volume on narrow tasks. Mid-size models (27–32B) are the current sweet spot for quality against VRAM. Large models (70B and above) lead on reasoning and cost several GPUs each.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Clinical reasoning assistant service AI Medical services Pricing

Input type — Clinical text — general

Models in this group take clinical or biomedical text and return text. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

Baichuan-M2-32B

Vendor: Baichuan Intelligence

What it does: holds a multi-turn clinical conversation and currently posts the strongest physician-graded scores of any open model. At 32B it runs quantised on one or two 48 GB cards, which makes it the best quality-per-GPU option on this page.

HealthBench 60.1 overall, 34.7 Hard, 91.5 Consensus — highest among compared open models and above most closed models (developer-reported, arXiv 2509.02208; not independently reproduced).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

gpt-oss-120b / gpt-oss-20b

Vendor: OpenAI (open-weight release)

What it does: answers clinical questions with a controllable reasoning budget — more thinking on hard cases, less on routine ones. The 120B version fits a single 80 GB GPU, so leading open-weight quality does not require a multi-GPU node.

Leading open-weight scores on the HealthBench public leaderboard (July 2026); 120b fits one 80 GB GPU (independent, llm-stats.com).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

MedGemma 27B text-it

Vendor: Google DeepMind / Google Research

What it does: answers medical questions and drafts clinical text, and is the strongest exam-benchmark performer here by a wide margin. The obvious default where licence terms permit your use.

MedQA 87.7 percent zero-shot, 89.8 best-of-5; MedMCQA 74.2; PubMedQA 76.8; MMLU-Med 87.0; AfriMed-QA 84.0; exceeds physician performance on AgentClinic-MedQA (developer-reported).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

HuatuoGPT-o1-8B / 70B / 72B

Vendor: FreedomIntelligence (CUHK-Shenzhen / SRIBD)

What it does: works through a differential step by step and shows its reasoning, which is what makes an answer reviewable rather than merely plausible. The 8B version is the best in its size class, so the reasoning does not have to cost 70B hardware.

8B: MedQA 72.6, MedMCQA 60.4, PubMedQA 79.2, average 63.9 across 7 benchmarks — best in its size class (independent, arXiv 2412.18925).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

Meditron-7B / 70B, Meditron3-8B / 70B

Vendor: EPFL (LiGHT lab) / Yale

What it does: answers from a base trained on medical literature and guidelines, which shows in how it phrases clinical reasoning. Meditron3 follows instructions more closely, which matters when output must fit a fixed template.

Meditron-70B about 70 percent MedQA; Meditron3 adds a Llama-3.1 base and better instruction following (independent).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

Med42-v2 8B / 70B

Vendor: M42 Health (Abu Dhabi)

What it does: is tuned for clinical instruction following, so it stays inside the format and scope you give it. Strong on examination-style questions at both sizes.

About 72 percent (8B) and about 80 percent (70B) MedQA; strong on USMLE-style sets (developer-reported).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

OpenBioLLM-8B / 70B

Vendor: Saama AI Research

What it does: answers biomedical and clinical questions and holds up well on published QA suites. Note that independent reproduction came in below the original developer claims — we benchmark it against your own material.

70B: MedQA 76.1, MedMCQA 74.7, PubMedQA 79.2 (independent reproduction; developer claims were higher).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

UltraMedical-8B / 70B

Vendor: Tsinghua University (UltraMedical project)

What it does: targets medical question answering specifically, and the 70B version is competitive with much larger general models on the standard sets.

70B: MedQA 82.2, MedMCQA 71.8; 8B: MedQA 71.1 (independent).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

BioMistral-7B

Vendor: Avignon Université / Nantes Université

What it does: runs on a single mid-range card, which puts a medical-tuned model within reach of high-volume, narrow tasks where a 70B model would never pay for itself. Accuracy is correspondingly modest.

MedQA 45.0, MedMCQA 40.2, PubMedQA 66.9 (independent, HuatuoGPT-o1 comparison table).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

Me-LLaMA

Vendor: Me-LLaMA authors

What it does: is a medical-domain language model for question answering, generation and embeddings from one deployment, useful where several text tasks share a pipeline.

Benchmark-dependent; evaluated across medical NLP and QA tasks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

MedAlpaca-7B / MedAlpaca-13B

Vendor: MedAlpaca authors

What it does: holds a conversational medical exchange, and is small and well understood enough to serve as a baseline before a larger model is justified.

Benchmark-dependent; conversational medical model, no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

PMC-LLaMA-7B / PMC-LLaMA-13B

Vendor: PMC-LLaMA authors

What it does: answers from a base trained on biomedical literature, which suits question answering grounded in published evidence rather than in hospital notes.

Benchmark-dependent; evaluated on medical QA tasks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Asclepius-Llama3-8B

Vendor: Stanford and community (Asclepius project)

What it does: summarises and answers questions over clinical notes, and is trained entirely on synthetic notes — so it is redistributable without the licence friction most clinical models carry.

Approaches GPT-3.5 on clinical note QA and summarisation; trained purely on synthetic notes so it is redistributable (developer-reported, ACL).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

MedFound-7B / MedFound-176B

Vendor: Chinese Academy of Sciences and collaborators

What it does: reads a clinical narrative and returns a diagnosis list with its reasoning. Developer evaluation reports accuracy competitive with clinicians across specialties, on internal data.

Diagnostic accuracy competitive with clinicians on internal multi-specialty evaluation (developer-reported, Nature Medicine 2025).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

Diabetica-7B / Diabetica-o1

Vendor: Shanghai Jiao Tong University

What it does: is built for one disease area and beats much larger general models inside it — the case for a specialist model when your scope is narrow.

Outperforms GPT-4 on diabetes-specific QA sets (developer-reported).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

Qwen3 (8B / 32B / 235B-A22B)

Vendor: Alibaba Qwen

What it does: is a strong general model, not a medical one: paired with retrieval over your own guidelines it closes most of the gap, and its permissive licence removes a common blocker.

Strong general reasoning; medical performance below dedicated medical tunes unless paired with guideline retrieval (independent).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

DeepSeek-R1 / R1-distill variants

Vendor: DeepSeek AI

What it does: produces long chain-of-thought reasoning, which is what complex differentials need. It is the heaviest option here, and the distilled variants exist for exactly that reason.

MedQA about 90 percent for the full model; heavy compute, strong chain-of-thought for complex differentials (independent).

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 420≈ 1,100

Input type — Clinical text — multilingual

Models in this group are built for languages other than English, or for several languages at once. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

Apollo / Apollo2 (0.5B–7B, 6 languages)

Vendor: FreedomIntelligence

What it does: answers medical questions across six languages at small sizes, with notably strong Arabic, Hindi, Spanish and French coverage — the practical choice for multilingual services on modest hardware.

Best-in-size multilingual medical QA; strong Arabic, Hindi, Spanish and French coverage (developer-reported).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

MMedLM-2

Vendor: OpenMedLab / MMedLM authors

What it does: answers medical questions across languages from one deployment, rather than requiring a separate model per language.

Benchmark-dependent; evaluated on multilingual medical QA benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

BiMediX

Vendor: BiMediX authors

What it does: works in English and Arabic together, which suits bilingual clinical services where switching models mid-conversation is not an option.

Benchmark-dependent; evaluated on bilingual English/Arabic medical tasks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

HuatuoGPT-II 7B / 13B / 34B

Vendor: FreedomIntelligence

What it does: handles Chinese clinical dialogue and surpassed commercial Chinese medical models at release, with three sizes to fit the hardware you have.

CMB 60.4 (7B) and 63.0 (13B), surpassing commercial Chinese medical models at release (independent, CMB).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

DISC-MedLLM-13B

Vendor: Fudan University

What it does: answers Chinese medical questions, and was among the leading open Chinese medical models at release.

CMB average about 39.8; second among open Chinese medical models at release (independent, CMB benchmark).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Zhongjing (Ling-Dan) 13B

Vendor: Zhejiang University / CMKRG

What it does: was built for multi-turn Chinese consultation quality specifically, and shipped with its own evaluation dataset — unusual transparency in this group.

Leading multi-turn Chinese consultation quality at release; CMtMedQA dataset released alongside (developer-reported, AAAI 2024).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

BenTsao (HuaTuo)

Vendor: HuaTuo / BenTsao authors

What it does: answers Chinese medical questions at a small size, useful as a baseline or for high-volume narrow tasks.

Benchmark-dependent; Chinese medical QA performance varies by dataset.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 3,200≈ 8,100

PULSE

Vendor: PULSE / OpenMedLab authors

What it does: answers Chinese medical questions across general clinical scope.

Benchmark-dependent; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Taiyi

Vendor: Taiyi authors

What it does: works across Chinese and English biomedical text, returning entities, labels and answers — a bilingual option for extraction as well as question answering.

Task-dependent across biomedical NLP and QA benchmarks.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 350≈ 1,200≈ 3,200

Choosing between them

Start from the language mix and the volume, then the reasoning depth you actually need — retrieval over your own guidelines closes more of the gap than a larger model does. Licensing varies widely here, from Apache-2.0 to bespoke terms that restrict medical claims, and we check it against your intended use before anything is deployed.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Clinical reasoning assistant service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for clinical reasoning assistant. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.