Summarization models read clinical text and write shorter clinical text. Fluency is not the problem — omission is. A summary that reads well and drops an allergy is worse than no summary at all.
Model size sets the trade-off. Large models follow a template closely, hold a long record in one pass and omit less, at the cost of several GPUs each. Smaller medical models run on a single card and suit short, well-structured documents where the input already resembles the output.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
MedGemma 27B
Vendor: Google DeepMind
What it does: drafts summaries and letters in your template and holds a long record in context. The strongest general option at a size a single server can run.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (notes/hour)
≈ 600
≈ 2,200
≈ 6,000
Meditron 3 70B
Vendor: EPFL
What it does: a large model trained on medical literature and guidelines, which shows in how it phrases clinical reasoning in a summary.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (notes/hour)
≈ 200
≈ 700
≈ 2,000
OpenBioLLM 70B
Vendor: Saama AI Research
What it does: a large biomedical model that performs well on clinical question answering, used where the summary must also answer specific questions about the record.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (notes/hour)
≈ 200
≈ 700
≈ 2,000
Med42-v2 70B
Vendor: M42 Health
What it does: a clinically aligned large model tuned to follow instructions closely, which matters when the output must match a fixed hospital template.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (notes/hour)
≈ 220
≈ 750
≈ 2,100
BioMistral 7B
Vendor: Avignon Université and Nantes Université
What it does: small enough to run on one mid-range card, suitable for short notes and high volume where the input is already structured.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (notes/hour)
≈ 1,800
≈ 6,500
≈ 18,000
Qwen2.5 72B (medical fine-tune)
Vendor: Alibaba Cloud
What it does: a strong general model fitted to medical text, notable for handling records in languages other than English.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (notes/hour)
≈ 180
≈ 650
≈ 1,800
Choosing between them
Record length, template complexity and how much editing your clinicians will accept decide the size of model you need. We benchmark candidates against documents your own clinicians have already written and signed, and we measure omission rather than readability. Cluster sizing then follows from your real daily volume.
At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for clinical note summarization. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.