Extraction turns clinical prose into fields. Finding the words is the easy part; deciding what they mean in context — present, ruled out, historical, someone else’s — is the work.
Three layers usually do it. A reading layer handles scanned material and preserves layout. An extraction layer names entities and their attributes. A terminology layer maps what was found onto a coding system with context flags. A large medical model can do all three at once at far higher cost per document, which is why it is normally reserved for the cases the layers above cannot resolve.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
GLiNER-biomed
Vendor: Knowledgator
What it does: names the entity types you ask for — conditions, medications, dosages, findings — without training a separate model per type. Fast and the usual starting point.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 12,000
≈ 45,000
≈ 120,000
MedCAT v2
Vendor: King’s College London
What it does: maps free-text mentions onto SNOMED CT and UMLS concepts and flags negation, history and family context. The terminology layer most coding work needs.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 9,000
≈ 32,000
≈ 90,000
Clinical-Longformer
Vendor: Northwestern University
What it does: a clinical language model that reads very long documents in one pass, for classification and extraction across a whole admission record.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 15,000
≈ 55,000
≈ 150,000
MedGemma 27B
Vendor: Google DeepMind
What it does: reads a whole record and returns structured fields including coding candidates with the supporting sentence. Expensive per document, and the option for cases the layers above cannot resolve.
Requirement
Minimum
Medium
High
GPU type
A100 80 GB
A100 80 GB ×2
H100 80 GB ×4
VRAM
80 GB
160 GB
320 GB
vCPUs
24
48
96
RAM
128 GB
256 GB
512 GB
Server
1× A100 SXM 80 GB
2× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (documents/hour)
≈ 600
≈ 2,200
≈ 6,000
LayoutLMv3
Vendor: Microsoft Research
What it does: reads text together with its position on the page, so scanned forms and tables are extracted as forms and tables rather than as a stream of words.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 6,000
≈ 22,000
≈ 60,000
Surya
Vendor: Vik Paruchuri
What it does: reads scanned and faxed records — text, lines and reading order — as the first step before any extraction runs.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 4,000
≈ 15,000
≈ 42,000
Choosing between them
Document quality decides whether you need a reading layer at all; your coding requirements decide whether terminology mapping is a separate step or a job for a large model. We run the shortlist over a sample of your own documents and report field-level accuracy — the number a coding or registry team can actually act on.
At the start of a project we run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. The estimates on this page are replaced with real figures, so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for medical record extraction. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.