Skip to main content

Model reference — AI Medical · Clinical language

AI models for Clinical coding support

The honest numbers here are modest and worth stating plainly: micro-F1 around 0.60 and precision-at-8 around 0.77 on the full ICD label space. That is not autonomous coding, and no released model is. It is a ranked shortlist that shortens a coder’s work.

Concept linking is stronger, with detection F1 0.84–0.93 across UK hospital corpora for the established terminology tool. Zero-shot extraction models sit between the two and beat prompting a general model at a fraction of the cost.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Clinical coding support service AI Medical services Pricing

Input type — Clinical documents

Models in this group take clinical free text and return codes, concepts or typed entities. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

PLM-ICD

Vendor: Academic community (CAML and PLM-ICD lineage)

What it does: ranks ICD codes from a discharge summary. Micro-F1 around 0.60 is not autonomous coding, and no released model is — it is a shortlist that shortens a coder’s work rather than replacing it.

Micro-F1 about 0.60, precision@8 about 0.77 on the MIMIC-III full ICD-9 label space (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

MedCAT / MedCAT v2

Vendor: King’s College London / UCLH

What it does: links free-text mentions to SNOMED CT and UMLS at F1 0.84–0.93 across UK hospital corpora, and flags negation, temporality and experiencer — so a ruled-out condition is not coded as present. Elastic-licensed, which we check against your use.

Concept detection F1 0.84–0.93 across UK hospital corpora (independent, Lancet Digital Health and npj Digital Medicine).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

scispaCy (en_core_sci_lg, en_ner_bc5cdr_md and others)

Vendor: Allen Institute for AI

What it does: extracts entities, abbreviations and UMLS concept IDs quickly and cheaply, and is the usual first layer in a coding pipeline.

NER F1 0.84–0.87 on BC5CDR and JNLPBA; entity linking accuracy about 0.75 to UMLS (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

medspaCy (ConText, Sectionizer, cTAKES-style rules)

Vendor: medspaCy community (UVA / VA)

What it does: adds assertion status and document sections through rules rather than learning, at F1 above 0.90 on the standard assertion set — and rules are auditable, which coding teams value.

Negation and assertion F1 above 0.90 on i2b2 assertion data (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

GLiNER-biomed / NuNER-medical

Vendor: Knowledgator and community

What it does: extracts whatever entity types you name at inference time, with no model per type, and beats prompting a general model at a fraction of the cost.

Zero-shot biomedical NER F1 0.55–0.70; beats prompting a general LLM at a fraction of the cost (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

Choosing between them

If the goal is revenue integrity, start with concept linking and your own rules on top — it is more accurate and more auditable than end-to-end code prediction. Use the code-ranking model as a shortlist layer, not a decision. Licences vary, including one Elastic-licensed component we check against your use.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Clinical coding support service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for clinical coding support. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.