Skip to main content

Model reference — AI Medical · Clinical language

AI models for Clinical text representation

These are the workhorses of clinical NLP and the numbers are well established: F1 around 0.87–0.90 on the standard clinical NER sets, with the larger clinical-corpus models leading on relation extraction and semantic similarity.

The choice is corpus and length. Biomedical-literature encoders are strong on published text, clinical-note encoders on hospital text, and long-document variants add a few points on tasks that span a whole record. One model in this group was trained on a 90-billion-word real clinical corpus, which shows on clinical rather than literature tasks.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Clinical text representation service AI Medical services Pricing

Input type — Clinical and biomedical text

Models in this group take text and return embeddings or task-specific labels. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

GatorTron (base 345M, medium 3.9B, large 8.9B) / GatorTronS

Vendor: University of Florida with NVIDIA

What it does: was trained on a 90-billion-word real clinical corpus, and it shows: best-in-class on clinical entity and relation extraction rather than on published literature, which is the distinction that matters for hospital text.

Best-in-class on clinical NER, relation extraction, semantic textual similarity and NLI over a 90-billion-word real clinical corpus (independent, npj Digital Medicine).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 2,000≈ 7,000≈ 18,000

BioBERT v1.2

Vendor: DMIS Lab / Korea University

What it does: is the long-standing biomedical encoder — a small, cheap and well-characterised base for literature-facing extraction.

Plus 0.6–2.8 F1 over BERT on biomedical NER, relation extraction and QA (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

Bio_ClinicalBERT / BioClinicalBERT

Vendor: MIT CSAIL / Alsentzer et al.

What it does: remains the default clinical baseline at F1 0.87–0.90 on the standard entity sets. Larger encoders beat it, and it is the sensible first benchmark before paying for them.

Strong baseline on i2b2 NER, F1 about 0.87–0.90; superseded by larger encoders but still the default clinical baseline (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

ClinicalBERT

Vendor: ClinicalBERT authors

What it does: encodes clinical notes for prediction and classification tasks, at a footprint small enough to run over a whole archive.

Task-dependent; performance varies by clinical prediction and NLP dataset.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

BiomedBERT (formerly PubMedBERT)

Vendor: Microsoft Research

What it does: is top-tier on the biomedical NLP benchmark suite and the better choice than a clinical encoder when your text is published literature.

Top-tier on the BLURB benchmark, average about 81–83 (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

BioLinkBERT

Vendor: BioLinkBERT authors

What it does: brings document links into the representation, which helps on biomedical question answering across connected sources.

Task-dependent; evaluated on biomedical QA and NLP benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (documents/hour)≈ 12,000≈ 42,000≈ 108,000

Clinical-Longformer / Clinical-BigBird

Vendor: Yale and community

What it does: reads up to 4096 tokens, so a whole admission record fits in one pass. That is worth 2–5 points on readmission and phenotyping, where truncation loses exactly the part the model needs.

Plus 2–5 points over ClinicalBERT on long-document tasks such as readmission and phenotyping (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

Clinical-T5

Vendor: Clinical-T5 authors / PhysioNet

What it does: generates as well as encodes, which suits summarisation and structured rewriting of clinical text in the same deployment.

Task-dependent; evaluated on MIMIC clinical text tasks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

BioGPT

Vendor: Microsoft Research

What it does: generates biomedical text and extracts relations, for literature-facing pipelines rather than clinical notes.

Task-dependent; evaluated on biomedical relation extraction and text-generation benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

Choosing between them

Match the encoder to your text: literature, clinical notes, or long records. Where a task spans an entire admission, a long-document model is worth the extra hardware; where it does not, the base encoders are several times cheaper per document. All of these fine-tune on a modest labelled set.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Clinical text representation service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for clinical text representation. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.