Skip to main content

Model reference — AI Medical · Clinical language

AI models for EHR outcome prediction

The reported gains are real but measured: AUROC improvements of 0.02–0.08 over gradient boosting across mortality, readmission and diagnosis-onset tasks. On a large population that is worth having; it is not a step change.

Next-event models report top-10 accuracy around 0.68–0.78 depending on site, and that site dependence is the point — these models are more sensitive to population and coding practice than any other group on this site.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

EHR outcome prediction service AI Medical services Pricing

Input type — Structured EHR timelines

Models in this group take longitudinal coded record data rather than free text; the rate is patient timelines per hour. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

CLMBR / MOTOR

Vendor: Stanford (Shah lab)

What it does: turns a coded patient timeline into embeddings and time-to-event risk scores, beating gradient boosting by 0.02–0.08 AUROC across mortality, readmission and diagnosis-onset — and the same embedding serves several endpoints.

AUROC gains of 0.02–0.08 over gradient boosting across mortality, readmission and diagnosis-onset tasks (independent, EHRSHOT).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

Foresight / ETHOS

Vendor: Community, including Google Health-derived open reimplementations

What it does: forecasts the next clinical events with probabilities, at top-10 accuracy around 0.68–0.78. Site dependence is high, which is why local training is part of the work rather than an optional extra.

Next-event prediction top-10 accuracy about 0.68–0.78 depending on site (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

EHRMamba

Vendor: EHRMamba authors

What it does: models long structured record sequences efficiently, which matters when a patient history runs to thousands of events.

Task-dependent; evaluated on EHR prediction tasks using MIMIC-IV.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (documents/hour)≈ 6,000≈ 21,000≈ 54,000

Choosing between them

Your data model decides the shortlist: OMOP or FHIR event streams suit the timeline foundation models directly. Expect local training rather than out-of-the-box use, and expect the honest comparison to be against your existing gradient-boosted baseline, which is what we benchmark against. Some weights in this group are released on request, which we handle for you.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

EHR outcome prediction service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for ehr outcome prediction. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.