Recall is the only metric that matters here, and the reported figures are strong: F1 about 0.97–0.99 on the standard de-identification set for transformer models, and PHI recall about 0.99 for the rule-heavy tool — at the cost of over-redaction, which is a trade most compliance teams accept.
Image de-identification splits in two. Header scrubbing is deterministic and solved. Burned-in pixel text needs OCR-based detection, with recall around 0.95–0.99 depending on the reading model used.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take clinical free text or structured fields. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
obi-deid-roberta-i2b2
Vendor: Obi and community
What it does: finds PHI spans at F1 0.97–0.99 on the standard set, and is the usual best balance between recall and keeping the clinical content you actually need.
F1 about 0.97–0.99 on i2b2 2014 de-identification (developer-reported).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (documents/hour)
≈ 12,000
≈ 42,000
≈ 108,000
Presidio
Vendor: Microsoft
What it does: covers free text and structured fields with configurable recognisers, at recall 0.90–0.97 on common categories. Local identifier formats and provider names always need tuning, which is part of the engagement.
Recall 0.90–0.97 for common PHI categories; needs clinical tuning for MRNs and provider names (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (documents/hour)
≈ 12,000
≈ 42,000
≈ 108,000
Philter (philter-ucsf)
Vendor: PhysioNet / OBI community
What it does: reaches about 0.99 PHI recall — the highest here — by over-redacting. That trade is the right one for some releases and unacceptable for others, and it is your call rather than ours.
PHI recall about 0.99 on UCSF notes at the cost of over-redaction (independent, JAMA Network Open).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (documents/hour)
≈ 12,000
≈ 42,000
≈ 108,000
Input type — DICOM studies
Models and tooling in this group take DICOM studies and return de-identified studies. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
CTP with pixel-burned-PHI detectors
Vendor: RSNA and community
What it does: scrubs DICOM headers deterministically and detects burned-in text on the pixels, which header scrubbing alone will always miss.
Header de-identification is deterministic; burned-in text detection recall about 0.95–0.99 depending on the OCR model (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (documents/hour)
≈ 12,000
≈ 42,000
≈ 108,000
Choosing between them
Decide first how much over-redaction you can live with: the highest-recall tool removes legitimate clinical content along with PHI. A transformer model tuned on your own note types is usually the better balance, and provider names and local identifier formats are the part that always needs local tuning. Imaging needs both header and pixel steps.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for clinical de-identification. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.