Document redaction removes sensitive content from a document and produces a copy that can be released — to a customer, a court, a regulator, or the public — with the removed material genuinely gone rather than merely hidden.
The distinction between hidden and gone is the whole point. A black rectangle drawn over a PDF leaves the text underneath, recoverable by anyone who copies it; a correct redaction removes the text and, where required, replaces the page with an image so nothing survives. That makes redaction a two-part problem: finding what must go, which needs models that locate sensitive content and its position on the page, and producing the output safely, which is engineering rather than modelling. We do both, and verify the result by searching the released file for what should no longer be in it.
Models in this group take text or whole documents as input: plain text, PDFs, scanned pages and office files. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page at 300 dpi redacted and verified.
Presidio
Vendor: Microsoft
What it does: finds personal data and structured identifiers and applies your chosen masking. The usual backbone, because it covers both the detection and the replacement in one framework.
Requirement
Minimum
Medium
High
GPU type
No GPU required
RTX 4090
A100 80 GB
VRAM
—
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
CPU instance, 4 vCPU
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 5,000
≈ 20,000
≈ 60,000
GLiNER
Vendor: Urchade Zaratiana and contributors
What it does: finds the categories specific to your case — a witness name, a supplier price, an unreleased product code — from a written list, without training data.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 3,000
≈ 12,000
≈ 34,000
LayoutLMv3
Vendor: Microsoft
What it does: locates sensitive content and its exact position on the page, which is what allows the area to be removed rather than the whole page suppressed.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 1,600
≈ 6,000
≈ 15,000
Surya
Vendor: Datalab
What it does: reads scanned pages in roughly ninety languages and returns each word with its coordinates, so scanned material can be redacted as precisely as born-digital text.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 700
≈ 3,000
≈ 6,000
PaddleOCR
Vendor: Baidu (PaddlePaddle)
What it does: a fast character reader that returns word positions, used as the reading stage under a detection model on large scanned volumes.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 3,000
≈ 12,000
≈ 34,000
Input type — Pictures
Models in this group take a page image or photograph as input, and return the regions to be removed. The three hardware tiers mean the same as above. The sample input here is one page image at about 2 megapixels with the regions to redact identified.
Qwen2.5-VL 7B / 72B
Vendor: Alibaba Cloud
What it does: reads a page image and identifies sensitive regions described in plain words — "any signature", "any photograph of a person", "any handwritten note" — which pattern-based detection cannot express.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (7B, reduced precision)
A100 80 GB (7B, full precision)
2× H100 80 GB (72B model)
VRAM
16 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (images/hour)
≈ 400
≈ 1,600
≈ 900 (72B model, higher accuracy)
Florence-2
Vendor: Microsoft
What it does: a small vision model that locates faces, signatures and stamps on a page quickly enough to run over an entire archive as a first pass.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 1,800
≈ 6,000
≈ 14,000
Choosing between them
Requirements differ sharply between a subject access request, a court exhibit and a public release. Our consultants review your obligations and a sample of your documents, then recommend the detection models, the output format, and the verification step that proves the released file is clean.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Document redaction. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.