Handwriting recognition reads handwritten text from a photograph or scan and returns it as typed characters — filled-in forms, engineers’ notes, delivery signatures, historical registers, clinical annotations.
It is markedly harder than reading print, and accuracy varies far more with the material than with the model. Neat block capitals on a ruled form read almost perfectly; a doctor’s cursive on a crumpled slip may not be readable by a person either. Two things reliably improve results: constraining what the model expects, since a field known to hold a date is far easier than free text, and correcting against a known list, since a surname checked against your customer database beats any general model. We measure accuracy per field rather than per page, because a form is usually only as good as its worst field.
Models in this group take a single image as input: a photograph, a scan or a screenshot. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page at 300 dpi carrying handwritten entries.
TrOCR
Vendor: Microsoft
What it does: reads handwritten lines and returns them as text, purpose-built for handwriting rather than adapted from print reading. The usual first choice, and it can be fitted to a specific hand or form.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 2,000
≈ 8,000
≈ 20,000
Surya
Vendor: Datalab
What it does: reads handwriting alongside print in roughly ninety languages and returns each word with its position, which suits mixed forms and multilingual archives.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 700
≈ 3,000
≈ 6,000
GOT-OCR 2.0
Vendor: StepFun
What it does: reads a page in one step including handwritten passages, and outputs structure as well as text. Small enough to deploy at a branch or depot.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 600
≈ 2,200
≈ 4,500
PaddleOCR
Vendor: Baidu (PaddlePaddle)
What it does: a fast reader with a handwriting mode and strong Chinese support, effective on high volumes of forms.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 3,000
≈ 12,000
≈ 34,000
Qwen2.5-VL 7B / 72B
Vendor: Alibaba Cloud
What it does: reads difficult handwriting by using context — inferring an unclear word from the sentence around it — which is where line-by-line readers fail. The most accurate option on hard material, and the slowest.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (7B, reduced precision)
A100 80 GB (7B, full precision)
2× H100 80 GB (72B model)
VRAM
16 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (pages/hour)
≈ 400
≈ 1,600
≈ 900 (72B model, higher accuracy)
InternVL 3
Vendor: OpenGVLab (Shanghai AI Laboratory)
What it does: the same context-driven reading with particular strength on dense handwritten tables and ledgers.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (2B model)
A100 80 GB (8B model)
4× H100 80 GB (38B model or larger)
VRAM
12 GB
40 GB
320 GB combined
vCPUs
8
16
48
RAM
32 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (pages/hour)
≈ 550
≈ 1,800
≈ 1,200
EasyOCR
Vendor: Jaided AI
What it does: a light reader covering eighty languages, useful as a quick first pass to find which pages contain handwriting at all before a heavier model is applied.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 4,000
≈ 16,000
≈ 45,000
Choosing between them
What is achievable depends on the writing itself, the capture quality and how much context constrains each field. Our consultants measure per-field accuracy on a sample of your own documents, then recommend a model, the constraints worth applying, and the fields that will still need a person to confirm them.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Handwriting recognition. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.