Document visual question answering answers a question about a page by looking at the page as an image — “what is the total on this invoice?”, “which box is ticked?”, “what does the chart show for March?” — rather than by reading text pulled out of it.
This matters because extraction throws away layout. A figure’s meaning often lives in its position: the column it sits under, the box it sits inside, the tick beside it. Models here read the pixels and the words together, so they answer correctly on scanned forms, slides, charts and stamped or handwritten documents where a text-only pipeline loses the structure that carried the meaning.
Models in this group take text or whole documents as input: plain text, PDFs, scanned pages and office files. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page at 300 dpi with one question asked about it.
Qwen2.5-VL 7B / 72B
Vendor: Alibaba Cloud
What it does: looks at a page image and answers questions about it in plain language, including questions about charts, stamps and handwriting. The 7B and 72B labels are model sizes in billions of parameters; the larger is more accurate and slower.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (7B, reduced precision)
1× A100 80 GB (7B, full precision)
2× H100 80 GB (72B model)
VRAM
16 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (pages/hour)
≈ 400
≈ 1,600
≈ 900 (72B model, higher accuracy)
InternVL 3
Vendor: OpenGVLab (Shanghai AI Laboratory)
What it does: reads page images and answers questions about them, with particular strength on charts and dense tables. It comes in several sizes; the 8-billion-parameter version is the usual balance for on-premise deployment.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (2B model)
1× A100 80 GB (8B model)
4× H100 80 GB (38B model or larger)
VRAM
12 GB
40 GB
320 GB combined
vCPUs
8
16
48
RAM
32 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (pages/hour)
≈ 550
≈ 1,800
≈ 1,200
LayoutLMv3
Vendor: Microsoft
What it does: combines the words on a page with their positions to answer questions about forms and structured documents. It is small and, once fitted to one recurring document type, more accurate on it than a much larger general model.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
1× A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 1,600
≈ 6,000
≈ 15,000
Donut
Vendor: NAVER Clova
What it does: answers questions straight from the page image with no OCR step at all, which removes a whole source of error on low-quality scans. Best when fitted to a single document type such as one supplier’s invoice.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
1× A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 900
≈ 3,200
≈ 8,000
Pix2Struct
Vendor: Google
What it does: a model trained on screenshots and rendered pages that answers questions about charts, user interfaces and infographics. Useful where the answer must be read off a graph rather than a line of text.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
1× A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 800
≈ 2,800
≈ 7,000
Input type — Pictures
Models in this group take a single image as input: a photograph of a page, a screenshot or a scanned form. The three hardware tiers mean the same as above. The sample input here is one photograph of a page or screen at about 2 megapixels, with one question asked about it.
Llama 3.2 Vision
Vendor: Meta
What it does: answers questions about photographs of documents, screens and signage in ordinary language. A sound general choice when the input is whatever staff happen to photograph rather than a controlled scan.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (11B model)
1× A100 80 GB (11B model)
2× H100 80 GB (90B model)
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (images/hour)
≈ 450
≈ 1,500
≈ 800
ColPali
Vendor: Illuin Technology
What it does: finds which page in a large archive answers a question by comparing page images directly, so the right slide or scanned form is retrieved before any question is answered about it.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
1× A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (images/hour)
≈ 900
≈ 3,500
≈ 7,000
Choosing between them
Accuracy depends on how much your documents rely on layout and how varied they are. Our consultants review a sample of your pages and the questions you need answered, then recommend either a small purpose-built model for one recurring document type or a general vision-language model for a varied intake.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Document visual question answering. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.