Skip to main content

Model reference — Documents · Pictures

AI models for Document visual question answering

Document visual question answering answers a question about a page by looking at the page as an image — “what is the total on this invoice?”, “which box is ticked?”, “what does the chart show for March?” — rather than by reading text pulled out of it.

This matters because extraction throws away layout. A figure’s meaning often lives in its position: the column it sits under, the box it sits inside, the tick beside it. Models here read the pixels and the words together, so they answer correctly on scanned forms, slides, charts and stamped or handwritten documents where a text-only pipeline loses the structure that carried the meaning.

Document visual question answering service AI models for documents AI models for pictures

Input type — Documents

Models in this group take text or whole documents as input: plain text, PDFs, scanned pages and office files. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page at 300 dpi with one question asked about it.

Qwen2.5-VL 7B / 72B

Vendor: Alibaba Cloud

What it does: looks at a page image and answers questions about it in plain language, including questions about charts, stamps and handwriting. The 7B and 72B labels are model sizes in billions of parameters; the larger is more accurate and slower.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (7B, reduced precision)1× A100 80 GB (7B, full precision)2× H100 80 GB (72B model)
VRAM16 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB200 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (pages/hour)≈ 400≈ 1,600≈ 900 (72B model, higher accuracy)

InternVL 3

Vendor: OpenGVLab (Shanghai AI Laboratory)

What it does: reads page images and answers questions about them, with particular strength on charts and dense tables. It comes in several sizes; the 8-billion-parameter version is the usual balance for on-premise deployment.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (2B model)1× A100 80 GB (8B model)4× H100 80 GB (38B model or larger)
VRAM12 GB40 GB320 GB combined
vCPUs81648
RAM32 GB64 GB256 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (pages/hour)≈ 550≈ 1,800≈ 1,200

LayoutLMv3

Vendor: Microsoft

What it does: combines the words on a page with their positions to answer questions about forms and structured documents. It is small and, once fitted to one recurring document type, more accurate on it than a much larger general model.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 40901× A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (pages/hour)≈ 1,600≈ 6,000≈ 15,000

Donut

Vendor: NAVER Clova

What it does: answers questions straight from the page image with no OCR step at all, which removes a whole source of error on low-quality scans. Best when fitted to a single document type such as one supplier’s invoice.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 40901× A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (pages/hour)≈ 900≈ 3,200≈ 8,000

Pix2Struct

Vendor: Google

What it does: a model trained on screenshots and rendered pages that answers questions about charts, user interfaces and infographics. Useful where the answer must be read off a graph rather than a line of text.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 40901× A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (pages/hour)≈ 800≈ 2,800≈ 7,000

Input type — Pictures

Models in this group take a single image as input: a photograph of a page, a screenshot or a scanned form. The three hardware tiers mean the same as above. The sample input here is one photograph of a page or screen at about 2 megapixels, with one question asked about it.

Llama 3.2 Vision

Vendor: Meta

What it does: answers questions about photographs of documents, screens and signage in ordinary language. A sound general choice when the input is whatever staff happen to photograph rather than a controlled scan.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (11B model)1× A100 80 GB (11B model)2× H100 80 GB (90B model)
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB200 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (images/hour)≈ 450≈ 1,500≈ 800

ColPali

Vendor: Illuin Technology

What it does: finds which page in a large archive answers a question by comparing page images directly, so the right slide or scanned form is retrieved before any question is answered about it.

RequirementMinimumMediumHigh
GPU typeRTX 40901× A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (images/hour)≈ 900≈ 3,500≈ 7,000

Choosing between them

Accuracy depends on how much your documents rely on layout and how varied they are. Our consultants review a sample of your pages and the questions you need answered, then recommend either a small purpose-built model for one recurring document type or a general vision-language model for a varied intake.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Document visual question answering service AI models for documents AI models for pictures Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Document visual question answering. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.