Skip to main content

Model reference — Documents

AI models for Multilingual translation

Translation converts documents, messages and interface text from one language into another, on your own hardware, so that material which cannot legally or commercially leave your network can still be read across the business.

The models divide by breadth and by quality. Purpose-built translation models cover two hundred languages including many with little training data available, and are small and fast. Large language models translate fewer languages but produce noticeably better prose in the major ones, and they can be told to keep a glossary, hold a formal register, or leave product names untranslated. Where terminology must be exact — legal, medical, engineering — the model is paired with your own term list, which it is required to follow.

Multilingual translation service AI models for documents

Input type — Documents

Models in this group take text or whole documents as input: plain text, PDFs, scanned pages and office files. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page of about 500 words translated into one target language.

NLLB-200

Vendor: Meta

What it does: translates between two hundred languages, including many that other models do not cover at all. The choice when breadth matters more than polish — internal documents, archives, low-resource languages.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (pages/hour)≈ 700≈ 2,800≈ 6,000

MADLAD-400

Vendor: Google

What it does: covers over four hundred languages and handles noisy input such as scanned text and informal messages without breaking down.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (pages/hour)≈ 600≈ 2,400≈ 5,200

Opus-MT / MarianMT

Vendor: University of Helsinki

What it does: small single-pair models — one per language direction — that run on almost any hardware. Where you translate only a few fixed pairs at volume, this is by far the cheapest option per page.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (pages/hour)≈ 4,000≈ 16,000≈ 45,000

Tower

Vendor: Unbabel and Instituto Superior Técnico

What it does: a translation-tuned language model that produces publication-quality prose in the major European languages and follows a glossary reliably.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (reduced precision)L40S 48 GB2× H100 80 GB
VRAM22 GB48 GB160 GB combined
vCPUs121648
RAM48 GB64 GB256 GB
Server1× RTX 4090 24 GB1× L40S 48 GB2× H100 SXM 80 GB
Rate (pages/hour)≈ 350≈ 1,100≈ 4,200

Qwen2.5 32B

Vendor: Alibaba Cloud

What it does: translates with the fluency of a general language model and holds context across a long document, so pronouns and terminology stay consistent from page to page.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (reduced precision)L40S 48 GB2× H100 80 GB
VRAM22 GB48 GB160 GB combined
vCPUs121648
RAM48 GB64 GB256 GB
Server1× RTX 4090 24 GB1× L40S 48 GB2× H100 SXM 80 GB
Rate (pages/hour)≈ 350≈ 1,100≈ 4,200

Llama 3.3 70B

Vendor: Meta

What it does: the highest-quality option for material that will be published, able to hold register, tone and a glossary across a whole document. Slower and dearer per page.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (pages/hour)≈ 250≈ 1,100≈ 4,500

SeamlessM4T v2

Vendor: Meta

What it does: translates text and speech in the same model, which matters when the same content must appear as a document and as an audio track.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (pages/hour)≈ 500≈ 2,000≈ 4,000

Choosing between them

The right model depends on your language pairs, whether the output is read internally or published, and how strict your terminology is. Our consultants review your material and glossary, then recommend a model per language pair and the review step for anything customer-facing.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Multilingual translation service AI models for documents Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Multilingual translation. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.