Translation converts documents, messages and interface text from one language into another, on your own hardware, so that material which cannot legally or commercially leave your network can still be read across the business.
The models divide by breadth and by quality. Purpose-built translation models cover two hundred languages including many with little training data available, and are small and fast. Large language models translate fewer languages but produce noticeably better prose in the major ones, and they can be told to keep a glossary, hold a formal register, or leave product names untranslated. Where terminology must be exact — legal, medical, engineering — the model is paired with your own term list, which it is required to follow.
Models in this group take text or whole documents as input: plain text, PDFs, scanned pages and office files. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one A4 page of about 500 words translated into one target language.
NLLB-200
Vendor: Meta
What it does: translates between two hundred languages, including many that other models do not cover at all. The choice when breadth matters more than polish — internal documents, archives, low-resource languages.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (pages/hour)
≈ 700
≈ 2,800
≈ 6,000
MADLAD-400
Vendor: Google
What it does: covers over four hundred languages and handles noisy input such as scanned text and informal messages without breaking down.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (pages/hour)
≈ 600
≈ 2,400
≈ 5,200
Opus-MT / MarianMT
Vendor: University of Helsinki
What it does: small single-pair models — one per language direction — that run on almost any hardware. Where you translate only a few fixed pairs at volume, this is by far the cheapest option per page.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (pages/hour)
≈ 4,000
≈ 16,000
≈ 45,000
Tower
Vendor: Unbabel and Instituto Superior Técnico
What it does: a translation-tuned language model that produces publication-quality prose in the major European languages and follows a glossary reliably.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (reduced precision)
L40S 48 GB
2× H100 80 GB
VRAM
22 GB
48 GB
160 GB combined
vCPUs
12
16
48
RAM
48 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× L40S 48 GB
2× H100 SXM 80 GB
Rate (pages/hour)
≈ 350
≈ 1,100
≈ 4,200
Qwen2.5 32B
Vendor: Alibaba Cloud
What it does: translates with the fluency of a general language model and holds context across a long document, so pronouns and terminology stay consistent from page to page.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (reduced precision)
L40S 48 GB
2× H100 80 GB
VRAM
22 GB
48 GB
160 GB combined
vCPUs
12
16
48
RAM
48 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× L40S 48 GB
2× H100 SXM 80 GB
Rate (pages/hour)
≈ 350
≈ 1,100
≈ 4,200
Llama 3.3 70B
Vendor: Meta
What it does: the highest-quality option for material that will be published, able to hold register, tone and a glossary across a whole document. Slower and dearer per page.
Requirement
Minimum
Medium
High
GPU type
2× RTX 4090 (reduced precision)
H100 80 GB
4× H100 80 GB
VRAM
48 GB combined
80 GB
320 GB combined
vCPUs
16
24
64
RAM
64 GB
128 GB
512 GB
Server
2× RTX 4090 24 GB
1× H100 SXM 80 GB
4× H100 SXM 80 GB
Rate (pages/hour)
≈ 250
≈ 1,100
≈ 4,500
SeamlessM4T v2
Vendor: Meta
What it does: translates text and speech in the same model, which matters when the same content must appear as a document and as an audio track.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (pages/hour)
≈ 500
≈ 2,000
≈ 4,000
Choosing between them
The right model depends on your language pairs, whether the output is read internally or published, and how strict your terminology is. Our consultants review your material and glossary, then recommend a model per language pair and the review step for anything customer-facing.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Multilingual translation. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.