This service produces subtitle files from a video: text divided into lines and cues, timed to the speech, with speaker labels where needed, ready to be delivered as SRT, WebVTT or burnt into the picture.
A subtitle is not a transcript. It must obey rules a transcript does not: a maximum number of characters per line and lines per cue, a minimum time on screen, a reading speed a viewer can keep up with, and line breaks that fall at sensible points in the sentence. Those rules, and the alignment that times each cue to the frame, are what separate usable subtitles from a wall of text. Translated subtitles add a further constraint, since a translation that is longer than the original must still fit the same time on screen.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video at 25 frames per second.
faster-whisper
Vendor: SYSTRAN
What it does: transcribes the speech that the subtitles are built from, quickly and in a hundred languages. The foundation of the pipeline.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 1,200
≈ 4,200
≈ 10,800
WhisperX
Vendor: University of Oxford (VGG)
What it does: aligns the words to the audio precisely, which is what makes cues appear and disappear with the speech rather than drifting.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 900
≈ 3,300
≈ 8,400
pyannote.audio 3
Vendor: pyannote (Hervé Bredin)
What it does: identifies speaker changes, which subtitle standards require to be marked when more than one person speaks.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 1,200
≈ 4,200
≈ 10,800
Canary
Vendor: NVIDIA
What it does: transcribes and translates in one pass, giving a foreign-language subtitle track without a separate translation stage.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 900
≈ 3,000
≈ 7,200
NLLB-200
Vendor: Meta
What it does: translates subtitle text into two hundred languages, including many with no commercial subtitling service available.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 2,400
≈ 9,000
≈ 20,000
Llama 3.3 70B
Vendor: Meta
What it does: rewrites translated lines to fit the character and reading-speed limits without losing meaning, which is the step that makes translated subtitles actually readable.
Requirement
Minimum
Medium
High
GPU type
2× RTX 4090 (reduced precision)
H100 80 GB
4× H100 80 GB
VRAM
48 GB combined
80 GB
320 GB combined
vCPUs
16
24
64
RAM
64 GB
128 GB
512 GB
Server
2× RTX 4090 24 GB
1× H100 SXM 80 GB
4× H100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 600
≈ 2,400
≈ 9,000
FFmpeg
Vendor: FFmpeg project
What it does: muxes the subtitle track into the file or burns it into the picture, at the standard your distribution requires.
Requirement
Minimum
Medium
High
GPU type
No GPU required
No GPU required
GPU-accelerated decode (NVENC/NVDEC)
VRAM
—
—
8 GB
vCPUs
2
8
16
RAM
4 GB
16 GB
32 GB
Server
CPU instance, 2 vCPU
CPU instance, 8 vCPU
1× RTX 4090 24 GB
Rate (video minutes processed per hour)
≈ 6,000
≈ 20,000
≈ 60,000
Choosing between them
The right pipeline depends on your delivery standard, languages and whether subtitles are reviewed before publication. Our consultants review your material and subtitle specification, then recommend the models and the cue rules, and where human review is warranted.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Automatic subtitles / captions. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.