Skip to main content

Model reference — Video

AI models for Automatic subtitles / captions

This service produces subtitle files from a video: text divided into lines and cues, timed to the speech, with speaker labels where needed, ready to be delivered as SRT, WebVTT or burnt into the picture.

A subtitle is not a transcript. It must obey rules a transcript does not: a maximum number of characters per line and lines per cue, a minimum time on screen, a reading speed a viewer can keep up with, and line breaks that fall at sensible points in the sentence. Those rules, and the alignment that times each cue to the frame, are what separate usable subtitles from a wall of text. Translated subtitles add a further constraint, since a translation that is longer than the original must still fit the same time on screen.

Automatic subtitles / captions service AI models for video files

Input type — Video

Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video at 25 frames per second.

faster-whisper

Vendor: SYSTRAN

What it does: transcribes the speech that the subtitles are built from, quickly and in a hundred languages. The foundation of the pipeline.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 1,200≈ 4,200≈ 10,800

WhisperX

Vendor: University of Oxford (VGG)

What it does: aligns the words to the audio precisely, which is what makes cues appear and disappear with the speech rather than drifting.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 900≈ 3,300≈ 8,400

pyannote.audio 3

Vendor: pyannote (Hervé Bredin)

What it does: identifies speaker changes, which subtitle standards require to be marked when more than one person speaks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 1,200≈ 4,200≈ 10,800

Canary

Vendor: NVIDIA

What it does: transcribes and translates in one pass, giving a foreign-language subtitle track without a separate translation stage.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 900≈ 3,000≈ 7,200

NLLB-200

Vendor: Meta

What it does: translates subtitle text into two hundred languages, including many with no commercial subtitling service available.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 2,400≈ 9,000≈ 20,000

Llama 3.3 70B

Vendor: Meta

What it does: rewrites translated lines to fit the character and reading-speed limits without losing meaning, which is the step that makes translated subtitles actually readable.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (video minutes processed per hour)≈ 600≈ 2,400≈ 9,000

FFmpeg

Vendor: FFmpeg project

What it does: muxes the subtitle track into the file or burns it into the picture, at the standard your distribution requires.

RequirementMinimumMediumHigh
GPU typeNo GPU requiredNo GPU requiredGPU-accelerated decode (NVENC/NVDEC)
VRAM8 GB
vCPUs2816
RAM4 GB16 GB32 GB
ServerCPU instance, 2 vCPUCPU instance, 8 vCPU1× RTX 4090 24 GB
Rate (video minutes processed per hour)≈ 6,000≈ 20,000≈ 60,000

Choosing between them

The right pipeline depends on your delivery standard, languages and whether subtitles are reviewed before publication. Our consultants review your material and subtitle specification, then recommend the models and the cue rules, and where human review is warranted.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Automatic subtitles / captions service AI models for video files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Automatic subtitles / captions. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.