AI models for Narrator to video - generate narration from text
This service writes narration onto a video from a script: the text is read aloud by a synthetic voice, timed to the picture, and mixed with the existing sound — which is how training material, product videos and multilingual versions get made without booking a studio.
The distinctive constraint is length. A shot lasts a fixed number of seconds, and the words must fit it; a script that runs long has to be shortened or spoken faster, and only one of those is acceptable. So a working pipeline pairs a speech model with a language model that adjusts the script to fit the available time. Two rules apply to every deployment: a cloned voice needs the documented consent of the person it belongs to, and generated audio should be watermarked so it can be identified as synthetic later.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of video narrated from about 150 words of script.
XTTS v2
Vendor: Coqui
What it does: clones a voice from a few seconds of sample and reads the script in it across seventeen languages, which is how one narrator’s voice carries a whole multilingual library.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 120
≈ 400
≈ 1,000
F5-TTS
Vendor: Shanghai Jiao Tong University
What it does: produces very natural narration with accurate emphasis and pacing, and is fast enough to regenerate a section on demand when the script changes.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 200
≈ 700
≈ 1,600
StyleTTS 2
Vendor: Columbia University
What it does: gives direct control over pace and emphasis, which is what allows narration to be stretched or compressed to fit a shot without sounding rushed.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 150
≈ 500
≈ 1,200
Kokoro
Vendor: Hexgrad
What it does: a compact model with good quality for its size, suited to narrating a large volume of short videos at low cost.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 600
≈ 2,000
≈ 6,000
Piper
Vendor: Rhasspy
What it does: a very small model producing clear, plain narration on hardware with no GPU. The right answer for high-volume internal and accessibility material.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 900
≈ 3,000
≈ 9,000
Llama 3.3 70B
Vendor: Meta
What it does: rewrites the script so each section fits the seconds available, without losing meaning. This is usually the difference between narration that fits the picture and narration that has to be re-recorded.
Requirement
Minimum
Medium
High
GPU type
2× RTX 4090 (reduced precision)
H100 80 GB
4× H100 80 GB
VRAM
48 GB combined
80 GB
320 GB combined
vCPUs
16
24
64
RAM
64 GB
128 GB
512 GB
Server
2× RTX 4090 24 GB
1× H100 SXM 80 GB
4× H100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 600
≈ 2,400
≈ 9,000
AudioSeal
Vendor: Meta
What it does: marks the generated narration as synthetic in a way that survives re-encoding, which several jurisdictions now expect of published material.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes narrated per hour)
≈ 3,000
≈ 10,000
≈ 26,000
FFmpeg
Vendor: FFmpeg project
What it does: mixes the narration with the existing soundtrack and writes the delivery file. No GPU needed.
Requirement
Minimum
Medium
High
GPU type
No GPU required
No GPU required
GPU-accelerated decode (NVENC/NVDEC)
VRAM
—
—
8 GB
vCPUs
2
8
16
RAM
4 GB
16 GB
32 GB
Server
CPU instance, 2 vCPU
CPU instance, 8 vCPU
1× RTX 4090 24 GB
Rate (video minutes narrated per hour)
≈ 6,000
≈ 20,000
≈ 60,000
Choosing between them
Model choice depends on how natural the voice must sound, how many languages you need, and how tightly the narration must fit the picture. Our consultants review your scripts and videos, then recommend a voice model, the script-fitting step and the consent and marking process.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Narrator to video - generate narration from text. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.