Skip to main content

Model reference — Video

AI models for Narrator to video - generate narration from text

This service writes narration onto a video from a script: the text is read aloud by a synthetic voice, timed to the picture, and mixed with the existing sound — which is how training material, product videos and multilingual versions get made without booking a studio.

The distinctive constraint is length. A shot lasts a fixed number of seconds, and the words must fit it; a script that runs long has to be shortened or spoken faster, and only one of those is acceptable. So a working pipeline pairs a speech model with a language model that adjusts the script to fit the available time. Two rules apply to every deployment: a cloned voice needs the documented consent of the person it belongs to, and generated audio should be watermarked so it can be identified as synthetic later.

Narrator to video - generate narration from text service AI models for video files

Input type — Video

Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of video narrated from about 150 words of script.

XTTS v2

Vendor: Coqui

What it does: clones a voice from a few seconds of sample and reads the script in it across seventeen languages, which is how one narrator’s voice carries a whole multilingual library.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 120≈ 400≈ 1,000

F5-TTS

Vendor: Shanghai Jiao Tong University

What it does: produces very natural narration with accurate emphasis and pacing, and is fast enough to regenerate a section on demand when the script changes.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 200≈ 700≈ 1,600

StyleTTS 2

Vendor: Columbia University

What it does: gives direct control over pace and emphasis, which is what allows narration to be stretched or compressed to fit a shot without sounding rushed.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 150≈ 500≈ 1,200

Kokoro

Vendor: Hexgrad

What it does: a compact model with good quality for its size, suited to narrating a large volume of short videos at low cost.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 600≈ 2,000≈ 6,000

Piper

Vendor: Rhasspy

What it does: a very small model producing clear, plain narration on hardware with no GPU. The right answer for high-volume internal and accessibility material.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 900≈ 3,000≈ 9,000

Llama 3.3 70B

Vendor: Meta

What it does: rewrites the script so each section fits the seconds available, without losing meaning. This is usually the difference between narration that fits the picture and narration that has to be re-recorded.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 600≈ 2,400≈ 9,000

AudioSeal

Vendor: Meta

What it does: marks the generated narration as synthetic in a way that survives re-encoding, which several jurisdictions now expect of published material.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes narrated per hour)≈ 3,000≈ 10,000≈ 26,000

FFmpeg

Vendor: FFmpeg project

What it does: mixes the narration with the existing soundtrack and writes the delivery file. No GPU needed.

RequirementMinimumMediumHigh
GPU typeNo GPU requiredNo GPU requiredGPU-accelerated decode (NVENC/NVDEC)
VRAM8 GB
vCPUs2816
RAM4 GB16 GB32 GB
ServerCPU instance, 2 vCPUCPU instance, 8 vCPU1× RTX 4090 24 GB
Rate (video minutes narrated per hour)≈ 6,000≈ 20,000≈ 60,000

Choosing between them

Model choice depends on how natural the voice must sound, how many languages you need, and how tightly the narration must fit the picture. Our consultants review your scripts and videos, then recommend a voice model, the script-fitting step and the consent and marking process.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Narrator to video - generate narration from text service AI models for video files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Narrator to video - generate narration from text. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.