Text to speech reads written text aloud. Voice cloning does it in a particular person’s voice, built either from a few seconds of their speech or from a longer recording session that produces a higher-quality result.
The practical uses are narration at volume — training material, product announcements, accessibility audio — and consistency, where the same voice must read thousands of items that change weekly and re-recording with a human is impractical. Two constraints govern every deployment. Consent is required: cloning a voice needs the documented permission of its owner, and we build that check into the process. And output should be marked, so that synthetic speech can be identified as such later; audio watermarking makes that possible.
Models in this group take written text and a voice sample as input, and produce spoken audio. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of generated speech from about 150 words of text.
XTTS v2
Vendor: Coqui
What it does: clones a voice from a few seconds of speech and reads text in it across seventeen languages, including reading in a language the original speaker never spoke. The most versatile general choice.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 120
≈ 400
≈ 1,000
F5-TTS
Vendor: Shanghai Jiao Tong University
What it does: produces very natural speech with accurate timing and emphasis from a short voice sample, and is quick enough for on-demand generation.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 200
≈ 700
≈ 1,600
StyleTTS 2
Vendor: Columbia University
What it does: generates speech with controllable style — pace, emphasis, emotion — which matters for narration that must hold a listener’s attention for an hour.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 150
≈ 500
≈ 1,200
Chatterbox
Vendor: Resemble AI
What it does: a recent model with strong expressive control and built-in marking of its output as synthetic, which suits published material.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 130
≈ 450
≈ 1,100
Fish Speech
Vendor: Fish Audio
What it does: multilingual generation with fast cloning from short samples, strong on Asian languages where several other models are weak.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 160
≈ 550
≈ 1,300
Piper
Vendor: Rhasspy
What it does: a very small model producing clear, plain speech on almost any hardware, including devices with no GPU. The right answer for announcements and accessibility audio at scale.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 900
≈ 3,000
≈ 9,000
Kokoro
Vendor: Hexgrad
What it does: a compact model with unusually good quality for its size, which makes it a practical middle option between the small and the heavy models.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 600
≈ 2,000
≈ 6,000
AudioSeal
Vendor: Meta
What it does: marks generated audio so it can be identified as synthetic afterwards, surviving recompression and re-recording. Not a voice model, but part of any responsible deployment.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio minutes generated per hour)
≈ 2,000
≈ 7,000
≈ 18,000
Choosing between them
Model choice follows from how natural the result must sound, how many voices you need, and whether generation happens in advance or on demand. Our consultants review your scripts and voice requirements, then recommend a model, the recording needed to build the voice, and the consent and marking process around it.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Voice cloning / custom voice TTS. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.