Skip to main content

Model reference — Audio

AI models for Voice cloning / custom voice TTS

Text to speech reads written text aloud. Voice cloning does it in a particular person’s voice, built either from a few seconds of their speech or from a longer recording session that produces a higher-quality result.

The practical uses are narration at volume — training material, product announcements, accessibility audio — and consistency, where the same voice must read thousands of items that change weekly and re-recording with a human is impractical. Two constraints govern every deployment. Consent is required: cloning a voice needs the documented permission of its owner, and we build that check into the process. And output should be marked, so that synthetic speech can be identified as such later; audio watermarking makes that possible.

Voice cloning / custom voice TTS service AI models for audio files

Input type — Audio

Models in this group take written text and a voice sample as input, and produce spoken audio. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of generated speech from about 150 words of text.

XTTS v2

Vendor: Coqui

What it does: clones a voice from a few seconds of speech and reads text in it across seventeen languages, including reading in a language the original speaker never spoke. The most versatile general choice.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 120≈ 400≈ 1,000

F5-TTS

Vendor: Shanghai Jiao Tong University

What it does: produces very natural speech with accurate timing and emphasis from a short voice sample, and is quick enough for on-demand generation.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 200≈ 700≈ 1,600

StyleTTS 2

Vendor: Columbia University

What it does: generates speech with controllable style — pace, emphasis, emotion — which matters for narration that must hold a listener’s attention for an hour.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 150≈ 500≈ 1,200

Chatterbox

Vendor: Resemble AI

What it does: a recent model with strong expressive control and built-in marking of its output as synthetic, which suits published material.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 130≈ 450≈ 1,100

Fish Speech

Vendor: Fish Audio

What it does: multilingual generation with fast cloning from short samples, strong on Asian languages where several other models are weak.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 160≈ 550≈ 1,300

Piper

Vendor: Rhasspy

What it does: a very small model producing clear, plain speech on almost any hardware, including devices with no GPU. The right answer for announcements and accessibility audio at scale.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 900≈ 3,000≈ 9,000

Kokoro

Vendor: Hexgrad

What it does: a compact model with unusually good quality for its size, which makes it a practical middle option between the small and the heavy models.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 600≈ 2,000≈ 6,000

AudioSeal

Vendor: Meta

What it does: marks generated audio so it can be identified as synthetic afterwards, surviving recompression and re-recording. Not a voice model, but part of any responsible deployment.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio minutes generated per hour)≈ 2,000≈ 7,000≈ 18,000

Choosing between them

Model choice follows from how natural the result must sound, how many voices you need, and whether generation happens in advance or on demand. Our consultants review your scripts and voice requirements, then recommend a model, the recording needed to build the voice, and the consent and marking process around it.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Voice cloning / custom voice TTS service AI models for audio files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Voice cloning / custom voice TTS. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.