Transcription turns recorded speech into written text with timings, so that calls, meetings, interviews and broadcasts become searchable records instead of files nobody can look inside.
Accuracy is quoted as word error rate — the share of words wrong — and it depends far more on your audio than on the model: a clean single-speaker recording transcribes almost perfectly, while a telephone call with two people talking over each other and background noise is a much harder problem. The practical decisions are which model, whether to clean the audio first, and whether transcription runs after the fact in bulk or live as the words are spoken, which is a different engineering problem with a different cost.
Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.
Whisper large-v3
Vendor: OpenAI
What it does: transcribes speech in roughly a hundred languages and is remarkably tolerant of accents, background noise and poor recording quality. The general-purpose default, and the benchmark others are measured against.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 4
≈ 12
≈ 30
Whisper large-v3-turbo
Vendor: OpenAI
What it does: a faster version of the same model with a small accuracy penalty, which is usually the better economic choice for large archives where a fraction of a percent of accuracy is not worth double the hardware.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 12
≈ 40
≈ 100
faster-whisper
Vendor: SYSTRAN
What it does: the same model re-engineered to run several times faster on the same card, with identical output. Where Whisper is the right model, this is usually the right way to run it.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 20
≈ 70
≈ 180
Parakeet TDT
Vendor: NVIDIA
What it does: an English transcription model built for speed, capable of transcribing an hour of audio in seconds on a single professional card. The choice for very large English archives and for live captioning.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 60
≈ 250
≈ 600
Canary
Vendor: NVIDIA
What it does: transcribes and translates across major European languages in one model, returning either the original words or an English version of them.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 15
≈ 50
≈ 120
wav2vec 2.0
Vendor: Meta
What it does: a small model that can be fitted to your own domain vocabulary — drug names, part numbers, place names — which is what fixes the specific words a general model keeps getting wrong.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 25
≈ 90
≈ 220
MMS (Massively Multilingual Speech)
Vendor: Meta
What it does: covers over a thousand languages, including many with no other transcription option. The choice when the language matters more than the accuracy achievable in it.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 10
≈ 35
≈ 90
Choosing between them
Model choice follows from your audio quality, your languages and whether transcription must be immediate or can run overnight. Our consultants measure word error rate on your own recordings, then recommend a model, a cleaning step if it pays for itself, and the throughput needed for your volume.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Audio to text - transcription. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.