Skip to main content

Model reference — Audio

AI models for Audio to text - transcription

Transcription turns recorded speech into written text with timings, so that calls, meetings, interviews and broadcasts become searchable records instead of files nobody can look inside.

Accuracy is quoted as word error rate — the share of words wrong — and it depends far more on your audio than on the model: a clean single-speaker recording transcribes almost perfectly, while a telephone call with two people talking over each other and background noise is a much harder problem. The practical decisions are which model, whether to clean the audio first, and whether transcription runs after the fact in bulk or live as the words are spoken, which is a different engineering problem with a different cost.

Audio to text - transcription service AI models for audio files

Input type — Audio

Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.

Whisper large-v3

Vendor: OpenAI

What it does: transcribes speech in roughly a hundred languages and is remarkably tolerant of accents, background noise and poor recording quality. The general-purpose default, and the benchmark others are measured against.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 4≈ 12≈ 30

Whisper large-v3-turbo

Vendor: OpenAI

What it does: a faster version of the same model with a small accuracy penalty, which is usually the better economic choice for large archives where a fraction of a percent of accuracy is not worth double the hardware.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 12≈ 40≈ 100

faster-whisper

Vendor: SYSTRAN

What it does: the same model re-engineered to run several times faster on the same card, with identical output. Where Whisper is the right model, this is usually the right way to run it.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 20≈ 70≈ 180

Parakeet TDT

Vendor: NVIDIA

What it does: an English transcription model built for speed, capable of transcribing an hour of audio in seconds on a single professional card. The choice for very large English archives and for live captioning.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 60≈ 250≈ 600

Canary

Vendor: NVIDIA

What it does: transcribes and translates across major European languages in one model, returning either the original words or an English version of them.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 15≈ 50≈ 120

wav2vec 2.0

Vendor: Meta

What it does: a small model that can be fitted to your own domain vocabulary — drug names, part numbers, place names — which is what fixes the specific words a general model keeps getting wrong.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 25≈ 90≈ 220

MMS (Massively Multilingual Speech)

Vendor: Meta

What it does: covers over a thousand languages, including many with no other transcription option. The choice when the language matters more than the accuracy achievable in it.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 10≈ 35≈ 90

Choosing between them

Model choice follows from your audio quality, your languages and whether transcription must be immediate or can run overnight. Our consultants measure word error rate on your own recordings, then recommend a model, a cleaning step if it pays for itself, and the throughput needed for your volume.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Audio to text - transcription service AI models for audio files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Audio to text - transcription. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.