This service removes background noise from a live audio feed as it passes — traffic, machinery, ventilation, keyboards, room echo, crowd noise — so that listeners hear the voice clearly and transcription works.
Live suppression is judged on two numbers: how much noise it removes and how much delay it adds. A model that cleans beautifully but adds half a second is useless in a conversation, so live work uses the smallest models that do the job, and the good ones add only a few milliseconds. There is also a trade-off to respect: aggressive suppression removes parts of the speech along with the noise, which can make a transcript worse even as the audio sounds cleaner. Which setting is right depends on whether a person or a machine is listening.
Models in this group take a live camera or stream as input and must keep pace with it in real time. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is how many camera streams or feeds one server of that tier can keep up with in real time, not a per-hour count: live work must fit inside the interval between frames, and a server that cannot keep pace drops frames rather than falling behind. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one live audio feed cleaned continuously at 16 kHz.
DeepFilterNet 3
Vendor: Friedrich-Alexander-Universität
What it does: removes steady noise in real time with a few milliseconds of delay and very little computing power. The live default, and enough for most situations.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (streams handled continuously)
≈ 20 streams
≈ 70 streams
≈ 180 streams
RNNoise
Vendor: Xiph.Org Foundation
What it does: a classical suppressor needing no GPU, effective on hum, fans and air handling, and cheap enough to run on every stream regardless.
Requirement
Minimum
Medium
High
GPU type
No GPU required
No GPU required
No GPU required
VRAM
—
—
—
vCPUs
2
4
16
RAM
4 GB
8 GB
32 GB
Server
CPU instance, 2 vCPU
CPU instance, 4 vCPU
CPU instance, 16 vCPU
Rate (streams handled continuously)
≈ 100 streams
≈ 300 streams
≈ 900 streams
MetricGAN+ (SpeechBrain)
Vendor: SpeechBrain
What it does: optimises directly for perceived quality, which suits a broadcast or conference feed where a person is listening and a little more delay is acceptable.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (streams handled continuously)
≈ 6 streams
≈ 22 streams
≈ 60 streams
MDX-Net
Vendor: Kuielab
What it does: separates the voice from music or another voice rather than merely reducing noise, which is the only thing that works when the interference is itself speech or music.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (streams handled continuously)
≈ 2 streams
≈ 8 streams
≈ 20 streams
Resemble Enhance
Vendor: Resemble AI
What it does: rebuilds badly degraded live audio — a poor telephone line, a failing microphone — at a delay too high for conversation but acceptable for monitoring.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (streams handled continuously)
≈ 1 stream
≈ 4 streams
≈ 10 streams
Silero VAD
Vendor: Silero
What it does: detects speech so suppression and everything after it run only when there is something to clean.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
4 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
8 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (streams handled continuously)
≈ 60 streams
≈ 200 streams
≈ 600 streams
Choosing between them
The right model and setting depend on your noise, your delay budget and whether the audio is for a listener or a transcription model. Our consultants test the candidates on your own feeds, measuring both listening quality and transcription accuracy, and recommend the configuration that improves your outcome.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Audio noise suppression. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.