Skip to main content

Model reference — Live feed

AI models for Audio noise suppression

This service removes background noise from a live audio feed as it passes — traffic, machinery, ventilation, keyboards, room echo, crowd noise — so that listeners hear the voice clearly and transcription works.

Live suppression is judged on two numbers: how much noise it removes and how much delay it adds. A model that cleans beautifully but adds half a second is useless in a conversation, so live work uses the smallest models that do the job, and the good ones add only a few milliseconds. There is also a trade-off to respect: aggressive suppression removes parts of the speech along with the noise, which can make a transcript worse even as the audio sounds cleaner. Which setting is right depends on whether a person or a machine is listening.

Audio noise suppression service AI models for live feed

Input type — Live feed

Models in this group take a live camera or stream as input and must keep pace with it in real time. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is how many camera streams or feeds one server of that tier can keep up with in real time, not a per-hour count: live work must fit inside the interval between frames, and a server that cannot keep pace drops frames rather than falling behind. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one live audio feed cleaned continuously at 16 kHz.

DeepFilterNet 3

Vendor: Friedrich-Alexander-Universität

What it does: removes steady noise in real time with a few milliseconds of delay and very little computing power. The live default, and enough for most situations.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (streams handled continuously)≈ 20 streams≈ 70 streams≈ 180 streams

RNNoise

Vendor: Xiph.Org Foundation

What it does: a classical suppressor needing no GPU, effective on hum, fans and air handling, and cheap enough to run on every stream regardless.

RequirementMinimumMediumHigh
GPU typeNo GPU requiredNo GPU requiredNo GPU required
VRAM
vCPUs2416
RAM4 GB8 GB32 GB
ServerCPU instance, 2 vCPUCPU instance, 4 vCPUCPU instance, 16 vCPU
Rate (streams handled continuously)≈ 100 streams≈ 300 streams≈ 900 streams

MetricGAN+ (SpeechBrain)

Vendor: SpeechBrain

What it does: optimises directly for perceived quality, which suits a broadcast or conference feed where a person is listening and a little more delay is acceptable.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (streams handled continuously)≈ 6 streams≈ 22 streams≈ 60 streams

MDX-Net

Vendor: Kuielab

What it does: separates the voice from music or another voice rather than merely reducing noise, which is the only thing that works when the interference is itself speech or music.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (streams handled continuously)≈ 2 streams≈ 8 streams≈ 20 streams

Resemble Enhance

Vendor: Resemble AI

What it does: rebuilds badly degraded live audio — a poor telephone line, a failing microphone — at a delay too high for conversation but acceptable for monitoring.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (streams handled continuously)≈ 1 stream≈ 4 streams≈ 10 streams

Silero VAD

Vendor: Silero

What it does: detects speech so suppression and everything after it run only when there is something to clean.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM4 GB24 GB80 GB
vCPUs4824
RAM8 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (streams handled continuously)≈ 60 streams≈ 200 streams≈ 600 streams

Choosing between them

The right model and setting depend on your noise, your delay budget and whether the audio is for a listener or a transcription model. Our consultants test the candidates on your own feeds, measuring both listening quality and transcription accuracy, and recommend the configuration that improves your outcome.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Audio noise suppression service AI models for live feed Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Audio noise suppression. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.