Skip to main content

Model reference — Audio

AI models for Call-center analytics

Call-center analytics processes every call rather than the one percent a supervisor has time to listen to, and reports what is happening across all of them: why customers are calling, how those calls end, where agents are struggling, and which calls broke a rule.

The pipeline transcribes each call, separates agent from customer, and then applies a set of judgements: the reason for the call, the sentiment on both sides and how it moved during the call, whether required disclosures were read, whether the issue was resolved, and whether the call resembles a rising cluster of complaints. The value is in the aggregate — a fault reported by forty customers in a morning is visible immediately — and in the fairness of measuring every agent on every call rather than on a handful of samples.

Call-center analytics service AI models for audio files

Input type — Audio

Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.

faster-whisper

Vendor: SYSTRAN

What it does: transcribes calls at the volume a contact centre produces, on telephone-quality audio. The foundation of every other measurement.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 20≈ 70≈ 180

pyannote.audio 3

Vendor: pyannote (Hervé Bredin)

What it does: separates agent from customer, without which sentiment, talk-time and interruption measures are meaningless.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 20≈ 70≈ 180

TitaNet

Vendor: NVIDIA

What it does: confirms which agent handled a call by voice, which matters where several agents share a line or a login.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 80≈ 300≈ 800

RoBERTa

Vendor: Meta

What it does: scores customer sentiment through the call, so a call that started badly and ended well is distinguished from one that went the other way.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 40,000 calls/hour≈ 160,000 calls/hour≈ 450,000 calls/hour

DeBERTa v3

Vendor: Microsoft

What it does: classifies the reason for the call against your own categories, fitted from calls your team has already coded. This is what produces reliable volume-by-reason reporting.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 30,000 calls/hour≈ 120,000 calls/hour≈ 340,000 calls/hour

Llama 3.1 8B

Vendor: Meta

What it does: checks each call against your compliance script — whether disclosures were read, whether the resolution was confirmed — and quotes the passage supporting its judgement.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090H100 80 GB
VRAM16 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB128 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× H100 SXM 80 GB
Rate (audio hours processed per hour)≈ 2,000 calls/hour≈ 6,000 calls/hour≈ 20,000 calls/hour

BGE-M3

Vendor: Beijing Academy of Artificial Intelligence

What it does: fingerprints calls so that a cluster of customers reporting the same new problem is detected as it emerges rather than in next month’s report.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 6,000 calls/hour≈ 30,000 calls/hour≈ 90,000 calls/hour

Choosing between them

What you get out depends on your call types, your compliance requirements and the reporting your supervisors will act on. Our consultants review a sample of your calls and your quality framework, then recommend the models, the judgements worth automating, and the reporting that makes them useful.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Call-center analytics service AI models for audio files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Call-center analytics. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.