Call-center analytics processes every call rather than the one percent a supervisor has time to listen to, and reports what is happening across all of them: why customers are calling, how those calls end, where agents are struggling, and which calls broke a rule.
The pipeline transcribes each call, separates agent from customer, and then applies a set of judgements: the reason for the call, the sentiment on both sides and how it moved during the call, whether required disclosures were read, whether the issue was resolved, and whether the call resembles a rising cluster of complaints. The value is in the aggregate — a fault reported by forty customers in a morning is visible immediately — and in the fairness of measuring every agent on every call rather than on a handful of samples.
Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.
faster-whisper
Vendor: SYSTRAN
What it does: transcribes calls at the volume a contact centre produces, on telephone-quality audio. The foundation of every other measurement.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 20
≈ 70
≈ 180
pyannote.audio 3
Vendor: pyannote (Hervé Bredin)
What it does: separates agent from customer, without which sentiment, talk-time and interruption measures are meaningless.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 20
≈ 70
≈ 180
TitaNet
Vendor: NVIDIA
What it does: confirms which agent handled a call by voice, which matters where several agents share a line or a login.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 80
≈ 300
≈ 800
RoBERTa
Vendor: Meta
What it does: scores customer sentiment through the call, so a call that started badly and ended well is distinguished from one that went the other way.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 40,000 calls/hour
≈ 160,000 calls/hour
≈ 450,000 calls/hour
DeBERTa v3
Vendor: Microsoft
What it does: classifies the reason for the call against your own categories, fitted from calls your team has already coded. This is what produces reliable volume-by-reason reporting.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 30,000 calls/hour
≈ 120,000 calls/hour
≈ 340,000 calls/hour
Llama 3.1 8B
Vendor: Meta
What it does: checks each call against your compliance script — whether disclosures were read, whether the resolution was confirmed — and quotes the passage supporting its judgement.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
H100 80 GB
VRAM
16 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
128 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× H100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 2,000 calls/hour
≈ 6,000 calls/hour
≈ 20,000 calls/hour
BGE-M3
Vendor: Beijing Academy of Artificial Intelligence
What it does: fingerprints calls so that a cluster of customers reporting the same new problem is detected as it emerges rather than in next month’s report.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 6,000 calls/hour
≈ 30,000 calls/hour
≈ 90,000 calls/hour
Choosing between them
What you get out depends on your call types, your compliance requirements and the reporting your supervisors will act on. Our consultants review a sample of your calls and your quality framework, then recommend the models, the judgements worth automating, and the reporting that makes them useful.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Call-center analytics. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.