This service turns a recorded meeting, call or podcast into a written record: a summary of what was discussed, the decisions taken, the actions agreed with their owners, and the points left unresolved.
It is a three-stage pipeline, and each stage matters. The audio is transcribed; the transcript is divided by speaker, so that an action can be attributed to the person who accepted it; and a language model then reads the whole conversation and writes the record. The output that people actually use is not a paragraph of prose but a structured list — decisions, actions, owners, dates — with each item linked to the moment in the recording it came from, so a disputed item can be checked in seconds.
Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.
faster-whisper
Vendor: SYSTRAN
What it does: transcribes the recording quickly and accurately, which is the stage everything downstream depends on. Errors here cannot be repaired later.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 20
≈ 70
≈ 180
pyannote.audio 3
Vendor: pyannote (Hervé Bredin)
What it does: divides the transcript by speaker so that statements and actions are attributed to the person who made them rather than floating free.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 20
≈ 70
≈ 180
WhisperX
Vendor: University of Oxford (VGG)
What it does: aligns the transcript to the audio to the word, which is what lets every line in the summary link back to the exact second it came from.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 15
≈ 55
≈ 140
Llama 3.3 70B
Vendor: Meta
What it does: reads the whole conversation and writes the record: summary, decisions, actions with owners, and open questions. The most reliable option where the record is relied upon.
Requirement
Minimum
Medium
High
GPU type
2× RTX 4090 (reduced precision)
H100 80 GB
4× H100 80 GB
VRAM
48 GB combined
80 GB
320 GB combined
vCPUs
16
24
64
RAM
64 GB
128 GB
512 GB
Server
2× RTX 4090 24 GB
1× H100 SXM 80 GB
4× H100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 25 meeting hours/hour
≈ 90 meeting hours/hour
≈ 300 meeting hours/hour
Qwen2.5 32B
Vendor: Alibaba Cloud
What it does: reads very long transcripts in one pass, so a three-hour board meeting is summarised without being split and losing the thread between parts.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (reduced precision)
L40S 48 GB
2× H100 80 GB
VRAM
22 GB
48 GB
160 GB combined
vCPUs
12
16
48
RAM
48 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× L40S 48 GB
2× H100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 30 meeting hours/hour
≈ 100 meeting hours/hour
≈ 330 meeting hours/hour
Mistral Small 3
Vendor: Mistral AI
What it does: a compact model for summarising a high volume of short calls, where cost per call matters more than depth.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
L40S 48 GB
H100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
12
16
32
RAM
48 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× L40S 48 GB
1× H100 SXM 80 GB
Rate (audio hours processed per hour)
≈ 60 meeting hours/hour
≈ 200 meeting hours/hour
≈ 600 meeting hours/hour
Choosing between them
The right pipeline depends on your meeting length and format, and on whether summaries are read for information or relied on as a record. Our consultants review your recordings and the record you need, then recommend the models for each stage and the format the output should take.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Podcast / meeting summarization. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.