Skip to main content

Model reference — Audio

AI models for Podcast / meeting summarization

This service turns a recorded meeting, call or podcast into a written record: a summary of what was discussed, the decisions taken, the actions agreed with their owners, and the points left unresolved.

It is a three-stage pipeline, and each stage matters. The audio is transcribed; the transcript is divided by speaker, so that an action can be attributed to the person who accepted it; and a language model then reads the whole conversation and writes the record. The output that people actually use is not a paragraph of prose but a structured list — decisions, actions, owners, dates — with each item linked to the moment in the recording it came from, so a disputed item can be checked in seconds.

Podcast / meeting summarization service AI models for audio files

Input type — Audio

Models in this group take an audio file as input: a recording, a call, an interview or a broadcast. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one hour of 16 kHz mono speech.

faster-whisper

Vendor: SYSTRAN

What it does: transcribes the recording quickly and accurately, which is the stage everything downstream depends on. Errors here cannot be repaired later.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 20≈ 70≈ 180

pyannote.audio 3

Vendor: pyannote (Hervé Bredin)

What it does: divides the transcript by speaker so that statements and actions are attributed to the person who made them rather than floating free.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 20≈ 70≈ 180

WhisperX

Vendor: University of Oxford (VGG)

What it does: aligns the transcript to the audio to the word, which is what lets every line in the summary link back to the exact second it came from.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (audio hours processed per hour)≈ 15≈ 55≈ 140

Llama 3.3 70B

Vendor: Meta

What it does: reads the whole conversation and writes the record: summary, decisions, actions with owners, and open questions. The most reliable option where the record is relied upon.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (audio hours processed per hour)≈ 25 meeting hours/hour≈ 90 meeting hours/hour≈ 300 meeting hours/hour

Qwen2.5 32B

Vendor: Alibaba Cloud

What it does: reads very long transcripts in one pass, so a three-hour board meeting is summarised without being split and losing the thread between parts.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (reduced precision)L40S 48 GB2× H100 80 GB
VRAM22 GB48 GB160 GB combined
vCPUs121648
RAM48 GB64 GB256 GB
Server1× RTX 4090 24 GB1× L40S 48 GB2× H100 SXM 80 GB
Rate (audio hours processed per hour)≈ 30 meeting hours/hour≈ 100 meeting hours/hour≈ 330 meeting hours/hour

Mistral Small 3

Vendor: Mistral AI

What it does: a compact model for summarising a high volume of short calls, where cost per call matters more than depth.

RequirementMinimumMediumHigh
GPU typeRTX 4090L40S 48 GBH100 80 GB
VRAM24 GB48 GB80 GB
vCPUs121632
RAM48 GB64 GB128 GB
Server1× RTX 4090 24 GB1× L40S 48 GB1× H100 SXM 80 GB
Rate (audio hours processed per hour)≈ 60 meeting hours/hour≈ 200 meeting hours/hour≈ 600 meeting hours/hour

Choosing between them

The right pipeline depends on your meeting length and format, and on whether summaries are read for information or relied on as a record. Our consultants review your recordings and the record you need, then recommend the models for each stage and the format the output should take.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Podcast / meeting summarization service AI models for audio files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Podcast / meeting summarization. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.