Model reference — AI Medical · Respiratory and acoustics
AI models for Health acoustics analysis
The foundation models here are the substantive development: one is trained on 300 million two-second clips, and both report beating task-specific baselines with small labelled sets.
Downstream accuracy is more modest and honestly reported. Adventitious lung sound classification sits around 0.60–0.65 on the standard benchmark, which remains unsolved, and spirometry interpretation agreement is 0.68–0.82 kappa against experts.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take short audio clips — cough, breath, lung sounds or speech; the rate is recorded hours per hour of wall-clock time. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
HeAR (Health Acoustic Representations)
Vendor: Google Research / Google Health
What it does: was trained on 300 million audio clips and turns a cough or breath into an embedding that a classifier reaches useful accuracy on with a small labelled set — which is what makes a TB or respiratory screening project affordable.
Trained on 300M two-second audio clips; leading embeddings for TB and COVID cough screening with small labelled sets (developer-reported).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours/hour)
≈ 90
≈ 315
≈ 810
OPERA
Vendor: University of Cambridge
What it does: beats task-specific baselines across 19 respiratory health tasks from one embedding model, under a permissive licence.
Beats task-specific baselines on 19 respiratory health tasks (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (audio hours/hour)
≈ 90
≈ 315
≈ 810
ICBHI respiratory sound classifiers
Vendor: Academic community (ICBHI 2017 lineage, CNN and AST fine-tunes)
What it does: classifies crackles and wheeze from stethoscope audio. The benchmark score of 0.60–0.65 is the state of the art, and this remains an unsolved problem — we will say so rather than oversell it.
ICBHI score about 0.60–0.65 for crackle and wheeze detection — still an unsolved benchmark (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (audio hours/hour)
≈ 200
≈ 700
≈ 1,800
PFT interpretation models
Vendor: Academic community (gradient boosting and ATS-rule hybrids)
What it does: classifies spirometry into obstructive, restrictive or mixed patterns at 0.68–0.82 kappa against expert interpretation, as the non-audio companion to the models above.
Agreement with expert interpretation 0.68–0.82 kappa (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (audio hours/hour)
≈ 200
≈ 700
≈ 1,800
Choosing between them
Start with an embedding model and your own labels — that is both the cheapest and the best-supported route in this area. Direct classifiers are worth benchmarking only where your recording conditions closely match the training set. We measure on your own recordings before sizing anything.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for health acoustics analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.