Model reference — AI Medical · Surgery and endoscopy
AI models for Endoscopy video analysis
Polyp segmentation reaches Dice 0.82–0.90 on the public sets and falls to 0.70–0.80 on data from another centre. That is the single most important number on this page.
Two approaches: a video foundation model that produces transferable features across polyp detection, segmentation and disease classification, and task-specific segmenters that return a mask directly and run fast enough for live use.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take a colonoscopy or endoscopy clip; the rate assumes short clips rather than whole procedures. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
EndoFM (Endo-FM)
Vendor: CUHK and academic consortium
What it does: produces spatio-temporal features from endoscopy video that transfer across polyp detection, segmentation and disease classification — worth 3–7 points over supervised baselines, and the better base if you are building several endpoints.
Plus 3–7 points over supervised baselines on polyp detection, segmentation and disease classification transfer (independent, MICCAI 2023).
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (clips/hour)
≈ 150
≈ 525
≈ 1,400
PraNet / Polyp-PVT / SAM-based polyp segmenters
Vendor: Academic community (Kvasir-SEG and CVC-ClinicDB lineage)
What it does: outlines polyps frame by frame, fast enough for live assistance. Dice 0.82–0.90 on the public sets falls to 0.70–0.80 across centres, which is the number to plan around.
Dice 0.82–0.90 on Kvasir-SEG; falls to 0.70–0.80 on cross-centre data (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (clips/hour)
≈ 400
≈ 1,400
≈ 3,600
Choosing between them
For live assistance, a fast task-specific segmenter is the right shape. For building several endpoints from one archive, the foundation model transfers better and costs more per clip. Either way we benchmark on your own centre’s footage first, because the cross-centre gap is large.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for endoscopy video analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.