Model reference — AI Medical · Surgery and endoscopy
AI models for Surgical workflow analysis
Phase recognition is the reliable part: 88–92 percent accuracy on the standard laparoscopic cholecystectomy benchmark. Action-triplet recognition — instrument, verb, target together — is much harder, with mAP around 0.30–0.40, and that is the honest state of the art.
Instrument segmentation sits in between, at IoU 0.70–0.85 on challenge data. Real-time intra-operative guidance is a documented gap in the open ecosystem; we will say so rather than size a pipeline for it.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take laparoscopic or surgical video, as clips or whole recordings. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
TeCNO / Rendezvous / SurgVLP
Vendor: University of Strasbourg (CAMMA)
What it does: labels the phase of an operation at 88–92 percent accuracy, which makes phase-duration analytics dependable today. Action-triplet recognition from the same family is research-grade, at mAP 0.30–0.40.
Phase recognition accuracy about 88–92 percent on Cholec80; triplet recognition mAP about 0.30–0.40 on CholecT50 (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (clips/hour)
≈ 150
≈ 525
≈ 1,400
EndoVis instrument segmentation models
Vendor: Academic community (EndoVis challenges)
What it does: segments instruments and anatomy per frame at IoU 0.70–0.85, for technique and safety research rather than live guidance.
IoU 0.70–0.85 for instrument segmentation on challenge data (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (clips/hour)
≈ 400
≈ 1,400
≈ 3,600
Choosing between them
If you want phase and duration analytics, the workflow models are ready for that today. If you want fine-grained action recognition, expect research-grade accuracy and plan a human in the loop. Segmentation models are a separate step and are sized for frame-rate processing.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for surgical workflow analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.