Detection and segmentation are different problems here. Detection — is there blood — reaches AUC above 0.95 on challenge data. Segmentation, which is what gives you a volume, sits around Dice 0.72, and that gap is the honest state of the open art.
No self-hostable open model matches the cleared commercial triage products for this indication. Where a cleared device is a requirement, those vendors do supply on-premise appliances; where the requirement is measurement and cohort work, these models are the practical option.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take a non-contrast head CT study. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
DeepBleed
Vendor: Yale and community
What it does: segments intracerebral haemorrhage and reports its volume. Dice 0.72 is the honest ceiling of the open art here, and it is a measurement tool rather than a triage device.
Dice about 0.72 for intracerebral haemorrhage volume segmentation (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3090
A100 80 GB
H100 80 GB
VRAM
24 GB
80 GB
80 GB
vCPUs
12
24
48
RAM
64 GB
128 GB
256 GB
Server
1× RTX 3090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (volumes/hour)
≈ 60
≈ 210
≈ 540
RSNA-ICH ensemble models
Vendor: RSNA challenge community (EfficientNet / SE-ResNeXt)
What it does: detects haemorrhage and its subtype per slice at AUC above 0.95 on challenge data — accurate enough to order a reading queue, which is the highest-value use of it.
Weighted log-loss challenge winners; AUC above 0.95 for any-haemorrhage detection on challenge data (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (volumes/hour)
≈ 200
≈ 700
≈ 1,800
Choosing between them
If the goal is ordering a queue, the detection ensembles are accurate enough and cheap to run. If the goal is a reported volume, the segmentation model is the one to use, with its accuracy limit understood. Both are sized here for 3D head CT volumes.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for stroke and haemorrhage detection. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.