Skip to main content

Model reference — AI Medical · Ophthalmology

AI models for Retinal image analysis

Published accuracy is strong — AUROC in the low 0.9s for referable diabetic retinopathy — and drops predictably with camera and population shift. That gap, not the headline number, is what a deployment has to plan for.

Three kinds of model apply. Foundation encoders produce embeddings for your own classifiers and transfer best across cameras. Task-specific graders return a severity grade directly. Quantification tools return vessel and disc measurements rather than a label. OCT needs 3D-capable models; 2D classifiers trained on one public set transfer poorly.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Retinal image analysis service AI Medical services Pricing

Input type — Fundus and OCT images

Models in this group take a colour fundus photograph or an OCT scan; 3D models take a full OCT volume. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

RETFound / RETFound-DINOv2 / RETFound-DINOv3

Vendor: Moorfields Eye Hospital / UCL

What it does: is the reference retinal foundation model: fine-tuned it reaches AUROC 0.82–0.94 on sight-threatening disease, and it also predicts incident heart failure and myocardial infarction above baseline — the clearest evidence for oculomics in an open model.

AUROC 0.822–0.943 for sight-threatening eye disease; diabetic retinopathy AUROC 0.943 on APTOS-2019; also predicts incident heart failure and MI above baselines (independent, Nature 2023).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

RetFiner-RETFound

Vendor: RetFiner authors

What it does: refines the RETFound representation, improving downstream accuracy without retraining from scratch.

Task-dependent; reported improvements are benchmark-specific.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

RetFiner-UrFound

Vendor: RetFiner authors

What it does: applies the same refinement to the UrFound base for retinal downstream tasks.

Task-dependent; reported improvements are benchmark-specific.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

RetFiner-VisionFM

Vendor: RetFiner authors

What it does: applies the same refinement to VisionFM, for teams already standardised on it.

Task-dependent; reported improvements are benchmark-specific.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

VisionFM

Vendor: Zhejiang University / Shanghai AI Lab, CUHK collaborators

What it does: covers eight ophthalmic modalities — fundus, OCT, slit-lamp, ultrasound, angiography and more — and matched or exceeded junior ophthalmologists on several diagnostic tasks.

Matches or exceeds junior ophthalmologists on several diagnostic tasks across 8 modalities (independent, NEJM AI 2024).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (images/hour)≈ 1,200≈ 4,200≈ 10,800

OCTCube / OCTCube-M

Vendor: University of Washington and community

What it does: reads the OCT volume rather than single B-scans, which is worth 0.02–0.06 AUROC over 2D baselines because pathology spreads through depth.

Outperforms 2D baselines on 3D OCT disease detection, AUROC +0.02–0.06 (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3090A100 80 GBH100 80 GB
VRAM24 GB80 GB80 GB
vCPUs122448
RAM64 GB128 GB256 GB
Server1× RTX 3090 24 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (images/hour)≈ 2,000≈ 7,000≈ 18,000

AutoMorph

Vendor: Moorfields / UCL

What it does: measures the vessels rather than labelling the disease — calibre, tortuosity, fractal dimension and disc metrics — which is what oculomics research runs on.

Vessel segmentation Dice about 0.80–0.83; artery/vein AUC about 0.93 (independent, TVST 2022).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

Diabetic retinopathy grading CNNs

Vendor: Community, EyePACS-trained (EfficientNet / Inception)

What it does: grades referable diabetic retinopathy at AUROC 0.94–0.98 on the public sets. Camera and population shift is where it loses accuracy, and where we benchmark first.

Referable DR AUROC 0.94–0.98 on EyePACS and Messidor; degrades with camera and population shift (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

OCT2017 classifiers

Vendor: Community (Kermany dataset lineage)

What it does: classifies four common retinal pathologies on a B-scan. The headline 96–98 percent is on its original test set only; expect a substantial drop elsewhere, so treat it as a starting baseline.

About 96–98 percent on the original test set; heavily over-fitted to that dataset, expect a large drop elsewhere (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090L40S 48 GB
VRAM12 GB24 GB48 GB
vCPUs4816
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (images/hour)≈ 30,000≈ 105,000≈ 270,000

Choosing between them

If you have labelled data and one endpoint, a task-specific grader is quickest. If you have several endpoints or a shifting camera fleet, a foundation encoder is the better base. Note that several of the strongest models here are released under non-commercial terms — we check licensing against your intended use before a pipeline is built.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Retinal image analysis service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for retinal image analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.