Skip to main content

Model reference — AI Medical · Dermatology

AI models for Skin lesion analysis

The strongest reported results in this area are notable: reader studies where the model beats clinicians on early melanoma detection, and about an eleven-point accuracy improvement for non-specialists. Both come with the usual caveat that performance falls on darker skin tones and out-of-distribution cameras.

Three families apply. Multi-disease models classify directly across a wide disease list. Embedding models support your own classifier heads with small label budgets. Vision-language and concept models add text-driven assessment and attribute-level explanation, and the conversational ones are explicitly not validated for triage.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Skin lesion analysis service AI Medical services Pricing

Input type — Skin photographs

Models in this group take a clinical, dermoscopic or total-body photograph. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

PanDerm

Vendor: Monash University / University of Oxford consortium

What it does: classifies across a wide disease list with risk stratification, and beat clinicians in reader studies on early melanoma detection while lifting non-specialist accuracy by about eleven points — the strongest case on this page for putting a model in a triage queue.

Beats clinicians in reader studies on early melanoma detection and improves non-specialist accuracy by about 11 percent (independent, Nature Medicine 2025).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

PanDerm_Base

Vendor: PanDerm authors

What it does: is the embedding-only version: a linear probe on your own labelled set turns it into an endpoint in hours rather than weeks.

Task-dependent; downstream fine-tuning or linear probing required for specific tasks.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

DermLIP_PanDerm

Vendor: PanDerm authors

What it does: pairs skin images with text, so lesions can be retrieved and labelled by description instead of by a fixed class list.

Task-dependent; evaluated on multiple dermatology classification and retrieval tasks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

Derm Foundation

Vendor: Google Health (Health AI Developer Foundations)

What it does: produces a 6144-dimension embedding covering 419 skin conditions, reaching clinician-comparable top-3 accuracy with small label budgets. Superseded by MedSigLIP and listed for continuity.

Supports 419 skin conditions; linear probes reach clinician-comparable top-3 accuracy with small label budgets (developer-reported). Legacy model — MedSigLIP is the current recommendation.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

MONET

Vendor: Stanford University

What it does: scores dermatological concepts — asymmetry, border, pigment network — without concept-level labels, which is what makes a classifier output reviewable and a dataset auditable.

Concept annotation AUC about 0.87–0.94 for dermatological attributes without concept-level labels (independent, Nature Medicine).

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 4,000≈ 14,000≈ 36,000

ISIC-trained EfficientNet / ConvNeXt ensembles

Vendor: ISIC challenge community

What it does: grades melanoma at AUROC 0.93–0.96 on ISIC data. Accuracy drops markedly on darker skin tones and unfamiliar cameras, which is the first thing we measure rather than the last.

Melanoma AUROC 0.93–0.96 on ISIC held-out data; drops markedly on darker skin tones and out-of-distribution cameras (independent).

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (images/hour)≈ 12,000≈ 42,000≈ 108,000

SkinGPT-4

Vendor: KAUST and community

What it does: discusses a photograph alongside a patient description and returns a differential in plain language. Correct in 60–79 percent of internal cases and explicitly not validated for triage.

Correct diagnosis category in about 60–79 percent of internal cases; not validated for triage (developer-reported).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (images/hour)≈ 1,200≈ 4,200≈ 10,800

Choosing between them

Skin tone and camera coverage in your own data matter more than the headline metric — that is the first thing we measure. For a single endpoint an embedding model plus a linear probe is usually both cheapest and most robust; for broad disease coverage the multi-disease models are the base. Several models here are non-commercial; we check licensing before building.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Skin lesion analysis service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for skin lesion analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.