Skip to main content

Model reference — AI Medical · Drug discovery

AI models for Protein structure prediction

The accuracy reference remains the complex-prediction models, with large gains over the previous generation for protein-ligand, protein-nucleic acid and antibody-antigen complexes. Language-model folding trades some accuracy — TM-score around 0.68 average — for an order-of-magnitude speed-up, which is what makes proteome-scale work affordable.

Design is the newer half: backbone generation against a motif, and sequence design for a given backbone. Both are generative, so success is reported as design-task outcome rather than as accuracy.

Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.

Protein structure prediction service AI Medical services Pricing

Input type — Protein sequence and structure

Models in this group take protein sequence, structure or design constraints. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.

AlphaFold 3

Vendor: Google DeepMind

What it does: predicts complexes — protein with ligand, nucleic acid or antibody — with large gains over the previous generation. Code is open, weights require an application and are non-commercial, which we resolve before a pipeline is built.

Large accuracy gains over AlphaFold 2 for protein-ligand, protein-nucleic acid and antibody-antigen complexes (independent, Nature 2024). Code open; weights require an application, non-commercial.

RequirementMinimumMediumHigh
GPU typeA100 80 GBA100 80 GB ×2H100 80 GB ×4
VRAM80 GB160 GB320 GB
vCPUs244896
RAM128 GB256 GB512 GB
Server1× A100 SXM 80 GB2× A100 SXM 80 GB4× H100 SXM 80 GB
Rate (sequences/hour)≈ 300≈ 1,100≈ 2,700

AlphaFold 2

Vendor: Google DeepMind

What it does: remains the accuracy reference for single-chain structure at near-experimental quality, and its licensing is far simpler than AlphaFold 3.

CASP14 near-experimental structure accuracy; accuracy is target-dependent rather than a single percentage.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (sequences/hour)≈ 800≈ 2,800≈ 7,200

OpenFold

Vendor: OpenFold Consortium

What it does: reproduces AlphaFold 2 performance closely under a licence you can actually deploy, which is usually the deciding factor.

Reported to closely reproduce AlphaFold 2 performance; structure accuracy is target-dependent.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (sequences/hour)≈ 800≈ 2,800≈ 7,200

ESMFold

Vendor: Meta AI

What it does: folds from sequence alone at an order of magnitude less compute than AlphaFold 2, trading some accuracy — which is what makes proteome-scale prediction affordable.

Structure prediction TM-score about 0.68 average, an order of magnitude faster than AlphaFold 2 (independent).

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (sequences/hour)≈ 800≈ 2,800≈ 7,200

ESM-2 (8M–15B) / ESM C

Vendor: Meta AI

What it does: turns a protein sequence into embeddings and contact maps, from 8M to 15B parameters. The base layer under most function-prediction work.

Task-dependent across protein prediction tasks; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (sequences/hour)≈ 2,500≈ 8,800≈ 22,500

ESM3-open-small (1.4B)

Vendor: EvolutionaryScale

What it does: generates as well as understands, taking sequence, structure and function as context — for design work rather than prediction. Non-commercial community licence.

Task-dependent across protein understanding and generation benchmarks. Non-commercial community licence.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (sequences/hour)≈ 800≈ 2,800≈ 7,200

ProtT5

Vendor: Rostlab

What it does: produces protein embeddings for function and structure-related prediction, a well-established and permissively licensed base.

Task-dependent across protein-function and structure-related benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (sequences/hour)≈ 2,500≈ 8,800≈ 22,500

Ankh

Vendor: Ankh authors

What it does: is a protein encoder tuned for efficiency, delivering competitive downstream accuracy at a smaller footprint.

Task-dependent across protein benchmarks.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (sequences/hour)≈ 2,500≈ 8,800≈ 22,500

ProteinMPNN

Vendor: Institute for Protein Design, University of Washington

What it does: designs the sequence for a given backbone, and is the standard tool for that step in a design pipeline.

Sequence-recovery and design metrics are benchmark-dependent; no single universal accuracy.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (sequences/hour)≈ 8,000≈ 28,000≈ 72,000

RFdiffusion

Vendor: Institute for Protein Design, University of Washington

What it does: generates protein backbones against a motif or constraint — the generative half of de novo design, paired with sequence design above.

Generative success is design-task dependent; no single scalar accuracy.

RequirementMinimumMediumHigh
GPU typeL40S 48 GBA100 80 GBH100 80 GB ×2
VRAM48 GB80 GB160 GB
vCPUs163264
RAM64 GB128 GB256 GB
Server1× L40S 48 GB1× A100 SXM 80 GB2× H100 SXM 80 GB
Rate (sequences/hour)≈ 800≈ 2,800≈ 7,200

ProGen2

Vendor: Salesforce Research

What it does: generates protein sequences from a prompt or family context, for library design rather than single-target work.

Task-dependent; generative quality evaluated through sequence and function measures rather than generic accuracy.

RequirementMinimumMediumHigh
GPU typeRTX 3090L40S 48 GBA100 80 GB
VRAM24 GB48 GB80 GB
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 3090 24 GB1× L40S 48 GB1× A100 SXM 80 GB
Rate (sequences/hour)≈ 2,500≈ 8,800≈ 22,500

Choosing between them

If you need the best structure for a handful of targets, use the complex models and accept the compute. If you need thousands of structures, language-model folding is the economical route. Licensing is the real constraint in this group — several weights are non-commercial or require an application, which we resolve before a pipeline is built.

Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.

Protein structure prediction service AI Medical services Pricing

From benchmark to production

Send a representative sample, your expected volume and your latency target for protein structure prediction. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.