The reported claim that matters commercially: affinity accuracy approaching free-energy perturbation at roughly a thousandth of the compute. If that holds on your targets, it changes what is affordable at the screening stage, which is exactly why we benchmark it on your own data first.
Structure quality is the other half. One model reports about 77 percent success on the standard ligand-docking benchmark without requiring a multiple sequence alignment at inference, which also removes a slow preprocessing step.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take a protein target with a ligand definition. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
Boltz-2
Vendor: MIT Jameel Clinic / Recursion and collaborators
What it does: predicts the complex and its binding affinity together, reportedly approaching free-energy perturbation accuracy at roughly a thousandth of the compute. If that holds on your targets it changes what is affordable at screening — which is why we test it on your data first.
Approaches FEP-level binding affinity accuracy at about 1000x lower compute; structure quality comparable to Boltz-1x (developer-reported).
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB ×2
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (molecules/hour)
≈ 2,500
≈ 8,800
≈ 22,500
Boltz-1
Vendor: Boltz team / MIT Jameel Clinic collaborators
What it does: predicts complex structure close to AlphaFold 3 quality under an MIT licence, which removes the licensing obstacle its peers carry.
Reported to approach AlphaFold 3 structural accuracy; benchmark-dependent.
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB ×2
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (molecules/hour)
≈ 2,500
≈ 8,800
≈ 22,500
Chai-1 / Chai-2
Vendor: Chai Discovery
What it does: co-folds protein with ligand at AlphaFold3-class quality and about 77 percent docking success without needing a sequence alignment at inference — removing a slow preprocessing step. Chai-1 weights are non-commercial by default.
AlphaFold3-class structure prediction; about 77 percent success on PoseBusters ligand docking without MSA at inference (developer-reported). Chai-1 weights non-commercial by default.
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB ×2
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (molecules/hour)
≈ 2,500
≈ 8,800
≈ 22,500
DiffDock
Vendor: MIT CSAIL / DiffDock authors
What it does: docks a ligand against a known structure and ranks the poses, lighter than the co-folding models when structure is already in hand.
Benchmark-dependent; docking success normally reported as the fraction of poses within RMSD thresholds.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (molecules/hour)
≈ 8,000
≈ 28,000
≈ 72,000
Choosing between them
If you need affinity as well as structure, that narrows the field to the models which predict it. If you need poses against a known structure, a dedicated docking model is lighter. Weights for one leading model are non-commercial by default, which we resolve against your intended use before building.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for binding affinity prediction. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.