Models in this group take protein sequence, structure or design constraints. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
AlphaFold 3
Vendor: Google DeepMind
What it does: predicts complexes — protein with ligand, nucleic acid or antibody — with large gains over the previous generation. Code is open, weights require an application and are non-commercial, which we resolve before a pipeline is built.
Large accuracy gains over AlphaFold 2 for protein-ligand, protein-nucleic acid and antibody-antigen complexes (independent, Nature 2024). Code open; weights require an application, non-commercial.
AlphaFold 2
Vendor: Google DeepMind
What it does: remains the accuracy reference for single-chain structure at near-experimental quality, and its licensing is far simpler than AlphaFold 3.
CASP14 near-experimental structure accuracy; accuracy is target-dependent rather than a single percentage.
OpenFold
Vendor: OpenFold Consortium
What it does: reproduces AlphaFold 2 performance closely under a licence you can actually deploy, which is usually the deciding factor.
Reported to closely reproduce AlphaFold 2 performance; structure accuracy is target-dependent.
ESMFold
Vendor: Meta AI
What it does: folds from sequence alone at an order of magnitude less compute than AlphaFold 2, trading some accuracy — which is what makes proteome-scale prediction affordable.
Structure prediction TM-score about 0.68 average, an order of magnitude faster than AlphaFold 2 (independent).
ESM-2 (8M–15B) / ESM C
Vendor: Meta AI
What it does: turns a protein sequence into embeddings and contact maps, from 8M to 15B parameters. The base layer under most function-prediction work.
Task-dependent across protein prediction tasks; no single universal accuracy.
ESM3-open-small (1.4B)
Vendor: EvolutionaryScale
What it does: generates as well as understands, taking sequence, structure and function as context — for design work rather than prediction. Non-commercial community licence.
Task-dependent across protein understanding and generation benchmarks. Non-commercial community licence.
ProtT5
Vendor: Rostlab
What it does: produces protein embeddings for function and structure-related prediction, a well-established and permissively licensed base.
Task-dependent across protein-function and structure-related benchmarks.
Ankh
Vendor: Ankh authors
What it does: is a protein encoder tuned for efficiency, delivering competitive downstream accuracy at a smaller footprint.
Task-dependent across protein benchmarks.
ProteinMPNN
Vendor: Institute for Protein Design, University of Washington
What it does: designs the sequence for a given backbone, and is the standard tool for that step in a design pipeline.
Sequence-recovery and design metrics are benchmark-dependent; no single universal accuracy.
RFdiffusion
Vendor: Institute for Protein Design, University of Washington
What it does: generates protein backbones against a motif or constraint — the generative half of de novo design, paired with sequence design above.
Generative success is design-task dependent; no single scalar accuracy.
ProGen2
Vendor: Salesforce Research
What it does: generates protein sequences from a prompt or family context, for library design rather than single-target work.
Task-dependent; generative quality evaluated through sequence and function measures rather than generic accuracy.