Models in this group take SMILES, molecular graphs or 3D coordinates. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
TxGemma-2B / TxGemma-9B / TxGemma-27B
Vendor: Google DeepMind
What it does: answers therapeutic questions in plain language as well as predicting properties, and beat or matched specialist models on 64 of 66 tasks in the standard therapeutics suite. Three sizes, and the 27B costs accordingly.
Beats or matches specialist models on 64 of 66 Therapeutics Data Commons tasks (developer-reported).
Vendor: IBM Research
What it does: is at or near state of the art on ten property benchmarks and produces representations your own regressors sit on top of.
State of the art or near it on 10 MoleculeNet benchmarks (developer-reported).
ChemBERTa-2 / ChemBERTa
Vendor: DeepChem and HuggingMolecules community
What it does: screens at very high throughput on modest hardware, at ROC-AUC 0.70–0.85 across ADMET tasks — usually the best cost per molecule for a first filter.
Competitive on MoleculeNet ADMET tasks, ROC-AUC 0.70–0.85 by task (independent).
Uni-Mol
Vendor: DP Technology
What it does: uses 3D coordinates rather than strings, which is what conformation-dependent properties and docking-adjacent work require.
Task-dependent across molecular property, conformation and docking benchmarks.
MegaMolBART
Vendor: NVIDIA
What it does: generates molecules as well as embedding them, for library expansion rather than scoring alone.
Task-dependent; evaluated on molecular representation and generation tasks.
MoleculeSTM
Vendor: MoleculeSTM authors
What it does: links molecules to text, so a library can be searched — and molecules edited — by description.
Task-dependent across molecule-text retrieval and molecular editing benchmarks.
GIT-Mol
Vendor: GIT-Mol authors
What it does: takes graphs, images and text together for captioning, question answering and property prediction from one model.
Task-dependent across molecular captioning, QA and property tasks.
nach0
Vendor: nach0 authors
What it does: works across chemistry and natural language in one model, for tasks that mix the two.
Task-dependent across chemistry-language tasks.
Vendor: Academic community (DDI-BioBERT lineage)
What it does: extracts drug-drug interaction statements from literature and labels at micro-F1 0.79–0.83 — cheap, well characterised pharmacovigilance text mining.
Micro-F1 0.79–0.83 on DDIExtraction-2013 (independent).