Reported F1 sits between 0.66 and 0.85 across depression, stress and suicidal-ideation detection datasets — wide, dataset-dependent, and nowhere near a level that would support autonomous decisions.
The explainable models are the more useful development: they approach state-of-the-art discriminative accuracy while also producing an explanation that annotators rate highly, which is what makes a flag reviewable. Counselling-dialogue models score higher on empathy ratings than general models and have no clinical outcome evidence at all.
Every model on this page runs as part of a managed AI pipeline in our GPU clusters, with a dedicated private cluster in our cloud or an on-premise installation where medical governance requires it. Output is decision support for a qualified professional to review, not a diagnosis.
Models in this group take social or clinical text and return classifications, optionally with explanations. Each table gives three hardware tiers — Minimum, the smallest configuration on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is what one server of that tier processes per hour. Use these figures for initial sizing only. Before production we benchmark your own data to confirm accuracy, latency, throughput and cost.
MentalBERT / MentalRoBERTa
Vendor: Georgia Institute of Technology and community
What it does: scores indicators across depression, stress and suicidal-ideation datasets at F1 0.66–0.85 — adequate for research and measurement, and nowhere near a level that supports an autonomous decision.
F1 0.66–0.85 across depression, stress and suicidal-ideation detection datasets (independent).
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
L40S 48 GB
VRAM
12 GB
24 GB
48 GB
vCPUs
4
8
16
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× L40S 48 GB
Rate (documents/hour)
≈ 12,000
≈ 42,000
≈ 108,000
MentaLLaMA-chat-7B / 13B / 33B
Vendor: University of Manchester and community
What it does: classifies and explains in the same output, with explanations annotators rate highly. Where a person will read and act on a flag, an explanation is what makes that possible.
Approaches state-of-the-art discriminative models while generating explanations rated highly by annotators (independent, WWW 2024).
Requirement
Minimum
Medium
High
GPU type
L40S 48 GB
A100 80 GB
H100 80 GB ×2
VRAM
48 GB
80 GB
160 GB
vCPUs
16
32
64
RAM
64 GB
128 GB
256 GB
Server
1× L40S 48 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (documents/hour)
≈ 700
≈ 2,500
≈ 6,300
SoulChat / PsyChat
Vendor: Community (Psych-LLM and SoulChat lineage)
What it does: produces empathic counselling-style replies, rated higher on empathy than general models with no clinical outcome evidence behind it. Not for unsupervised crisis or triage use, and we will not build it that way.
Higher empathy and helpfulness ratings than general LLMs in human evaluation; no clinical outcome evidence (developer-reported). Not for unsupervised crisis or triage use.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
L40S 48 GB
A100 80 GB
VRAM
24 GB
48 GB
80 GB
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 3090 24 GB
1× L40S 48 GB
1× A100 SXM 80 GB
Rate (documents/hour)
≈ 2,000
≈ 7,000
≈ 18,000
Choosing between them
For research and measurement, the classifiers are adequate and cheap. For anything a person will read and act on, use an explainable model so the reasoning is visible. For crisis pathways, use none of them autonomously — the design constraint is human review, and we build to it.
Accuracy figures above are those the producers and independent evaluations report, on their own test sets. They are a shortlist tool, not a prediction of what you will see. At the start of a project we run a short proof of concept on a sample of your own data, which replaces them with real figures — so the cost and the schedule for the full engagement are known before anything is committed.
Send a representative sample, your expected volume and your latency target for mental health text analysis. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster — with the cost per unit of work known before you commit.