Skip to main content

Model reference — Documents

AI models for Content moderation - text

Text moderation checks messages, comments, reviews and uploaded copy against your own policy and flags what breaks it — abuse, threats, sexual content, self-harm, fraud, hate speech — before it reaches other people.

A moderation model produces a category and a severity, not a simple yes or no, because the action taken differs: some content is blocked outright, some is held for review, some is published with a warning. Two things decide whether a deployment works. The first is that your policy is written down precisely enough for a model to apply. The second is the balance between missing violations and blocking innocent posts — a threshold that is a business decision, not a technical one, and one we set with you and then measure.

Content moderation - text service AI models for documents

Input type — Documents

Models in this group take text as input: a message, comment, review or document checked against your policy. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one message of about 100 words checked against a written policy.

Llama Guard 3

Vendor: Meta

What it does: checks text against a written safety policy and returns which category it breaks and how severely. Built for moderation specifically, so it can be pointed at your own policy wording rather than a fixed list of categories.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090H100 80 GB
VRAM16 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB128 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× H100 SXM 80 GB
Rate (messages/hour)≈ 3,000≈ 9,000≈ 30,000

ShieldGemma

Vendor: Google

What it does: the same policy-driven checking in a smaller model, tuned to be cautious. Suited to a first-pass filter running on every message before a larger model looks at what it flags.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090H100 80 GB
VRAM16 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB128 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× H100 SXM 80 GB
Rate (messages/hour)≈ 4,000≈ 12,000≈ 36,000

Detoxify

Vendor: Unitary

What it does: scores text for toxicity, insult, threat and identity attack. Very small and very fast, which makes it the practical choice for filtering high-volume comment streams in real time.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (messages/hour)≈ 40,000≈ 160,000≈ 450,000

DeBERTa v3

Vendor: Microsoft

What it does: fitted to your own policy categories from examples your moderators have already actioned, so it enforces your decisions rather than a general standard. The most accurate option where moderation history exists.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (messages/hour)≈ 30,000≈ 120,000≈ 340,000

XLM-RoBERTa

Vendor: Meta

What it does: applies the same policy across a hundred languages with one model, which prevents a policy being enforced strictly in English and loosely everywhere else.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (messages/hour)≈ 26,000≈ 100,000≈ 300,000

Mistral Small 3

Vendor: Mistral AI

What it does: judges long or ambiguous text — a three-paragraph post, a sarcastic review — and explains which sentence broke the policy, which is what a moderator needs in order to uphold or overturn the decision.

RequirementMinimumMediumHigh
GPU typeRTX 4090L40S 48 GBH100 80 GB
VRAM24 GB48 GB80 GB
vCPUs121632
RAM48 GB64 GB128 GB
Server1× RTX 4090 24 GB1× L40S 48 GB1× H100 SXM 80 GB
Rate (messages/hour)≈ 1,300≈ 3,800≈ 11,000

Choosing between them

Moderation is judged on two error rates at once, and the right balance depends on your audience and your legal exposure. Our consultants review your policy and a sample of your traffic, then recommend a model, the thresholds for block, hold and allow, and the volume of human review the result implies.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Content moderation - text service AI models for documents Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Content moderation - text. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.