Skip to main content

Model reference — Documents

AI models for Private chatbot / enterprise assistant

A private chatbot answers staff questions in your own words, from your own material — policies, manuals, past tickets, product data — and runs entirely inside your network, so no question and no document is ever sent to an outside provider.

The engine is a large language model. Its size, measured in billions of parameters (written 8B, 70B and so on), sets both the quality of the answers and the hardware needed to serve them. A smaller model answers straightforward questions quickly on modest hardware; a larger model reasons over longer material and follows complicated instructions, at a higher cost per answer. Company knowledge is supplied by retrieval — the relevant passages are found and handed to the model with each question — rather than by retraining it.

Private chatbot / enterprise assistant service AI models for documents

Input type — Documents

Models in this group take text as input: the user’s question plus passages retrieved from your own documents. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one question answered in about 300 words, with roughly 4,000 words of retrieved context supplied to the model.

Llama 3.3 70B

Vendor: Meta

What it does: a general-purpose assistant model that follows instructions well and writes clear prose in a dozen major languages. The usual first choice for a company-wide assistant where answer quality matters more than cost per answer.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)1× H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 120≈ 600≈ 3,000

Llama 3.1 8B

Vendor: Meta

What it does: a small assistant model that answers short factual questions and summarises retrieved passages. Fast and cheap enough to put in front of a whole workforce, but it loses accuracy on long chains of reasoning.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 40901× H100 80 GB
VRAM16 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB128 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× H100 SXM 80 GB
Rate (answers/hour)≈ 500≈ 1,400≈ 6,000

Qwen2.5 72B

Vendor: Alibaba Cloud

What it does: an assistant model with unusually broad language coverage, including Chinese, Japanese, Korean, Arabic and most European languages. Chosen when the same assistant must serve offices in several regions.

RequirementMinimumMediumHigh
GPU type2× RTX 4090 (reduced precision)1× H100 80 GB4× H100 80 GB
VRAM48 GB combined80 GB320 GB combined
vCPUs162464
RAM64 GB128 GB512 GB
Server2× RTX 4090 24 GB1× H100 SXM 80 GB4× H100 SXM 80 GB
Rate (answers/hour)≈ 110≈ 560≈ 2,800

Mistral Small 3

Vendor: Mistral AI

What it does: a mid-sized model that fits on a single professional card while still handling multi-step instructions. A good balance when one server must serve a department without a queue forming.

RequirementMinimumMediumHigh
GPU typeRTX 40901× L40S 48 GB1× H100 80 GB
VRAM24 GB48 GB80 GB
vCPUs121632
RAM48 GB64 GB128 GB
Server1× RTX 4090 24 GB1× L40S 48 GB1× H100 SXM 80 GB
Rate (answers/hour)≈ 260≈ 700≈ 2,200

Gemma 2 27B

Vendor: Google

What it does: a compact model tuned for safe, on-topic replies, which makes it suitable for assistants that face customers as well as staff. It refuses out-of-scope requests more readily than most models its size.

RequirementMinimumMediumHigh
GPU typeRTX 4090 (reduced precision)1× L40S 48 GB2× H100 80 GB
VRAM20 GB48 GB160 GB combined
vCPUs121648
RAM48 GB64 GB256 GB
Server1× RTX 4090 24 GB1× L40S 48 GB2× H100 SXM 80 GB
Rate (answers/hour)≈ 200≈ 620≈ 2,400

Phi-4 14B

Vendor: Microsoft

What it does: a small model trained heavily on reasoning tasks, so it handles arithmetic and step-by-step logic better than its size suggests. Suited to assistants that answer questions about figures and rules.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 40901× L40S 48 GB
VRAM16 GB24 GB48 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× L40S 48 GB
Rate (answers/hour)≈ 380≈ 1,100≈ 3,200

DeepSeek-V3

Vendor: DeepSeek

What it does: a very large model that activates only part of itself for each question, so it reasons at close to frontier quality while costing less per answer than its total size implies. Reserved for assistants doing genuinely hard analytical work.

RequirementMinimumMediumHigh
GPU typeNot practical below 8 cards8× H100 80 GB16× H100 80 GB
VRAM640 GB combined1,280 GB combined
vCPUs96192
RAM1 TB2 TB
Server8× H100 SXM 80 GB2× 8× H100 SXM 80 GB
Rate (answers/hour)≈ 900≈ 2,000

Choosing between them

The right size depends on how hard the questions are and how many people ask them at once. Our consultants review the material the assistant must cover, the reasoning the answers require and your expected number of daily users, then recommend a model and a serving configuration to match.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Private chatbot / enterprise assistant service AI models for documents Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Private chatbot / enterprise assistant. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.