Captioning writes a description of what an image shows, at whatever length and in whatever style you specify — a short line of alternative text for accessibility, a paragraph for a catalogue entry, or a structured note recording what an inspection photograph contains.
Two uses dominate. The first is accessibility and search: every image in an archive gets a description, which makes the archive searchable in words and usable by screen readers. The second is turning photographs into records — an engineer photographs a site and the caption becomes the written observation, following a template you define. Captions are generated from the image alone, so anything not visible in it will not appear; where a caption must include equipment identifiers or locations, those come from your own data rather than from the model.
Models in this group take a single image as input: a photograph, a scan or a screenshot. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one photograph at about 2 megapixels.
Qwen2.5-VL 7B / 72B
Vendor: Alibaba Cloud
What it does: writes accurate descriptions of photographs and follows detailed instructions about what to mention and what to leave out. The most generally capable option we deploy.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (7B, reduced precision)
A100 80 GB (7B, full precision)
2× H100 80 GB (72B model)
VRAM
16 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (images/hour)
≈ 400
≈ 1,600
≈ 900 (72B model, higher accuracy)
InternVL 3
Vendor: OpenGVLab (Shanghai AI Laboratory)
What it does: the same task with particular strength on technical images — charts, diagrams, equipment — where a general model gives a vague description.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (2B model)
A100 80 GB (8B model)
4× H100 80 GB (38B model or larger)
VRAM
12 GB
40 GB
320 GB combined
vCPUs
8
16
48
RAM
32 GB
64 GB
256 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
4× H100 SXM 80 GB
Rate (images/hour)
≈ 550
≈ 1,800
≈ 1,200
Llama 3.2 Vision
Vendor: Meta
What it does: writes fluent descriptions in several languages, useful where captions must be published in more than one market.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (11B model)
A100 80 GB (11B model)
2× H100 80 GB (90B model)
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (images/hour)
≈ 450
≈ 1,500
≈ 800
LLaVA 1.6
Vendor: University of Wisconsin–Madison and Microsoft Research
What it does: a well-established captioning and visual question answering model, lighter to run than the newest options and adequate for straightforward description.
Requirement
Minimum
Medium
High
GPU type
RTX 4090 (reduced precision)
A100 80 GB
2× H100 80 GB
VRAM
16 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
200 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× H100 SXM 80 GB
Rate (images/hour)
≈ 700
≈ 2,200
≈ 1,400
Florence-2
Vendor: Microsoft
What it does: a small model that produces short, factual captions at high volume, which is exactly what alternative text for a large image library needs.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 1,800
≈ 6,000
≈ 14,000
BLIP-2
Vendor: Salesforce
What it does: an earlier captioning model, cheap and predictable, still a sound choice for bulk captioning where brevity is acceptable.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (images/hour)
≈ 2,400
≈ 8,000
≈ 18,000
Choosing between them
The right model depends on caption length and precision, whether a template must be followed, and your volume. Our consultants review your images and the captions you need, then recommend a model and the instructions that hold it to your format.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Image-to-text captioning. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.