Video semantic search finds the moment in an archive that matches a description — "a forklift reversing near the loading bay", "the part where they discuss the Bristol contract" — and returns the file and timecode.
As with audio, two indexes serve two kinds of question. Speech is transcribed and indexed as text, which answers questions about what was said precisely and cheaply. Pictures are indexed by fingerprinting sampled frames against written descriptions, which answers questions about what was shown. The frame sampling rate decides both the cost of indexing and the granularity of the results: sampling once per second finds a two-second event, sampling once per minute will miss it entirely.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video at 25 frames per second.
InternVideo2
Vendor: Shanghai AI Laboratory
What it does: fingerprints stretches of video so a search matches an action or event across frames, not a single still. The strongest option for searching what happens.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 40
≈ 150
≈ 380
X-CLIP
Vendor: Microsoft
What it does: matches a written description against a short clip, which suits searching for a described activity in operational footage.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 120
≈ 450
≈ 1,100
SigLIP
Vendor: Google
What it does: fingerprints sampled frames for text-based search, accurate and cheap enough to index a very large archive frame by frame.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 700
≈ 2,600
≈ 7,000
CLIP
Vendor: OpenAI
What it does: the established frame-fingerprinting model, lighter and widely supported, adequate where searches are for objects and scenes rather than subtle actions.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 900
≈ 3,400
≈ 9,000
faster-whisper
Vendor: SYSTRAN
What it does: transcribes the speech so what was said becomes searchable text with timecodes back into the footage.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 1,200
≈ 4,200
≈ 10,800
BGE-M3
Vendor: Beijing Academy of Artificial Intelligence
What it does: indexes the transcripts by meaning, so a search finds the passage that answers the question rather than only matching typed words.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 6,000 pages/hour
≈ 30,000 pages/hour
≈ 90,000 pages/hour
BGE Reranker v2-M3
Vendor: Beijing Academy of Artificial Intelligence
What it does: re-sorts candidate moments by how well they match the query, which puts the right clip first rather than tenth.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (query-passage pairs/hour)
≈ 40,000
≈ 150,000
≈ 400,000
Grounding DINO
Vendor: IDEA Research
What it does: locates the described object within the retrieved frame, so a result points at the forklift rather than at the whole shot.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 40
≈ 160
≈ 400
Choosing between them
The right design depends on whether your searches are about speech, imagery or both, how long the events you look for last, and the size of the archive. Our consultants review your footage and the searches you need, then recommend the sampling rate, the models and the index structure.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Video semantic search. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.