Skip to main content

Model reference — Video

AI models for Video semantic search

Video semantic search finds the moment in an archive that matches a description — "a forklift reversing near the loading bay", "the part where they discuss the Bristol contract" — and returns the file and timecode.

As with audio, two indexes serve two kinds of question. Speech is transcribed and indexed as text, which answers questions about what was said precisely and cheaply. Pictures are indexed by fingerprinting sampled frames against written descriptions, which answers questions about what was shown. The frame sampling rate decides both the cost of indexing and the granularity of the results: sampling once per second finds a two-second event, sampling once per minute will miss it entirely.

Video semantic search service AI models for video files

Input type — Video

Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video at 25 frames per second.

InternVideo2

Vendor: Shanghai AI Laboratory

What it does: fingerprints stretches of video so a search matches an action or event across frames, not a single still. The strongest option for searching what happens.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 40≈ 150≈ 380

X-CLIP

Vendor: Microsoft

What it does: matches a written description against a short clip, which suits searching for a described activity in operational footage.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 120≈ 450≈ 1,100

SigLIP

Vendor: Google

What it does: fingerprints sampled frames for text-based search, accurate and cheap enough to index a very large archive frame by frame.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 700≈ 2,600≈ 7,000

CLIP

Vendor: OpenAI

What it does: the established frame-fingerprinting model, lighter and widely supported, adequate where searches are for objects and scenes rather than subtle actions.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM6 GB24 GB80 GB
vCPUs4824
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 900≈ 3,400≈ 9,000

faster-whisper

Vendor: SYSTRAN

What it does: transcribes the speech so what was said becomes searchable text with timecodes back into the footage.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 1,200≈ 4,200≈ 10,800

BGE-M3

Vendor: Beijing Academy of Artificial Intelligence

What it does: indexes the transcripts by meaning, so a search finds the passage that answers the question rather than only matching typed words.

RequirementMinimumMediumHigh
GPU typeRTX 3060RTX 4090A100 80 GB
VRAM8 GB24 GB80 GB
vCPUs61224
RAM16 GB32 GB64 GB
Server1× RTX 3060 12 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 6,000 pages/hour≈ 30,000 pages/hour≈ 90,000 pages/hour

BGE Reranker v2-M3

Vendor: Beijing Academy of Artificial Intelligence

What it does: re-sorts candidate moments by how well they match the query, which puts the right clip first rather than tenth.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (query-passage pairs/hour)≈ 40,000≈ 150,000≈ 400,000

Grounding DINO

Vendor: IDEA Research

What it does: locates the described object within the retrieved frame, so a result points at the forklift rather than at the whole shot.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes processed per hour)≈ 40≈ 160≈ 400

Choosing between them

The right design depends on whether your searches are about speech, imagery or both, how long the events you look for last, and the size of the archive. Our consultants review your footage and the searches you need, then recommend the sampling rate, the models and the index structure.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Video semantic search service AI models for video files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Video semantic search. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.