AI models for video files — face recognition, tracking, transcription, subtitles and summarization | super-gpu
Skip to main content

Input type — Video files

AI models for video files

Sixteen services that read recorded footage and return structured output — who and what is in it, what was said, and a cleaner file back. One clip over an API, or a whole archive in a batch, on the same dedicated cluster.

The right model for the job

Our AI advisors may assist you to select the best model or pipeline for the process or offer the most efficient hardware to perform the process.

Of course, you can bring your own model to run on our hardware.

In case you like to use your on-premise servers cluster or a certain cloud we may provision the needed services and build the pipeline.

One model chosen from the catalogue, matched to your workload and hardware

AI services for video files

Our service catalogue include vast selection of services.

Part of the services based on stable AI models and part of the services based on our native code.

Each service can use several models and methods. After a short discussion we can decide on the right pipeline for the process.

{{ filterNote }}

People and faces

Face recognition
Faces matched against a face database you hold and control.
Crowd counter
People counted in dense scenes where bodies overlap.
People entrance / exit counting
Movement across a line or doorway counted in each direction.
Queue length / dwell-time analytics
How long people wait and how long they stay, per zone.

Objects and text

General object detection and tracking
Objects located, labelled and followed from frame to frame.
Video OCR / text recognition
On-screen and in-scene text read across the footage, with timecodes.

Audio and speech

Video to audio extraction
The audio track separated from the video as a clean file.
Video audio to text
Speech in the footage transcribed, with timings and speaker turns.
Narrator to video - add supplied narrator audio
Narration you supply mixed into the video and aligned to it.
Narrator to video - generate narration from text
A narration track generated from your script and laid over the video.
Automatic subtitles / captions
Timed subtitle files, burned in or delivered alongside.

Understanding and search

Video summarization
A long recording reduced to a written brief or a short cut.
Video semantic search
Find the moment by describing it, across a whole archive.

Quality and provenance

Video upscale resolution
Resolution raised on low-quality or legacy footage.
Video frame interpolation / slow motion
Frames generated between frames for smoother or slowed playback.
Digital watermark / provenance embedding
Origin and ownership marked into the footage so it survives re-encoding.

How to provide us video files?

We support any OS / Encoding and format for video files.

You can upload your video files with SCP (behind VPN) to a URL we provide.

You can provide us URL, MQ or any API and we will pull the video files directly from your server.

You can send a NAS or a DAS to our location.

You can use our AI services with an API.

How to get the results?

We can provide the results with SCP (behind VPN) directly to your server, to use your API for results, send you SCP (behind VPN) details to download the results or send you back the NAS/DAS you provided with the results.

Other input types