Skip to main content

Model reference — Video

AI models for Video upscale resolution

Video upscaling raises footage to a higher resolution and reconstructs detail as it does so — standard definition archive to high definition, high definition to 4K — so that old material can be shown on modern screens.

Video is harder than still images for one reason: consistency between frames. A still-image model applied frame by frame invents slightly different detail each time, which the eye sees as shimmering and crawling texture. Models built for video look at several neighbouring frames together so the reconstructed detail stays stable as the picture moves. As with images, the added detail is invented and the result is a reconstruction, not a recovery — excellent for presentation, unsuitable as evidence.

Video upscale resolution service AI models for video files

Input type — Video

Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of standard-definition video enlarged four times.

BasicVSR++

Vendor: Nanyang Technological University

What it does: uses neighbouring frames together so reconstructed detail stays stable as the picture moves. The best general quality for archive restoration.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 2≈ 8≈ 20

RealBasicVSR

Vendor: Nanyang Technological University

What it does: the same approach trained on realistically degraded footage — compressed, noisy, analogue-sourced — which is what most archive material actually is.

RequirementMinimumMediumHigh
GPU typeRTX 4090A100 80 GB2× A100 80 GB
VRAM20 GB80 GB160 GB combined
vCPUs81632
RAM32 GB64 GB128 GB
Server1× RTX 4090 24 GB1× A100 SXM 80 GB2× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 2≈ 7≈ 18

Real-ESRGAN video pipeline

Vendor: Tencent ARC Lab

What it does: applies a fast still-image upscaler frame by frame with temporal smoothing. Far cheaper, at the cost of some shimmer on fine texture.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 12≈ 45≈ 110

SwinIR

Vendor: ETH Zürich

What it does: reconstructs texture very faithfully frame by frame, used where a short section must look its best rather than a whole archive being processed.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 4≈ 14≈ 32

GFPGAN

Vendor: Tencent ARC Lab

What it does: restores faces in the upscaled footage, which matters because a viewer’s attention goes to faces and a general upscaler leaves them soft.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 16≈ 55≈ 130

CodeFormer

Vendor: Nanyang Technological University

What it does: the face restoration option with a faithfulness control, so faces can be sharpened gently for archive work or strongly for display.

RequirementMinimumMediumHigh
GPU typeRTX 3090RTX 4090A100 80 GB
VRAM12 GB24 GB80 GB
vCPUs81224
RAM32 GB48 GB96 GB
Server1× RTX 3090 24 GB1× RTX 4090 24 GB1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)≈ 12≈ 42≈ 100

FFmpeg

Vendor: FFmpeg project

What it does: decodes, reassembles and encodes the footage around the upscaling step, and is where the delivery format and bitrate are set.

RequirementMinimumMediumHigh
GPU typeNo GPU requiredNo GPU requiredGPU-accelerated decode (NVENC/NVDEC)
VRAM8 GB
vCPUs2816
RAM4 GB16 GB32 GB
ServerCPU instance, 2 vCPUCPU instance, 8 vCPU1× RTX 4090 24 GB
Rate (video minutes upscaled per hour)≈ 1,200≈ 4,000≈ 12,000

Choosing between them

The right model depends on your source quality, the enlargement factor, whether faces feature prominently and how much processing time you can allow — costs here are per minute of footage, and they are significant. Our consultants test the candidates on your own material and recommend one with a realistic estimate of the total run.

At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.

Video upscale resolution service AI models for video files Pricing

From benchmark to production

Share a representative sample, expected volume, latency target and deployment location for Video upscale resolution. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.