Video upscaling raises footage to a higher resolution and reconstructs detail as it does so — standard definition archive to high definition, high definition to 4K — so that old material can be shown on modern screens.
Video is harder than still images for one reason: consistency between frames. A still-image model applied frame by frame invents slightly different detail each time, which the eye sees as shimmering and crawling texture. Models built for video look at several neighbouring frames together so the reconstructed detail stays stable as the picture moves. As with images, the added detail is invented and the result is a reconstruction, not a recovery — excellent for presentation, unsuitable as evidence.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of standard-definition video enlarged four times.
BasicVSR++
Vendor: Nanyang Technological University
What it does: uses neighbouring frames together so reconstructed detail stays stable as the picture moves. The best general quality for archive restoration.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 2
≈ 8
≈ 20
RealBasicVSR
Vendor: Nanyang Technological University
What it does: the same approach trained on realistically degraded footage — compressed, noisy, analogue-sourced — which is what most archive material actually is.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 2
≈ 7
≈ 18
Real-ESRGAN video pipeline
Vendor: Tencent ARC Lab
What it does: applies a fast still-image upscaler frame by frame with temporal smoothing. Far cheaper, at the cost of some shimmer on fine texture.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 12
≈ 45
≈ 110
SwinIR
Vendor: ETH Zürich
What it does: reconstructs texture very faithfully frame by frame, used where a short section must look its best rather than a whole archive being processed.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 4
≈ 14
≈ 32
GFPGAN
Vendor: Tencent ARC Lab
What it does: restores faces in the upscaled footage, which matters because a viewer’s attention goes to faces and a general upscaler leaves them soft.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 16
≈ 55
≈ 130
CodeFormer
Vendor: Nanyang Technological University
What it does: the face restoration option with a faithfulness control, so faces can be sharpened gently for archive work or strongly for display.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes upscaled per hour)
≈ 12
≈ 42
≈ 100
FFmpeg
Vendor: FFmpeg project
What it does: decodes, reassembles and encodes the footage around the upscaling step, and is where the delivery format and bitrate are set.
Requirement
Minimum
Medium
High
GPU type
No GPU required
No GPU required
GPU-accelerated decode (NVENC/NVDEC)
VRAM
—
—
8 GB
vCPUs
2
8
16
RAM
4 GB
16 GB
32 GB
Server
CPU instance, 2 vCPU
CPU instance, 8 vCPU
1× RTX 4090 24 GB
Rate (video minutes upscaled per hour)
≈ 1,200
≈ 4,000
≈ 12,000
Choosing between them
The right model depends on your source quality, the enlargement factor, whether faces feature prominently and how much processing time you can allow — costs here are per minute of footage, and they are significant. Our consultants test the candidates on your own material and recommend one with a realistic estimate of the total run.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Video upscale resolution. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.