AI models for Video frame interpolation / slow motion
Frame interpolation creates new frames between existing ones. That allows smooth slow motion from footage shot at an ordinary frame rate, conversion between frame rate standards, and smoother playback of animation and screen recordings.
The model works out how everything in the picture moved between two frames and draws the intermediate positions. It does this well for smooth, predictable motion and less well where something appears from behind something else, where motion is very fast, or across a cut — which is why cuts must be detected and left alone. The invented frames are plausible, not recorded, so interpolated footage is presentation material: it should not be used to measure speed or timing, and for analysis the original frame rate must be kept.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video doubled from 25 to 50 frames per second.
RIFE
Vendor: Megvii
What it does: generates intermediate frames quickly with good quality, fast enough to process long footage or even to run live. The general default.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes interpolated per hour)
≈ 40
≈ 150
≈ 380
FILM
Vendor: Google
What it does: handles large motion between frames far better, which is what makes extreme slow motion — eight times or more — possible without visible tearing.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes interpolated per hour)
≈ 8
≈ 30
≈ 75
BasicVSR++
Vendor: Nanyang Technological University
What it does: used alongside interpolation where the footage also needs resolution and detail improved, so both are done in one consistent pass.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (video minutes interpolated per hour)
≈ 2
≈ 8
≈ 20
Real-ESRGAN video pipeline
Vendor: Tencent ARC Lab
What it does: sharpens the interpolated result frame by frame, useful where the source is soft and interpolation would otherwise emphasise its softness.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes interpolated per hour)
≈ 12
≈ 45
≈ 110
FFmpeg
Vendor: FFmpeg project
What it does: detects cuts so interpolation is not attempted across them, and handles decoding and encoding at the target frame rate. No GPU needed.
Requirement
Minimum
Medium
High
GPU type
No GPU required
No GPU required
GPU-accelerated decode (NVENC/NVDEC)
VRAM
—
—
8 GB
vCPUs
2
8
16
RAM
4 GB
16 GB
32 GB
Server
CPU instance, 2 vCPU
CPU instance, 8 vCPU
1× RTX 4090 24 GB
Rate (video minutes interpolated per hour)
≈ 1,200
≈ 4,000
≈ 12,000
Choosing between them
Which model fits depends on your source frame rate, the slow-motion factor and how much fast or occluded motion the footage contains. Our consultants test the candidates on your own material and recommend one, with a realistic estimate of processing time.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Video frame interpolation / slow motion. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.