Upscaling increases an image’s resolution and, unlike simple enlargement, reconstructs plausible detail as it does so — turning a small or soft picture into one that can be printed, projected or examined.
The essential caveat is that the added detail is invented. These models produce what such a picture probably looked like, not what it certainly did, which makes them excellent for presentation and publishing and unsuitable as evidence. An upscaled licence plate is not a reading of that plate. Where the purpose is legal or forensic, the original must be retained and cited, and the upscaled version identified as a reconstruction.
Models in this group take a single image as input: a photograph, a scan or a screenshot. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one 1-megapixel image enlarged four times.
Real-ESRGAN
Vendor: Tencent ARC Lab
What it does: enlarges photographs up to four times with a clean, sharp result and no visible artefacts. The general default, and the fastest of the good options.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 1,200
≈ 4,000
≈ 9,000
SwinIR
Vendor: ETH Zürich
What it does: reconstructs fine texture more faithfully than the faster models, which matters for fabric, print and vegetation where invented detail is obvious.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 400
≈ 1,400
≈ 3,200
HAT
Vendor: University of Macau
What it does: the highest-quality option, recovering detail the others lose, at several times the cost per image. Reserved for material that will be printed large.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (images/hour)
≈ 150
≈ 550
≈ 1,300
BSRGAN
Vendor: ETH Zürich
What it does: trained on realistically degraded images — compressed, blurred, noisy — so it handles poor real-world sources better than models trained on clean downscales.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 700
≈ 2,400
≈ 5,500
StableSR
Vendor: Nanyang Technological University
What it does: generates detail for very degraded sources where other models have too little to work with. The most inventive, and therefore the least faithful.
Requirement
Minimum
Medium
High
GPU type
RTX 4090
A100 80 GB
2× A100 80 GB
VRAM
20 GB
80 GB
160 GB combined
vCPUs
8
16
32
RAM
32 GB
64 GB
128 GB
Server
1× RTX 4090 24 GB
1× A100 SXM 80 GB
2× A100 SXM 80 GB
Rate (images/hour)
≈ 60
≈ 220
≈ 500
ESRGAN
Vendor: Chinese Academy of Sciences
What it does: the earlier generation, lighter to run and still adequate for modest enlargement at high volume.
Requirement
Minimum
Medium
High
GPU type
RTX 3090
RTX 4090
A100 80 GB
VRAM
12 GB
24 GB
80 GB
vCPUs
8
12
24
RAM
32 GB
48 GB
96 GB
Server
1× RTX 3090 24 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (images/hour)
≈ 2,000
≈ 7,000
≈ 16,000
Choosing between them
Which model suits depends on your source material, the enlargement factor, and whether faithfulness or apparent sharpness matters more. Our consultants test the candidates on your own images and recommend one, along with the disclosure the output should carry.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Image upscaling / super-resolution. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.