This service measures how long people wait and where they linger: queue length at a till or gate, waiting time from joining to being served, and dwell time in a zone such as an aisle, showroom or platform.
The measurements come from tracking people through defined zones and timing their presence. That makes the definition of the zones as important as the models — where the queue starts, which area counts as being served — and those are drawn once with your operations team on the actual camera views. Nobody is identified: a track exists only while a person is in view and is discarded afterwards, which is what allows waiting time to be measured continuously without holding personal data.
Models in this group take a recorded video file as input and are applied frame by frame. Each specification table gives three hardware tiers — Minimum, the smallest setup on which the model runs correctly; Medium, the usual production configuration; and High, a configuration sized for peak volume. The rate is the number of sample inputs processed per hour on that hardware. Use these rates for initial sizing. Before production, benchmark your own data to validate accuracy, latency, throughput and cost. The sample input here is one minute of 1080p video at 25 frames per second.
YOLO11
Vendor: Ultralytics
What it does: detects people on each frame as the basis for every timing measurement. Accurate and fast enough to process many cameras on one card.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 1,600
≈ 6,000
≈ 16,000
RT-DETR
Vendor: Baidu (PaddlePaddle)
What it does: produces steadier detections in crowded frames, which improves timing accuracy where a queue is tightly packed.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 700
≈ 2,600
≈ 7,000
ByteTrack
Vendor: Huazhong University of Science and Technology
What it does: maintains each person’s track through the zone, which is what produces a waiting time rather than a headcount. The default tracker.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 1,200
≈ 4,000
≈ 10,000
BoT-SORT
Vendor: Tel Aviv University
What it does: holds a track through occlusion, which matters in a queue where people are continually hidden behind one another and a lost track means a lost measurement.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 800
≈ 2,800
≈ 7,000
RTMPose
Vendor: Shanghai AI Laboratory (OpenMMLab)
What it does: reads body orientation, which distinguishes someone queueing from someone walking past or browsing nearby — the main source of false queue length.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
8 GB
24 GB
80 GB
vCPUs
6
12
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 900
≈ 3,200
≈ 8,500
CSRNet
Vendor: University of Illinois
What it does: estimates crowd density as a cross-check on queue length at peak, when individual tracking becomes unreliable.
Requirement
Minimum
Medium
High
GPU type
RTX 3060
RTX 4090
A100 80 GB
VRAM
6 GB
24 GB
80 GB
vCPUs
4
8
24
RAM
16 GB
32 GB
64 GB
Server
1× RTX 3060 12 GB
1× RTX 4090 24 GB
1× A100 SXM 80 GB
Rate (video minutes processed per hour)
≈ 500
≈ 1,800
≈ 4,800
Choosing between them
What is measurable depends on your camera coverage and how orderly the queues are. Our consultants review your views, define the zones with you, validate the timings against manual observation, and recommend the models and the reporting.
At the start of a project we may run a short proof of concept on a sample of your own data, measuring the accuracy and the throughput the model actually achieves on your material. That replaces the estimates on this page with real figures, so the cost and the schedule for the full engagement are known before it is committed.
Share a representative sample, expected volume, latency target and deployment location for Queue length / dwell-time analytics. We benchmark the shortlisted models, recommend the lowest-cost GPU configuration that meets the target, and scale it from pilot capacity to a dedicated production cluster.