AI models for pictures — document, face, object and image services | super-gpu
Skip to main content

Input type — Pictures

AI models for pictures

Twenty-four services that read a still image and return structured output — from a scanned form to a crowded square. One image over an API, or a whole archive in a batch, on the same dedicated cluster.

The right model for the job

Our AI advisors may assist you to select the best model or pipeline for the process or offer the most efficient hardware to perform the process.

Of course, you can bring your own model to run on our hardware.

In case you like to use your on-premise servers cluster or a certain cloud we may provision the needed services and build the pipeline.

One model chosen from the catalogue, matched to your workload and hardware

AI services for pictures

Our service catalogue include vast selection of services.

Part of the services based on stable AI models and part of the services based on our native code.

Each service can use several models and methods. After a short discussion we can decide on the right pipeline for the process.

{{ filterNote }}

Documents and forms

Handwriting recognition
Handwritten notes, ledgers and filled fields read into text.
Invoice / receipt extraction
Supplier, dates, line items and totals as fields, not free text.
Form extraction
Key-value pairs pulled from structured and semi-structured forms.
Table extraction
Printed tables recovered as rows and columns, ready for a spreadsheet.
Document classification
Each page sorted into your own document types.
Document visual question answering
Ask a question of a page and get the answer with the region it came from.
Document redaction
Names, numbers and chosen regions masked before a document is shared.

Faces and people

Face recognition
Faces matched against a face database you hold and control.
Face verification (1:1)
Two images compared to confirm they show the same person.
Face detection / face count
Faces located and counted, with no identity involved.
Crowd counter
People counted in dense scenes where bodies overlap.
Pose estimation (human)
Body keypoints per person, for posture, gesture and activity work.

Objects and scenes

General object detection
Common objects located and labelled with a box and a confidence.
Open-vocabulary object detection
Describe what to find in words and it is found, without retraining.
Image segmentation
Pixel-level masks per region, for area and coverage measurement.
Image classification
A whole image tagged against your own label set.
Visual similarity / duplicate image detection
Near-identical and re-edited copies found across a large library.
Industrial defect / anomaly detection
Departures from a known-good reference flagged on the line.

Description and image quality

Image-to-text captioning
A written description of the scene, for search, archives and alt text.
Image upscaling / super-resolution
Resolution raised on small or compressed source images.
Face restoration
Blurred and degraded faces recovered from low-quality frames.

Hidden data and provenance

Image steganography — hide text in photo
A message carried inside a photograph, invisible to the eye.
Image steganography — hide small binary data in photo
A small payload such as a key or signature embedded in the pixels.
Digital watermark / provenance embedding
Origin and ownership marked into the image so it survives re-encoding.

How to provide us pictures?

We support any OS / Encoding and format for images.

You can upload your images with SCP (behind VPN) to a URL we provide.

You can provide us URL, MQ or any API and we will pull the images directly from your server.

You can send a NAS or a DAS to our location.

You can use our AI services with an API.

How to get the results?

We can provide the results with SCP (behind VPN) directly to your server, to use your API for results, send you SCP (behind VPN) details to download the results or sent you back the NAS/DAS you provided width the results.

Other input types