Not a platform, not a demo — a working system trained on your images and wired into your process.
| Reading the structure of a scan or photograph | which page this is, what is on it, what is missing |
| Finding required elements | a signature, a stamp, a marking, a seal |
| Comparing against a reference | what differs from the sample, which part is absent |
| Counting and locating | how many objects, and where exactly in the frame |
Vision is sold as a capability — the machine can see. It is really a choice between four methods, and that choice is not made on what is technically possible.
| Method | What it does | When to take it | Cost of error |
|---|---|---|---|
| Classification | Assigns the whole image to a class | Few classes, and they do not overlap | A mistake is corrected on the next screen. Cheapest to label, cheapest to fix |
| Detection | Finds an object and where it sits in the frame | Position and count matter, not just presence | A miss costs more than a false positive: you can dismiss a box that should not be there, not one that was never drawn |
| Segmentation | Traces the outline pixel by pixel | Area, shape or boundary is the answer | By far the most expensive to label, so the precision has to be worth it |
| Reference comparison | Finds what differs from a known-good sample | A reference sample exists and is stable | A false alarm is cheap, a miss is not — so it is tuned deliberately over-sensitive |
The first question every buyer asks, and the one no vendor page answers. There is no single number, but there is arithmetic.
What matters is not how many images you have but how many you have per cell — one class under one set of capture conditions. Lighting, angle, camera, background and format each split the data again, so the same archive gives opposite answers depending on how many cells it must cover:
A narrow task under controlled lighting works on a few hundred per class; the same accuracy across a dozen classes shot on whatever phone was nearest needs an order of magnitude more, and the difference is not the model. Three things move the arithmetic and we use all three: transfer learning, so the model learns your task rather than learning to see; synthetic and augmented data where the variation is mechanical rather than semantic; and a phased launch with a person in the loop, where each reviewer correction becomes a label. Whether your data is clean and accessible enough to start at all is the broader question, in data readiness for AI.
Client names withheld under NDA. See full case studies →
Scanned maintenance packs arriving as hundreds of pages of mixed forms. The system classifies each page into one of six document types, detects missing signatures and stamps, finds unfilled checklist cells and verifies date-and-signature pairing, then returns an annotated report before a specialist signs off — advisory by construction. Read the case →
This is computer vision applied to documents, and we would rather say so than let it read as something else. The mechanics — reading structure, finding a required element, detecting what is absent, comparing against a reference — carry over to photographs and other media, but the engagement was document control under an aviation rulebook and that is the one we can describe. Where the same mechanic meets a regulated paper process it lives in document intelligence; this page is the vision layer underneath.
| We build | We decline |
|---|---|
| Reading structure and elements on scans and photographs · detecting what is absent · comparing against a reference sample · classifying and routing an incoming stream of images · counting and locating objects in a frame | Machine vision on moving equipment and navigation — a different discipline, and a different cluster: Physical AI · biometrics and face recognition · medical diagnosis from images · industrial inspection of manufactured parts, which we have not done and will not present as if we had |
The last one is worth stating plainly. Visual quality control of physical products is the largest industrial application of vision and several vendors competing for this page sell it — we have not shipped it. The mechanics above are real; that deployment is not ours to claim.
Three phases. The second is a product you can buy on its own.
| Phase | What we do | What we need from you | What you get |
|---|---|---|---|
| Assessment | Image types, capture conditions, and what counts as an error — agreed with whoever does the checking today | A sample of real images, awkward ones included | An estimate of achievable accuracy on your data and the labelling volume it implies |
| Design | Method chosen against cost of error, metrics fixed, and the point where a human stays in the loop | Sign-off on acceptance criteria before anything is trained | An AI Solution Blueprint — from $30,000, credited against the build, yours to implement with anyone |
| Build | Training, integration, and monitoring of quality in production | A named owner of the process | The system, plus a written procedure for retraining as conditions drift |
Where they are stored, who can reach them, and a closed perimeter where that is required — open-weight models on your own hardware, nothing leaving the network. Details in security and compliance.
Pinned and dated versions, a reproducible build, and an honest position on drift: a vision model degrades as cameras, lighting and forms change, so the retraining procedure ships with it rather than being a later conversation. Operational detail in MLOps.
Fixed price against a scope agreed before work starts. No time and materials, and the achievable accuracy stated in the Blueprint rather than discovered in month four.
It depends on cells, not on the total. A cell is one class under one set of capture conditions, and what matters is images per cell. Five thousand images across three classes and one condition gives about 1,670 per cell and is workable; the same five thousand across twelve classes and four conditions gives about 104 and is thin. So the answer is a question back: how many classes, and how many conditions?
On your images, which is the only place the question can be answered. We measure it on a sample of your real data, awkward cases included, and state the figure before you commit to a build. Treat any number quoted before a vendor has seen your images with suspicion.
Five to six weeks to a specification with acceptance criteria — the Blueprint, which you can buy on its own. A build on top typically runs four to six months depending on labelling volume and integration depth. First results on your own images come inside the assessment, not at the end.
Transfer learning, synthetic and augmented data, and a phased launch where each reviewer correction becomes a label. None of the three fixes a task with more classes and conditions than the data can ever cover.
Fine-tune, in almost every case. Training from scratch means teaching a model to see before it learns your task, and there is no payback for that when the general capability already exists and is good.
Yes — open-weight models on your own hardware, images and inference inside your network, no path out where that is required. Whether you should own the hardware is a separate question with an arithmetic answer, usually no on cost alone and yes on residency.
Where this vision layer meets a regulated paper process — extraction, checks and evidence chains.
Explore →The pipeline vision sits inside, and why field accuracy and document accuracy are different numbers.
Read →The rest of the practice: prediction, anomalies, speech and operations simulation.
Explore →Pick a single class of image and the errors that matter on it. We assess what is achievable on your own data, state the accuracy and the labelling volume, and price the build — inside the Blueprint, from $30,000, credited against it.
Not sure a project is the right first step? Free AI Readiness Score — 3 minutes.
Not a quote? Ask a question