Document capture, one method per document type

Most of this market sells one method and fits your documents to it. We count your types, measure how much each varies, and pick per type — often a template that costs almost nothing to run.

from $30,000Solution Blueprint: routing table, fixed price and timeline, in 5–6 weeks
from $150,000typical build — first set of document types, 4–6 months
Fixed pricequoted against the Blueprint scope — no time and materials
On-prem or cloudopen-weight models inside your perimeter, GPUs in the same contract

The matrix we choose by

One axis does most of the work, and it is not the one the market talks about. Not how modern the method is — how much the form varies.

Form variabilityWhat it looks likeMethodWhy not the others
FixedA form you designed. A statutory return. A standard claim form. The layout is the same every time.Template extractionNo training data, near-zero running cost, identical output forever, and it fails loudly when the layout changes rather than quietly guessing. A model here is worse on every axis, including accuracy.
Varies within known limitsInvoices from two hundred suppliers. Bills of lading. Certificates from a known set of issuers.Trainable extractionToo much variation to template, enough regularity to learn. A vision model also works and costs more per page for accuracy you did not need.
ArbitraryDeal packs. Medical records. Mixed correspondence with attachments — structure you learn only by reading.Vision-language modelNothing else copes. Templates have no anchor and trainable extraction has no stable feature. You pay in cost, latency and reproducibility — knowingly.

Two further axes shape what gets built around the method rather than the method itself. Cost of error sets the confidence thresholds, the cross-field validation and how much goes to a human. Reproducibility decides the cases you will have to explain: a template returns the same answer permanently, a hosted model may not be able to tell you in a year what it read today.

That is a wider set of processes than the word “compliance” suggests. The useful question is not whether you are regulated but who will ask you to justify an answer — an auditor, a counterparty disputing an invoice, a losing bidder, your own management, or a court. A regulation is one of those and not the most frequent.

Almost every real document estate spans all three rows. The output of our first phase is therefore a routing table, not a single method — which is exactly the answer a single-product vendor cannot give you. The full matrix, with the field-error arithmetic →

When we tell you not to do this

We say no to capture projects more often than this market does. If any of these describes you, the honest answer is that it does not pay back, and we would rather say it now.

Volume below the thresholdUnder a few thousand documents a year the labelling and build effort will not pay back, whatever the model quality. We do the arithmetic with you before quoting — documents times minutes saved has to beat build plus run plus the cost of being wrong.
Nothing is digitised yetIf the documents are still in filing cabinets, the first project is scanning and indexing, not extraction. That is a real project with a real supplier, and it is not us.
Nobody owns the processFiles arrive by email, portal and post, and no single person is accountable for them. Automation has nothing stable to attach to. Fix ownership first.
You control the formIf the document is your own and it is badly designed, ten minutes of form design beats a training set. Add a barcode, fix the field order, make the boxes real boxes. We will point this out even though it removes most of the project.

How the work runs

Four phases, each with what we need from you. The first one is the product — everything after it is priced against what the first one found.

PHASE 01
Estate audit and routing tableWe count the document types, measure how much each varies, sample the worst files rather than the best, and assign a method per type. Output is a routing table, a build-ready specification, a fixed price and a timeline. This phase is a product you buy on its own — the AI Solution Blueprint, from $30,000, credited against the build. The spec is yours; build it with us, in-house or elsewhere.You provide
  • 50–100 real documents per major type, worst cases included
  • Annual volumes per type, however rough
  • What a wrong field costs, field by field
PHASE 02
Method per type, on real dataTemplates for the fixed types, trainable extraction where there is enough regularity, a vision model for the residue. Measured at field level, weighted by correction cost, against a held-out set from your production history.You provide
  • A few hundred examples per trainable type
  • One named owner for the documents
  • The target schema, or an hour to define it with us
PHASE 03
Validation, routing and the exception queueCross-field checks, confidence thresholds set against your error costs, and a review queue that a person can actually work. Plus the record of what was read, by which version, and who confirmed it.You provide
  • Validation rules that already exist in policy
  • A named owner for the exception queue
  • Retention and residency requirements
PHASE 04
Integration and handoverWiring into the ERP, DMS or database with reprocessing and monitoring; deployment into your environment; acceptance against the criteria set in phase 01.You provide
  • Integration contacts and a test environment
  • An acceptance owner who can sign
  • A decision on where it runs — your cloud or on-prem

The same method under four rulebooks

What makes the matrix trustworthy is that the underlying mechanic — classify the page, find the fields, check completeness, flag what a human must see — recurs unchanged under very different regulation. Client names withheld under NDA. See full case studies →

Aviation · MRO

Checking a maintenance package against a regulation

An autonomous computer-vision service that classifies each page of an MRO documentation pack into six document types, detects missing signatures and stamps, finds unfilled checklist cells and verifies date-and-signature pairing on work orders, then returns a JSON report and annotated pages — before a specialist signs off. Fixed forms, so templates and detectors, not a language model. Read the case →

6 document types classified4 classes of check on one packCLI + API delivery
Immigration tech · arbitrary documents

Drafting from inconsistent document sets

The other end of the matrix: large, unpredictable document sets read by a retrieval-augmented pipeline that drafts structured memoranda with citations back to source. No two packs alike, so no template could have worked. Read the case →

80% of routine drafting automated−45% processing time+30% throughput
Lending · identity documents

Reading identity documents at scale

Document authenticity checks, face match and liveness on identity documents — a known set of layouts from a known set of issuers, which is the middle row of the matrix and the reason trainable extraction fits it. Read the case →

−75% fraudulent applications+35% conversion

The fourth rulebook is GMP: the same completeness-and-signature mechanic applied to pharmaceutical batch records, covered in AI batch record review.

No surprises

Your data

The whole pipeline can run on-premises or air-gapped on open-weight models, with no document leaving your network, and the GPU hardware quoted in the same contract. Details in security and compliance and private AI infrastructure.

Your answers

Wherever an answer may be questioned later — by an auditor, a supplier, a bidder or your own management — we pin model versions and prefer deterministic templates, so the same document produces the same output in a year, after your model has been retrained twice and a hosted API has moved on without telling you.

Your budget

Fixed price against a routing table agreed before the build starts. No time and materials, no discovery that bills indefinitely. We do not quote a range before that table exists, because a number produced before anyone has counted your document types is a guess.

Frequently asked questions

How is a data capture project scoped and priced?

Scope first, price second, and both numbers are published. The specification is a product: the AI Solution Blueprint, one system, 5–6 weeks, from $30,000, credited in full against the build if implementation starts within 90 days. It counts your document types, measures how much each one varies, prices the errors and returns a routing table, a fixed price and a timeline. Capture builds of this kind typically start around $150,000 over 4–6 months for a first set of document types.

What accuracy can we expect?

The honest answer is that it depends on the field, not on the document, and that anyone quoting a single number is quoting character-level accuracy rather than field-level. A field is wrong if any character in it is wrong, so at a 2% character error rate a thirty-character line is wrong about 45% of the time. We measure at field level, weighted by what a correction costs, on a sample drawn from your worst real documents rather than a clean set.

Do we need a vision model for this?

Often not, and we will say so. If the layout is fixed — a form you designed, a statutory return — a coordinate template is cheaper, faster, exact and reproducible, and a model is worse on every axis including accuracy. Vision models earn their cost on documents whose structure you only learn by reading them. Most estates need a mix, which is why the output of the first phase is a routing table rather than a single method.

How much labelled data do you need from us?

None for fixed-form document types, because templates do not learn. For trainable extraction, realistically a few hundred examples per document type, and we would rather label them from your production history than have you prepare a clean set. For the vision-model path, a labelled evaluation set rather than a training set — a few dozen documents per type, chosen to include the bad ones.

Can it run entirely on our own infrastructure?

Yes. The whole pipeline can run on-premises or air-gapped with open-weight models, and we can quote the GPU hardware in the same contract. It also buys reproducibility wherever an answer may be questioned later: you pin the model version, so the same document produces the same answer in a year's time, which a hosted API cannot promise.

Our documents are handwritten. Will this work?

Partly, and the distinction matters. Detecting that a signature box is filled, a stamp is present or a checklist cell is empty is reliable, and that is often the whole requirement. Reading arbitrary handwritten prose into structured fields is not, and we design those paths with human review as the default rather than as an exception.

Related practices

Tell us what your documents look like

An engineer replies, not an account manager. You get back a routing table and a fixed price.

Want the spec first? AI Solution Blueprint — from $30,000, credited against the build.

We reply within one business day. Prefer email? sales@haink.org