Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Software & AI / How to Choose an AI Development Company (Checklist & Red Flags)

Knowledge / Software & AI

Software & AI · Vendor selection · Written and maintained by Haink’s AI adoption team · Updated August 2026 · 8 min read

How to Choose an AI Development Company

The AI services market is full of teams that can build a demo. Far fewer can ship a system that is accurate, maintainable and trustworthy in production. The single best predictor is a production track record with measured outcomes — followed by evaluation discipline, honesty about AI's limits, senior engineers doing the actual work, and deployment flexibility. Use the checklist below to tell them apart.

Key takeaways

Written by a vendor, so read it that way. Haink sells AI development, which makes this page an interested document. The scoring sheet below is the one we would want a buyer to apply to us, including the criteria we would score badly on with some clients — small team, no 24/7 managed service, and we decline work we think should not be built. If a supplier hands you a selection guide with no disclosure of that kind, that is itself a data point.

The scoring sheet

Most vendor checklists average everything into one number, which quietly lets a good score on culture cancel out a failure on production experience. Split it instead.

Four gating criteria — pass or fail

A miss here ends the conversation regardless of how well the rest scores. These are not trade-offs.

GateWhat passing looks likeWhy it cannot be traded
Has shipped this class of systemAt least one comparable system in live production that they still operate or supported through handoverDemo-to-production is where most of the difficulty is. A vendor who has not crossed it does not know what they do not know
Can describe how they measure qualityA concrete answer about evaluation sets, what they score and how regressions are caught — not “we test thoroughly”Without measurement, every later conversation about quality is opinion against opinion
Will name who does the workNamed people, their seniority, and how much of their time you getA senior pitch with junior delivery is the most common failure in this market and the hardest to detect afterwards
Can meet your deployment constraintA real answer on residency, on-premises or air-gapped operation if you need it — before contract, not duringA constraint discovered in month three is a restart, not a change request

Then score the rest, 0–3 each

Criterion03
Depth of production track recordOne pilotSeveral comparable systems, still running, with named outcomes
Evaluation discipline“We test it”Evaluation set built early, scored on every change, metrics reported separately
Understanding of your domainGeneric AI talkAsks about your exceptions, your edge cases and who signs
Willingness to say noEvery idea is a great ideaHas told you which part not to build, and why
Integration experienceStandalone tools onlyHas shipped into systems like yours, with the authentication and audit story
Cost transparency“It depends”Published prices or a fixed-price specification phase, and a run-cost model
Handover and operabilityOnly they can run itDocumentation, runbooks and a path for your team to take it over
References you can actually callLogos on a slideA customer who will take a call and answer awkward questions
24 points available. 20+ proceed 16–19 proceed, but buy a paid specification phase before the build <16 keep looking — and if two shortlisted vendors score alike, buy a small paid piece from each rather than deciding on paper Any gating failure ends it at any score.

How to test the numbers they quote

This industry runs on statistics nobody sources. The test is one question, asked politely, and it separates vendors faster than any capability matrix: “which case study is that number in, and who measured it?”

What you hearWhat to askWhat a good answer sounds like
“85% of AI projects never reach production”Where does that figure come from?Honestly, it is folklore — a 2017 remark about big data projects. Here is our own ratio instead
“We cut manual review by 50%”Which engagement, measured how, over what period?A named case with the measurement method — or an admission that the number is indicative
“99% accuracy”Per field or per document? On which documents?Field-level, and here is what that implies per document at your field count
“It will pay back in six months”Show me the arithmetic with your assumptions visibleA model you can change the inputs of, not a conclusion
The one that should end a meeting

A vendor who quotes a number, cannot say where it came from, and does not appear troubled by that. It is not the wrong statistic that matters — everyone has repeated one. It is the absence of the instinct to check, because the same instinct is what catches a model quietly getting worse in production eight months from now.

What to look for

CriterionWhy it matters
Production track recordShipped systems with measured outcomes prove they can finish, not just prototype
Evaluation disciplineMeasuring accuracy and catching regressions is what makes AI trustworthy
Honesty about AI's limitsTelling you when not to use AI signals judgment over hype
Senior engineersAvoids junior staffing pyramids behind a senior pitch
Deployment flexibilityCloud or on-premises, with a real answer on data privacy
Total cost of ownershipThey plan for run cost and infrastructure, not just the build

Questions worth asking

Red flags

Why the infrastructure question matters

AI systems run on hardware, and sizing it wrong is expensive in both directions — over-provisioned GPUs idle, under-provisioned ones throttle the product. The useful test is not whether a vendor sells hardware; it is whether they can answer “what will this cost to run at our volume?” with arithmetic rather than a shrug, and whether they know that the answer is usually to rent rather than buy.

A vendor pushing you toward owned infrastructure without asking about your sustained throughput is selling hardware, not sizing a workload — the break-even against a hosted API sits near 340 tokens per second maintained around the clock, and most workloads are nowhere near it. Ask them what their number is and how they got it; the working is in on-premises versus cloud LLM deployment.

Disclosure, again: Haink supplies GPU hardware as well as software, so we have an obvious interest in that question. It is also why the page above tells you the answer is usually to rent.

Fixed scope vs iterative delivery

Be wary of large fixed-scope contracts for AI work, where the right design only becomes clear once you see real data and real usage. Strong partners favor a short discovery phase and iterative delivery — a narrow first use case that reaches working results in weeks — so you steer the roadmap with evidence and can stop or pivot cheaply.

Building AI software on your own infrastructure?

Model, pipeline and GPUs under one contract — tell us the use case and we'll scope it.

Talk to our engineers   Prefer email? sales@haink.org

What to read next

Related Resources

Frequently Asked Questions

How do I choose an AI development company?

Look for a production track record with measurable outcomes, evaluation discipline, honesty about when not to use AI, senior engineers doing the work, deployment flexibility (cloud or on-premises), and attention to total cost of ownership — not just the build.

What questions should I ask an AI development partner?

Ask for production systems with measured results and references, how they evaluate accuracy and catch regressions, when they would advise against AI, whether it can run privately, exactly who will do the work, and what it costs to run at your volume.

What are red flags in an AI vendor?

Promising to train a custom foundation model when retrieval would suffice, no plan to measure quality, never pushing back on weak use cases, ignoring inference cost, and a senior pitch with junior delivery.

Why does infrastructure matter when choosing an AI partner?

AI runs on hardware; wrong-sized infrastructure is costly. A partner who quotes right-sized GPUs with the software gives one accountable vendor and run cost matched to real throughput, and makes on-premises deployment a real option.

How should we score AI vendors?

Split gating criteria from scored ones rather than averaging everything. Four things are pass or fail: they have shipped this class of system into production, they can describe concretely how they measure quality and catch regressions, they will name who does the work and how senior they are, and they can meet your deployment constraint before contract rather than during. Score the remaining eight criteria zero to three each — track record depth, evaluation discipline, domain understanding, willingness to say no, integration experience, cost transparency, handover, and callable references. Above twenty of twenty-four, proceed; sixteen to nineteen, buy a paid specification phase first; below sixteen, keep looking. Any gating failure ends it at any score.

How do I check the statistics an AI vendor quotes?

Ask which case study the number is in and who measured it. Most of the figures in this market do not survive that question: the widely repeated claim that 85% of AI projects never reach production traces to a 2017 remark about big data projects, not to research on models. When you hear an accuracy figure, ask whether it is per field or per document, because those differ enormously — 99% per field is about 82% per document at twenty fields. What matters is less the wrong statistic, which everyone has repeated, than whether the vendor shows any instinct to check.

Should AI projects be fixed-scope or iterative?

Iterative is usually better for AI, because the right design emerges once you see real data and usage. Favor a short discovery phase and a narrow first use case over a large fixed-scope contract.

Sources and scope. This page is written by a supplier of AI development and hardware, and should be read as an interested document; the scoring sheet is the one we would want applied to us. Thresholds are judgement rather than research — the point of the instrument is that gating criteria are not averaged away, not that twenty is a magic number. The provenance of the 85% figure is discussed in MLOps: getting models to production; the field-versus-document accuracy arithmetic is in how AI document processing works.

Reviewed: August 2026.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsHong Kong · Dubai · Singapore · Mainland China · Delaware (USA)