Haink KnowledgeCase StudiesAbout Contact sales
Home / Knowledge / Software & AI / How AI Claims Processing Works

Software & AI · Insurance operations · Written and maintained by Haink’s AI adoption team · Updated August 2026 · 17 min read

How AI claims processing works — and which parts pay for themselves

Every guide to AI in claims lists the same fifteen use cases: intake and triage, image processing, fraud detection, predictive severity, automated settlement, telematics, subrogation, photo analysis, chatbots, litigation prediction, compliance monitoring, lifecycle tracking, settlement optimisation, trend analytics, leakage detection. The list is accurate. It is also useless, because a list of fifteen things AI can do to a claim is not an answer to the only question an operations director actually has, which is where do I start and what will it cost me if it is wrong.

This page answers that instead. Three things make it different from the standard treatment: the fifteen use cases collapse into three groups once you sort them by the cost of an error; several of them are better done by a rule than by a model, and we say which; and there is a regulatory question underneath claims automation that the rest of this market does not mention at all.

The regulatory one is worth stating up front, because it is counter-intuitive. A claim contains two different decisions. Whether the claim is genuine is a fraud question. Whether the claim is owed is a merit question. They have different error costs, different evidence, and — as it turns out — different positions relative to the EU AI Act. Systems that fuse them into a single score make both problems harder to defend.

Where the time actually goes in a claim

Before automating anything, it is worth being precise about which stage is consuming the hours, because it is almost never the one people assume.

StageWhat consumes the timeNature of the work
Notification (FNOL)Re-keying from email, portal, phone notes and post into one recordMechanical
Intake and classificationWorking out what type of claim this is and which pack appliesMechanical
CompletenessFinding what is missing, chasing it, waiting, checking againMechanical, and where the calendar time goes
VerificationIdentity, policy in force, cover on the date, duplicate submissionsMechanical, plus a fraud judgment
AssessmentApplying policy wording to facts; valuing the lossJudgment, partly arithmetic
SettlementDeciding, communicating, payingJudgment

Four of the six rows are assembly. In most claims operations the assessment itself takes a competent handler minutes; the claim sits in a queue for days because something is missing and nobody noticed until a human opened the file. That gap between handling time and elapsed time is where automation earns its money, and it has nothing to do with automating a decision.

Fifteen use cases, three groups

Sort the same list by what an error costs and it stops being a list.

GroupWhat is in itCost of an errorVerdict
AssemblyIntake, classification, extraction from documents and photographs, completeness checks, cross-document consistency, duplicate detection, policy-in-force verificationA correction. Caught by the next check or by the handler.Clears the threshold almost always. Start here.
RoutingTriage and prioritisation, straight-through eligibility, complexity scoring, assignment to the right handler, exception queuesA delay, and occasionally a claim that sat with the wrong person.Worth it once assembly is stable and the criteria are written down.
DecisionMerit assessment, quantum, settlement approval, declineThe payment, a complaint, an ombudsman referral, or litigation.Usually the last thing to automate, and often not at all.

The arithmetic behind that verdict is not complicated, and almost nobody writes it down.

Cost of manual handling = claims × hours per claim × loaded hourly cost Cost of automation = build + run + (error rate × cost per error × claims)

The second term is where claims differs from ordinary document automation. In invoice processing the cost per error is uniform and small. In claims it is neither. Wrongly paying and wrongly declining are not symmetric. Overpayment costs the amount. A wrongful decline costs the amount plus the complaint handling plus, in some markets, a regulatory finding and a remediation exercise across every similar claim you decided the same way — which is why the error you should be optimising against is the one with the legal tail, not the one with the bigger unit cost.

Apply the formula group by group and the answer for most insurers is a split one. Assembly wins easily: volume is high, errors are cheap and self-correcting, and the work is genuinely mechanical. The decision layer frequently does not win — not because a model cannot assess a claim, but because once you price the wrongful-decline tail properly, most of the remaining value is speed, and speed is available more cheaply by automating everything up to the decision and handing a complete, verified file to a human who then takes minutes instead of days.

The policy is the reference standard — and it splits in two

Every claim is measured against a document you already have: the policy. That is what makes claims different from most document automation, where the reference standard has to be invented. Here it is written down, signed and binding.

But it does not all read the same way. Some clauses are checked literally against the text; others require the circumstances to be characterised before the text can be applied at all. The line between those two is the automation boundary in claims, and it runs through the policy rather than around it.

What is being checkedHowExample
Deductible, limit, sub-limitMechanical — arithmetic against the scheduleIs the claimed amount above the excess and below the limit?
Cover in force, waiting periodMechanical — date comparisonWas the policy live on the loss date, and had the waiting period elapsed?
Territorial scope, scheduled itemsMechanical — lookup against a named listIs this address, this vehicle, this item on the schedule?
Exclusion by named, stated causeMechanical, provided the cause is stated and unambiguousPolicy excludes frost damage; the report says frost.
— the automation boundary —
Exclusion requiring the cause to be characterisedInterpretation — the fact must be classified before the clause appliesFlood or a burst pipe? Wear and tear or accidental damage?
Standards of conduct and condition precedentsInterpretation, and frequently contestedDid the insured take “reasonable care”? Was notification “as soon as practicable”?
Causation between covered peril and lossInterpretation — often the whole disputeDid the covered event cause this part of the loss, or something else did?

Everything above the line is a rule applied to a schedule. It needs no training data, produces the same answer permanently, fails loudly when a field is missing, and can be shown to a complainant as a calculation. Automate it completely and you have removed a large share of the handling time without touching a single judgement.

Everything below the line is a machine classifying facts, not applying a policy — and that is a different act with different accountability. A system that decides “this was wear and tear” has not read the policy more carefully; it has made the contested call and hidden it inside an extraction step. That is the failure mode worth designing against, because it looks like accuracy right up until someone disputes it.

A practical consequence for scoping. Ask which clauses in your most common policy wordings sit above the line and which sit below. In most personal and small commercial lines the majority of decisions turn on clauses above it — which is why the assembly-and-arithmetic layer pays back and the judgement layer usually does not. If your book is weighted the other way, that is a signal to automate less, not to buy a better model.

Models still earn their place, and above the line as much as below it: reading a scanned repair estimate so the arithmetic has numbers to work with, matching photographs to a described loss, spotting that the date on the invoice contradicts the date on the report, and ranking claims by likelihood of needing attention. The general framework for when a deterministic rule beats a learned one — and the arithmetic behind it — is set out in document data capture methods, where that question is the subject rather than a digression.

The regulatory question nobody in this market answers

Search the competing guides to AI in claims and you will find no mention of the EU AI Act at all. That is a gap, but the reason it is a gap is not the one you would guess — the answer is not that claims automation is high-risk. It is that the answer is genuinely unobvious, and being unable to state it is itself a problem.

Annex III lists what counts as high-risk. Read point 5 closely:

ProvisionWhat it actually coversRelevance to claims
Annex III, 5(c)“Risk assessment and pricing in relation to natural persons in the case of life and health insurance”This is underwriting and pricing. A settlement decision is neither.
Annex III, 5(a)Systems used by or on behalf of public authorities to evaluate eligibility for essential public benefits, and to “grant, reduce, revoke, or reclaim” themThe legislator did describe deciding a claim — but only for public bodies.
Annex III, 5(b)Creditworthiness of natural persons, excluding systems for detecting financial fraudNot claims. Relevant if your group also lends.
Claims handling by a private insurer — not listed anywhere in Annex III.

Notice what that comparison shows. The drafters knew perfectly well how to write “an AI system that grants, reduces or revokes a claim”, because they wrote exactly that in 5(a) — and confined it to public authorities. Private claims handling was not overlooked; it was left off.

Not listed is not the same as not your problem. Under Article 6 a provider who considers an Annex III-adjacent system not to be high-risk must document that assessment before placing it on the market, and Article 80 gives authorities a procedure for challenging exactly that self-classification. The practical standard is simple: if you cannot write two defensible pages explaining why your claims system is not high-risk, it probably is. That document is cheap to produce while the system is being designed and expensive to reconstruct afterwards.

Three ways a claims system becomes high-risk by accident

All three are architectural, and all three are decided long before anyone asks the compliance question.

  1. You are acting on behalf of a public authority. Insurers administering statutory health schemes, state-backed benefits or public reimbursement are inside 5(a) on its own terms, whatever their corporate form. This is the single most commonly missed route into the high-risk regime in insurance.
  2. Your claims model feeds pricing. The moment claims outcomes drive individual risk re-rating or renewal pricing in life or health cover, you have reached 5(c) through the back door. A feedback loop that looked like good data science is now a classification event.
  3. One model does everything. If a single score covers fraud likelihood, merit and future risk, no component can be classified separately, so the strictest classification applies to all of it — and to the document stack feeding it. This is precisely the fusion problem that also decides the answer in AI underwriting under the EU AI Act, where the fraud exclusion survives only for systems that do fraud and nothing else.

And one rule applies today, with no ambiguity and no deadline to wait for: GDPR Article 22 governs decisions based solely on automated processing that produce legal or similarly significant effects. An automated decline or reduction of a claim is squarely within it, and has been since 2018. For most insurers this bites well before any Annex III question does.

For the deadline picture and the full Annex III breakdown — including how Regulation (EU) 2026/1744 moved high-risk obligations to 2 December 2027 — see the underwriting analysis, which covers the same instrument from the pricing side.

The same question outside the EU

The Annex III analysis above assumes a European insurer. Outside the EU the classification question mostly disappears — no comparable regime lists claims handling as high-risk — and is replaced by a narrower one that bites sooner: is this an automated decision about an individual, and can you show your working? For a claims stack that is the more consequential question anyway.

WhereInstrumentStatusWhat it asks of a claims stack
DIFC, DubaiRegulation 10 of the DIFC Data Protection Law — personal data processed through autonomous and semi-autonomous systemsIn force since 1 September 2023, full enforcement from 1 January 2026A system deciding about an individual without human involvement in each decision is a High Risk Processing Activity: assess and document before processing, tell the claimant an autonomous system is involved, and give them enough to object.
SingaporeMAS FEAT principles and the MAS Guidelines on AI Risk Management (consultation closed 31 January 2026)Supervisory expectations, tested at inspectionInsurance models explainable enough for meaningful challenge and customer recourse; documented fairness assessment.
Hong KongHKMA circulars on big data analytics and AI (2019) and on generative AI in customer-facing use (November 2024)In forceGovernance and explainability proportionate to risk; consumer-protection expectations wherever a model touches the customer.
Mainland ChinaPIPL, Article 24In force since 2021Transparency and fairness in automated decisions, no unreasonable differential treatment on terms, and a right to an explanation and to refuse a decision made solely by automated means.
Saudi ArabiaPDPL with the SDAIA AI adoption frameworkIn forceHuman review checkpoints for automated decisions in high-risk contexts.

Two things follow. First, the Gulf is ahead of Europe here, not behind it. An insurer inside the DIFC has been under full enforcement since January 2026, while the EU pushed its high-risk obligations to December 2027. Treating the region as the light-touch option is working from a map two years out of date.

Second, an automated decline is the trigger everywhere. None of these regimes cares much about intake, extraction or completeness checking — the group where the money is. All of them care the moment a system decides an outcome for a person without a human in the loop. Which is the same architectural line the EU analysis arrives at from a different direction, and the reason the recommendation does not change with the jurisdiction: automate the assembly, keep the decision reachable by a human, and keep the evidence.

A reference architecture, with the boundary marked

Six stages. What matters is not the stack but which stages can be separated from which, because separability is what preserves your classification and your ability to change one part without revalidating the rest.

The evidence chain deserves emphasis because it is cheap on day one and painful later. If a complaint arrives eight months after settlement, you need to reconstruct what the system saw, which model version read it, what it extracted and who reviewed the result. Systems that store the outcome and discard the evidence cannot do this, and no amount of retrospective documentation fixes it. Where the data cannot leave your perimeter, that chain has to live inside it — see cloud versus private AI and security and compliance.

Why the pilot does not reach production

The failure mode in claims is specific enough to name.

The pilot ran on a prepared subset. Someone assembled a few hundred complete, legible, well-formed claims. The model did well. Production then delivered photographs taken at night, PDFs that are scans of faxes, schedules that were never attached, the same claim submitted twice through two channels, and handwriting. The gap is not model quality; it is that the pilot never met the input distribution.

Nobody owned the data. If notifications arrive across email, a portal, a broker feed and the post, and no single person is accountable for the record, automation has nothing stable to attach to. Fix ownership before building — this is the most common cause of death for these projects, and it is the general pattern in why AI pilots fail.

The straight-through criteria were written after the model. If “straight-through” means “whatever the model was confident about”, you have no policy, you have a threshold. Write the criteria first, in the language of the claims policy, and let the model report against them.

The real constraint was the claims policy. If the organisation does not agree on how a class of claim should be treated, automation will produce contested decisions faster and at scale. That is a governance problem wearing a technology costume — the general case is in when not to adopt AI.

Frequently asked questions

Is AI claims handling high-risk under the EU AI Act?

Claims handling is not listed in Annex III. Point 5(c) covers risk assessment and pricing in life and health insurance, which is the underwriting side, not settlement. Point 5(a) does cover deciding to grant, reduce, revoke or reclaim benefits, but only where the system is used by or on behalf of public authorities. So for a private insurer the usual answer is not high-risk — and under Article 6 that is a classification you have to document and be able to defend, not an assumption you get to make quietly.

How can a claims system end up high-risk by accident?

Three ways. You administer a statutory or public scheme, which puts you inside 5(a) as acting on behalf of a public authority. Your claims model feeds back into individual risk re-rating or renewal pricing in life or health cover, which is 5(c) reached through the back door. Or one model scores fraud, merit and pricing together, so no component can be classified separately from the others.

Which parts of claims processing are worth automating first?

The assembly, not the judgment. Intake, classification, completeness checks, extraction, cross-document consistency and duplicate detection are high-volume work where an error costs a correction. That group almost always clears the payback threshold. The settlement decision usually does not, because the cost of a wrong decision is high enough that most of the value of automation is speed, and speed can be bought more cheaply upstream of the decision.

When is a deterministic rule better than a model in claims?

Whenever the criterion is written down. Deductibles, limits, sub-limits, exclusions, waiting periods and policy wording are arithmetic and text lookup, not inference. A model that learns them from historical settlements will approximate them, drift from them, and have to be explained when it disagrees with the policy. Reserve models for what is genuinely learned: reading messy documents, resolving inconsistencies across sources, and ranking claims by likelihood of needing attention.

We underwrite in the Gulf or Asia, not the EU. Does any of this apply?

The classification question mostly disappears — no comparable regime lists claims handling as high-risk — and is replaced by a narrower one that bites sooner. DIFC Regulation 10 has been in full enforcement since 1 January 2026 and treats a system that decides about an individual without human involvement as a High Risk Processing Activity. Singapore's MAS expects insurance models to be explainable enough for meaningful challenge. China's PIPL Article 24 gives a right to refuse a decision made solely by automated means. The common trigger everywhere is an automated decline, not the document work upstream of it.

Does GDPR apply to automated claims decisions?

Yes, and it applies now regardless of AI Act classification. Article 22 governs decisions based solely on automated processing that produce legal or similarly significant effects, and an automated decision to decline or reduce a claim is squarely within that. This is the rule that bites first for most insurers, well before any Annex III question.

Why do claims automation pilots fail to reach production?

Because the pilot runs on a prepared subset. Someone assembles a clean set of complete claims, the model performs well, and then production delivers photographs taken at night, PDFs that are scans of faxes, missing schedules and duplicate submissions. If intake is fragmented across email, portals and post with no single data owner, fix that before building anything — it is the most common cause of death for these projects.

This page explains what the regulation says and what follows from it for system design. It is not legal advice, and it does not replace an assessment by your own counsel or compliance function.

Related Resources

Knowing which parts to automate is the easy half

Building assembly and routing as separable services, with an evidence chain that survives a complaint eight months later, is the other half. That is the build, and it has its own page.

See how we build it →   Free AI Readiness Score

Sources. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Annex III points 5(a)–(c) and Articles 6 and 80 · Regulation (EU) 2026/1744 (Digital Omnibus on AI) · Regulation (EU) 2016/679 (GDPR), Article 22.

Regulatory status verified: August 2026. Next review: November 2026.

Haink
info@haink.org

Winning House
72–76 Wing Lok Street
Sheung Wan, Hong Kong

© 2026 Haink. All rights reserved.  ·  Privacy Policy  ·  TermsHong Kong · Dubai · Singapore · Mainland China · Delaware (USA)