The Real Difference Between a Claims Automation Pilot and a Claims Automation System
Key Takeaways
-
A claims automation pilot proves a model can process a claim under controlled conditions. A claims automation system proves it can process the claim it has never seen before, escalate correctly when it isn't sure, and leave a record a regulator would accept. These are different engineering problems, not different degrees of the same one.
-
Gartner expects more than 40% of agentic AI projects to be cancelled before the end of 2027 - not primarily because the models underperform, but because of escalating costs, unclear business value and inadequate risk controls (Gartner, 2025).
-
A December 2025 McKinsey survey found 88% of organisations now use AI somewhere in the business, but only 7% had scaled it fully across the enterprise - the gap between "we use it" and "we run on it" is the gap this article is about (McKinsey & Company, 2025).
-
Regulators in the US, EU, Australia and the global standard-setting body for insurance supervisors have converged on the same three non-negotiables for AI-assisted claims decisions: explainability, human escalation and documented audit trails - none of which show up in a pilot's accuracy score (NAIC, 2023; EIOPA, 2025; IAIS, 2025; APRA, 2026).
-
Not every claim needs the same weight of governance. Low-value, high-volume, low-dispute claim types can run well on lighter controls (Gaca, 2026). The exception proves the rule rather than replacing it - and knowing which is which is itself a governance decision.

Introduction
At some point in the last eighteen months, a Head of Claims told their board that the organisation had "gone live" with AI-powered claims automation. The numbers looked good: a meaningful share of claims touched by the model, cycle times down, a straight-through-processing rate worth putting in the annual report.
Then a claim came in that didn't look like the ones in the demo. Maybe it was a coverage question the model had never been configured to recognise. Maybe a customer disputed a decision and an ombudsman asked for the reasoning behind it. Maybe an auditor asked a simpler question than anyone expected: show me the record of why this specific claim was declined, who reviewed it, and what version of the model made the call.
That is the moment a pilot and a system reveal they were never the same thing.
This is not a piece about whether claims automation works. The evidence that it can reduce cycle time and cost is well established, and source[code] has written about the broader economics of AI in claims and underwriting elsewhere. This piece is narrower and, for a Head of Claims sitting in front of a board or a regulator, more urgent: how do you know whether what you have built is a demo that performs well, or an operating system that can be trusted?
The core thesis is simple to state and hard to act on: a pilot proves a model can process a claim. A production system proves it can process the claim it has never seen before, escalate correctly when it is unsure, and produce an audit trail a regulator would accept. The features that make that difference are almost entirely unglamorous - exception handling, escalation paths, audit logging - and almost never the features a vendor leads with in a sales demo.
The evidence: why "it worked in the pilot" keeps not surviving contact with production

Three independent bodies of evidence point at the same gap, from three different angles.
First, scale. In its December 2025 global AI survey, McKinsey found that 88% of organisations reported using AI somewhere in the business - up ten percentage points on the prior year - while only 7% of respondents said AI had been fully scaled across their organisation (McKinsey & Company, 2025). That is not a claims-specific figure, but it describes the exact shape of the problem this article addresses: adoption is nearly universal, operational maturity is rare, and the distance between the two is where most "AI-powered" claims initiatives are quietly living.
Second, project mortality. Gartner's June 2025 forecast projects that more than 40% of agentic AI projects will be cancelled before the end of 2027. Crucially, the reasons Gartner cites are not model accuracy: "escalating costs, unclear business value and inadequate risk controls" (Gartner, 2025). Gartner analyst Anushree Verma put it plainly: most agentic AI projects today "are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied." The models are not the bottleneck. The surrounding operating environment is.
Third, and most directly relevant to Australian and APAC carriers: regulators are watching the exact seam this article is about. In its April 2026 letter to industry, the Australian Prudential Regulation Authority observed that regulated entities are "moving beyond experimentation from internal productivity use cases to customer facing applications," while governance "has not matured at the same pace" (APRA, 2026). That is a regulator telling the industry, in writing, that the pilot-to-production gap is not a private engineering concern - it is a supervisory one.
None of these findings are about whether the AI is accurate on average. They are about whether the system around it can be trusted when it is wrong, uncertain, or facing something it has not seen before - the entire subject of this article.
The Head of Claims' actual decision problem
Boards rarely ask a Head of Claims "is the model accurate?" They ask "are we live?" - a binary question that a well-run pilot can answer "yes" to long before the organisation has built anything resembling a production system.
Three structural pressures make that binary framing dangerous:
KPI reporting measures the easy path, not the hard one. Straight-through-processing rate, cycle-time reduction and customer satisfaction scores all describe what happens when the model is confident and right. None of them describe what happens when it is confident and wrong, or unsure and left to guess. A dashboard can show excellent numbers while the exception path - the part that determines whether the system is trustworthy - remains genuinely undeveloped.
Vendor demonstrations are performed on curated conditions. A proof of concept run against a clean, historical claims dataset says very little about performance against next quarter's fraud pattern, a new product line, a catastrophe surge, or a regulatory change mid-cycle. IAIS supervisory guidance makes the relevant distinction explicit: AI systems should be "designed to fail safely or escalate to human intervention" precisely when they encounter conditions outside their design envelope (IAIS, 2025, ¶73). A demo, by construction, rarely tests that condition.
Regulators examine a different question than the one procurement usually asks. Procurement processes tend to score vendors on demo accuracy and unit cost. Regulators ask whether the insurer can "meaningfully explain the outcomes" of the AI systems it uses (IAIS, 2025, ¶69), whether it has kept "appropriate records of the training and testing data and the modelling methodologies" (EIOPA, 2025, §3.23), and whether it has "escalation procedures, involving relevant staff" (EIOPA, 2025, §3.30). None of that is visible in a pilot's headline accuracy metric - and all of it is what a Head of Claims will actually be asked to produce if a claim decision is contested.
The real decision problem, then, is not "does the model work?" It is: has anyone actually tested and documented what happens when it doesn't?
What most institutions get wrong
Across the claims automation initiatives that reach this stage, the same pattern recurs, and it is worth naming precisely because it is not a failure of ambition or budget.
They optimise the acceptance path and treat the exception path as an afterthought. Accenture's February 2026 insurance predictions put this directly: the fix is to "redesign roles so humans are a control point, not a formality - clear approval thresholds, exception handling, audit trails, and escalation paths for high-impact decisions" (Accenture, 2026). Most claims automation programmes do the opposite by default: the acceptance path gets engineered first because it is what a demo needs to show, and the exception path is left as "someone will review it," which is not a design.
They treat human review as scaffolding to be automated away, rather than a permanent architectural control. This is a subtle but consequential error. A pilot's success criteria often implicitly assume the human-in-the-loop step will shrink over time as the model improves. But the regulatory expectation is the opposite: human oversight is a structural requirement of the system, not a temporary crutch for an immature model (EIOPA, 2025, §3.29-3.30; APRA, 2026).
Procurement scores the wrong thing. Vendor selection processes that weight demo accuracy heavily and audit-log ownership or escalation design lightly are, in effect, buying a pilot and hoping it becomes a system later. As Adam Gaca of Future Processing put it, reflecting on UK claims transformation programmes: "UK insurers rarely lack AI ideas. What is usually missing is a baseline, a named business owner and a clear definition of what better looks like" (Gaca, 2026).
They retrofit governance after scale-up, which is politically and technically harder than building it in. Once a model is processing a meaningful share of claim volume, adding confidence calibration, structured escalation and complete event logging means changing a system that is already relied upon operationally - under scrutiny, usually after an incident has already made the case for you.
The Production Floor Standard: three checkpoints that separate a demo from an operating system
source[code] uses a simple mental model with clients assessing where a claims automation initiative actually sits: the difference between a demo stage and a production floor. A demo stage is built to perform well under a script. A production floor has to handle whatever walks onto it, including the things nobody scripted for.
We call the assessment the Production Floor Standard - three checkpoints, each with its own maturity tiers, that test the parts of a claims automation initiative a pilot's accuracy score cannot show you.

Checkpoint 1 - The Unseen-Claim Checkpoint (generalisation)
Does the system handle claim types, combinations and edge cases it was not explicitly trained or configured for - not just statistical noise inside a familiar distribution, but genuinely novel patterns: a new product line, an unusual coverage interaction, a fraud pattern that didn't exist at training time?
Checkpoint 2 - The Escalation Checkpoint (calibrated uncertainty)
Does the system know when it doesn't know? Low-confidence or high-impact cases should route to a named human reviewer, with the context needed to act, inside a defined service window - not fall through to a silent default of auto-approve, auto-deny, or an undefined queue.
Checkpoint 3 - The Audit Checkpoint (evidentiary completeness)
Would the record of one specific claim decision - inputs, model version, confidence score, rationale, and any human override - survive a regulator's file review or an ombudsman complaint, on demand, without a special data-archaeology exercise to reconstruct it?

Most initiatives that describe themselves as "live" sit at Tier 1 on all three checkpoints. That is a perfectly reasonable place to be mid-programme. It is not a production system, and the gap matters most exactly when it is tested - during a dispute, an audit or a catastrophe surge, not during a steering committee update.
Pilot vs. production system, side by side

Business and technology implications
The Production Floor Standard is not only a governance exercise - it changes real decisions.
Vendor contracts need different clauses. Audit-log ownership, export rights, model-version retention and liability for AI-influenced decisions belong in the contract, not a post-incident negotiation. A vendor that cannot commit to exportable, per-decision audit records is a Tier 0-1 tool regardless of demo accuracy.
Organisational design has to treat escalation review as a permanent claims function, not a role that shrinks as the model "gets better." Regulators expect this as a standing control, not a temporary bridge (EIOPA, 2025, §3.29-3.31).
The technology stack needs components a pilot rarely includes: an orchestration layer separating deterministic business rules from model inference, a confidence-calibration mechanism the model can be measured against, and event logging built for a decision's full lifecycle - not an application log for debugging.
The cost curve favours building this in early. Retrofitting exception handling, escalation and audit logging onto a system already carrying real claim volume means re-engineering a production dependency under commercial and regulatory pressure at once - harder than designing it into the initial build.
A fair counterargument: not every claim needs the full standard
It would overstate the case to argue every claim requires Tier 3 governance before it can be trusted. That is not what the evidence shows, and source[code]'s own client conversations bear this out.
For low-value, high-volume, low-dispute claim types - small motor glass claims, or parametric products like flight-delay payouts - straight-through processing with lighter governance genuinely works. Gaca (2026) draws exactly this distinction: low-value, high-volume claims are well suited to heavy automation because the cost of an individual error is small, bounded and reversible, while specialty or high-value claims should keep AI in a decision-support role rather than a decision-making one.
Regulators effectively endorse this proportionality themselves. EIOPA's opinion explicitly allows that "stronger guardrails and increased human oversight" can substitute for full explainability where the underlying complexity is low (EIOPA, 2025, §3.27) - governance calibrated to stakes, not a uniform standard applied regardless of claim value.
This does not weaken the Production Floor Standard; it clarifies how to apply it. The three checkpoints should be applied risk-proportionately - rigorously for claim types with genuine dispute potential, regulatory exposure or customer harm, and more lightly for claim types where the cost of being wrong is genuinely small. The governance failure is not automating simple claims lightly. It is applying that same light-touch thinking to complex, high-value, or dispute-prone claims because the demo made it look easy.
What leaders should do next
1. Audit your current initiative against the three checkpoints, not its headline accuracy or STP metric. Where does it actually sit - Tier 0, 1, 2 or 3 - on generalisation, escalation and audit completeness?
2. Test on adversarial and out-of-distribution claims, not demo sets. If nobody can show how the system behaves on a claim type it wasn't built for, you don't yet know how it behaves in production.
3. Instrument audit logging before scaling volume, not after. Retrofitting is the expensive path, usually undertaken under scrutiny rather than on your own timeline.
4. Name an accountable owner for the escalation function, with an SLA and board visibility - not an informal "the team will handle it."
5. Segment claim types by required governance tier rather than one uniform standard. Low-value, low-dispute claims can run lighter; complex or high-value claims cannot.
6. Change what you report to the board. Replace "% of claims automated" with "% correctly escalated" and "audit completeness rate" - the figures that actually predict whether the system survives a regulator's or ombudsman's scrutiny.
The source[code] perspective
Across the claims transformation work source[code] has supported for banks, insurers and insurtechs in APAC, Australia and the Gulf, the pattern above is close to universal: the model is rarely the reason an initiative stalls between pilot and production. The reason is that exception handling, escalation design and audit logging are unglamorous, they don't demo well, and they are consistently the last things engineered rather than the first.
We don't think that is a criticism of the teams building these systems - it's a predictable consequence of how pilots get funded and demonstrated. A pilot is judged on whether it can show a working claim being processed. A system has to be judged on whether it can be trusted with the claim nobody planned for. Those are different design briefs, and treating them as the same one is where most claims automation initiatives quietly stop being progress and start being risk.
If your organisation has a claims automation initiative that has been called "live" for more than two quarters, the Production Floor Standard above is a reasonable place to start an honest internal assessment - before a regulator, an ombudsman, or a hard claim does it for you. We'd welcome the conversation. Talk to us here!
Conclusion
The gap between a claims automation pilot and a claims automation system is not a matter of degree - more data, more training, a slightly higher accuracy score. It is a different engineering and governance problem: can the system handle what it hasn't seen, know when to ask for help, and prove what it did. Every regulator now converging on this question - the NAIC, EIOPA, APRA, and the IAIS speaking for supervisors globally - is asking Heads of Claims to answer it with evidence, not a go-live announcement.
The organisations that treat exception handling, escalation and audit logging as core engineering from day one will be the ones still describing their claims automation as "live" when someone official asks them to prove it.
Frequently Asked Questions
What is the real difference between a claims automation pilot and a production system? A pilot demonstrates that a model can process a claim correctly under controlled or curated conditions. A production system demonstrates that it can handle a claim it has never seen before, escalate to a human when its confidence is low, and produce a complete, auditable record of the decision. The second requires engineering - exception handling, escalation routing, audit logging - that a pilot's accuracy score does not test.
Why do so many claims AI pilots fail to reach production? Gartner projects that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls as the primary reasons - not model inaccuracy (Gartner, 2025). Most claims pilots are demonstrated, funded and judged on accuracy alone, which leaves the operational and governance infrastructure needed to scale undeveloped.
What does a regulator actually want to see in an AI claims audit trail? Based on published guidance from the IAIS, EIOPA, NAIC and APRA, regulators expect: a meaningful explanation of how a specific decision was reached, evidence of human oversight and defined escalation procedures, and retained records covering the data, model version, and methodology behind the decision - producible on demand, not reconstructed after the fact.
Do all claim types need the same level of governance? No. Low-value, high-volume, low-dispute claim types - such as small motor glass claims or parametric payouts - can run well with lighter governance, because the cost of an individual error is small and reversible. Higher-value, complex or dispute-prone claims warrant the full standard. Applying uniform, minimal governance across all claim types regardless of stakes is the actual risk.
How long does it typically take to move from pilot to production-grade claims automation? There is no universal timeline, and vendors who quote one without knowing your claim mix, data quality and existing governance maturity should be treated cautiously. What is consistent across the evidence is that organisations which design exception handling, escalation and audit logging in from the start reach production-grade maturity faster than those that retrofit these capabilities after scaling volume on an accuracy-only pilot.
What should a Head of Claims ask a vendor before buying claims automation software? Beyond accuracy on a test set: How does the system behave on claim types outside its training distribution? What confidence threshold triggers escalation, and to whom? Is the full per-decision audit record - inputs, model version, rationale, overrides - exportable on demand? Who owns liability for an AI-influenced decision that is later disputed? A vendor unable to answer these clearly is selling a pilot.
Reference List
Accenture (2026) '5 predictions for the insurance industry in 2026', Accenture Insurance Blog, 2 February. Available at: https://insuranceblog.accenture.com/5-insurance-predictions-2026 (Accessed: 18 September 2026).
APRA (Australian Prudential Regulation Authority) (2026) Letter to industry: Artificial Intelligence (AI), 30 April. Available at: https://www.apra.gov.au/news-and-publications/apra-letter-industry-artificial-intelligence-ai (Accessed: 18 September 2026).
EIOPA (European Insurance and Occupational Pensions Authority) (2025) Opinion on Artificial Intelligence Governance and Risk Management, EIOPA-BoS-25-360, 6 August. Available at: https://www.eiopa.europa.eu/publications/opinion-artificial-intelligence-governance-and-risk-management_en (Accessed: 18 September 2026).
Gaca, A. (2026) 'Why a successful claims AI pilot can still fail in production', Insurance Edge, 20 July. Available at: https://insurance-edge.net/2026/07/20/why-a-successful-claims-ai-pilot-can-still-fail-in-production/ (Accessed: 18 September 2026).
Gartner (2025) Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, press release, 25 June. Available at: https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 (Accessed: 18 September 2026).
IAIS (International Association of Insurance Supervisors) (2025) Application Paper on the Supervision of Artificial Intelligence, July. Available at: https://www.iais.org/uploads/2025/07/Application-Paper-on-the-supervision-of-artificial-intelligence.pdf (Accessed: 18 September 2026).
McKinsey & Company (2025) 'AI at work but not at scale', McKinsey Week in Charts, 10 December. Available at: https://www.mckinsey.com/featured-insights/week-in-charts/ai-at-work-but-not-at-scale (Accessed: 18 September 2026).
NAIC (National Association of Insurance Commissioners) (2023) Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, December. Summary available at: https://content.naic.org/insurance-topics/artificial-intelligence (Accessed: 18 September 2026).