AI agents for commercial insurance: what works, and what it takes

Back to Insights

The right mix of AI and human labor can lower cost, speed production, and raise quality on insurance work. All three at once is a vendor-shaped claim. Here is what has to be true for it to hold.

Cheaper, faster, better is exactly the shape of a claim you should not accept from anyone selling you something. Any two of the three is a real result. All three, asserted without a mechanism behind it, is a pitch.

So here is the mechanism, and at the end I grade each of the three separately.

My vantage point is narrow and worth stating plainly. As part of an engagement with a frontier AI lab, I build the commercial insurance tasks that frontier models are tested against and evaluate what they produce. That does not make me a model researcher. It does mean I have watched current systems attempt real insurance work at volume against a standard, rather than watching a demo. The general version of this argument is in the right mix of AI and human labor; this is the insurance version.

Why insurance is a harder target

Submissions arrive unstructured, as an email with nine attachments, some scanned. Coverage language is legal language, where one exclusion changes the answer and jurisdiction changes it again. And the cost of being confidently wrong is not a bad survey score, it is a coverage decision.

That last point shapes everything. These models fail in a particular way. They produce fluent, well-formed, wrong answers, and they do not reliably know when they have done it. Stanford's RegLab found general-purpose models hallucinating on 69 to 88 percent of specific, verifiable legal questions, and found the models often failed to correct a user's incorrect assumption.

That tested general models without retrieval, so it is not a fair test of a purpose-built system. The follow-on study is the useful one. Purpose-built, retrieval-backed legal research tools from LexisNexis and Thomson Reuters hallucinated between 17 and 33 percent of the time, against 43 percent for GPT-4 on the same queries.

Read together, those two results give you the design principle. Grounding the model in the right documents cuts the error rate substantially. It does not reach zero. So the question is never whether to trust the output. It is where you put the human so the remaining error lands somewhere survivable.

Where that lands, task by task

FunctionWhat an agent does wellWhere the human staysPrimary gain
Submission intake and triageExtract, normalize, check appetite, route, flag gapsException queue onlyCycle time
Underwriting supportRisk summaries, exposure extraction, referral drafts, prior similar risksThe decision, above a set complexity or premium thresholdCapacity
Claims and FNOLFacts of loss capture, severity triage, documentation, fraud pattern detectionThe call itself, and coverage determinationCycle time and detection
Policy servicing and auditCertificates, endorsements, renewal prep, audit exceptionsSampling, plus full review on anything filed or regulatedVolume per head

The pattern in that table is not about which tasks are easy. It is about how fast a mistake surfaces. A mis-extracted field on an intake shows up at the next step and costs minutes to fix. A wrong coverage opinion shows up at claim time and costs a great deal more. Put agents where errors surface fast, keep people where they surface late.

Intake is where I would start almost every time. Accenture's longitudinal survey with The Institutes found the average underwriter spending 70 percent of their time on non-underwriting activities, 40 percent of it administrative. Intake and rekeying is a large share of that. Moving it does not cut headcount so much as it returns capacity to the work that prices risk.

Claims deserves the most caution. The intake mechanics automate well and fraud detection is a genuine machine advantage, since pattern recognition across a whole book is something no human can do at that scale. However the call from someone who has just had a loss is tone and judgment, not text generation.

What separates a real build from a wrapper

Grounding. The agent answers from the actual policy, submission, and guidelines, not from what the model absorbed in training. That is the difference between the 43 percent error rate and the 17 percent one.

Scope. Narrow agents against one document set beat general agents pointed at a whole function. Every expansion widens the failure surface.

Checkpoints. Decide where a human sees the output based on the cost of being wrong, not on a target automation rate.

Evaluation. Real cases with known correct answers, scored against a standard you set before you saw the output, run again as the models and your own documents change. If a vendor cannot show you their evaluation set and what it scores on your kind of work, they have not measured it. Which means neither have you.

Who captures the gain

Much of this work already sits with providers rather than carriers, which creates an owner problem. If a BPO deploys agents against work it bills by the seat or the transaction, the provider keeps the gain until the next renegotiation.

Three questions for any provider pitching an AI-enabled insurance operation. Show me the evaluation, what it scores and how often you run it. Show me where the human checkpoints are and who staffs them. Show me how pricing changes as automation rises, because if it does not, the savings are not coming to me.

For a provider, the honest version of this is a strong position rather than a threat. Demonstrated quality on automated work, priced to share the gain, is worth more than a lower rate. Cost is never the wage rate alone, and it is not the automation rate alone either. I unpack the full stack in cost is never the wage rate alone.

Grading the claim

Faster is the surest of the three. Cycle time on intake, triage, servicing, and documentation comes down materially, and it comes down first.

Better is achievable, however through coverage rather than replacement. Scoring every file instead of a sample, catching exceptions a human queue would miss, giving a newer adjuster the context a veteran has. The gain comes from the machine seeing everything, not deciding everything.

Cheaper takes the most discipline and is the easiest to fake. It is real when freed capacity gets redeployed and the exception queue is staffed well enough that errors do not flow downstream. It disappears the moment you measure automation rate instead of resolved work. Deloitte's 2026 outlook notes that 90 percent of insurance executives agree on the urgency of preparing people for human and machine collaboration while only 25 percent have acted on it. That gap is where these programs quietly stall.

There is no perfect answer here; it is a mix, designed for the specific book, the specific documents, and the specific cost of being wrong. If you are being pitched an AI-enabled insurance operation and want an honest read on whether the claims hold up, it is worth at least having the conversation.

Sources

Figures cited above, with the caveat that survey dates and definitions vary by source. Treat them as directional.

  • Underwriter time allocation: Accenture and The Institutes P&C Underwriting Survey, conducted 2021, reported by Michael Reilly, Accenture Insurance Blog, February 2022. insuranceblog.accenture.com
  • Legal hallucination rates in general-purpose models: Dahl, Magesh, Suzgun, and Ho, "Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models," Journal of Legal Analysis, January 2024. reglab.stanford.edu
  • Hallucination rates in retrieval-backed legal tools: Magesh et al., "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools," Journal of Empirical Legal Studies, April 2025. onlinelibrary.wiley.com
  • Workforce readiness figures: Deloitte, 2026 Global Insurance Outlook, Deloitte Center for Financial Services, October 9, 2025. deloitte.com

Worth at least having the conversation.

Tell us what you are trying to fix, scale, or evaluate. We will give you an honest read on whether we can help.

Contact us