
A life insurer ran its products on a legacy policy administration system, and every product change meant changing that system. Leadership wanted to know whether AI agents could carry one product from application through the first policy anniversary, and whether they could do it accurately enough to stand comparison with the system they might replace. The insurer wanted evidence before it committed to a modernization path.
The proof of concept runs on Amazon Bedrock, which the insurer chose after comparing projected run costs for this workload across platforms. Our engineers built the plumbing that makes agents testable: batch processing that can rerun without duplicating a transaction, a test harness that runs the same applications through the agents and the legacy system and compares the results, a local AWS emulation so the team can test without touching live services, and a generator that produces ACORD TXLife 103 new business messages from the e-application's XML schema.
Product managers mapped every question on the application to its regulatory basis, including NAIC Model 275 on annuity suitability, and recorded why each question exists. A review of one of the insurer's application packages, 169 questions across 11 forms, found that 55 asked for a fact already collected on another form, 72 could be prefilled from data the insurer already held instead of typed by the selling agent, 29 needed a business justification to stay, and four could be cut outright. Those findings went into the functional requirements, so the AI agents will process the application the insurer should have instead of the one it inherited.
The proof of concept is scoped at eight to twelve weeks for one product. Engineers and product managers work from the same requirements crosswalk, so every agent behavior traces to a requirement and every requirement traces to a regulation or a business rule. The functional requirements are complete, and the technical specification for the agent build is being written from them.
The insurer is making its modernization decision on evidence. It has a platform choice grounded in a cost comparison, a harness that will show whether agents match the legacy system on identical applications, and an application that is shorter and better justified before any agent touches it. The requirements and the redesigned application carry forward whichever way the comparison comes out.
Client details in this case study are generalized, and in places combined across engagements, to protect confidentiality. This work sits at the center of our policy administration system and AI implementation practices, and our article on agentic AI in insurance covers where agents fit next to the core system. We are glad to walk through comparable work under NDA.