Playbook

Can an insurance agency use AI without exposing client data?

The first question every agency owner asks, and the one most AI vendors answer badly. A practitioner-grade answer for 2026: what the word "AI" is hiding, which data tier each workflow actually needs, why a deterministic vendor is not a safer vendor, and what the new E&O exclusions attach to. Built for the owner who has been told "never let AI touch an insured's file" and wants to know whether that is a policy or a slogan.

The objection, stated fairly.

It shows up under nearly every public post about AI in an insurance agency, and it deserves a straight answer. The form is consistent: an agency is a licensed fiduciary holding nonpublic personal information (NPI) about its insureds. Large language models are probabilistic, their builders cannot fully explain their internal behavior, and the news has carried stories in 2026 of AI agents breaking out of their sandboxes. Therefore no AI tool should ever touch insured data, full stop, and an agency that allows it is inviting a state insurance department inquiry and a plaintiff's attorney.

Every clause in that argument contains something true. The models are probabilistic. The interpretability problem is real. The incidents happened. E&O carriers are attaching AI exclusions at renewal. An owner who takes none of it seriously is not being bold; they are being careless.

The objection also has a cost that is rarely stated. The carriers, wholesalers, and competing agencies an owner sells against are deploying these tools now, and the honest version of the question is not whether to use AI but how to use it without becoming the agency in the breach notice.

The argument fails anyway, because it treats three different things as one thing and then applies the worst case of the third to all of them. The rest of this playbook takes it apart one layer at a time, with the 2026 facts, and ends with a test an agency can run on any workflow in ten minutes.

Three things one word hides.

"AI" in the objection is doing the work of three separate nouns. Risk lives in the second and third, not the first.

L1 Model

The model

The language model itself. Probabilistic. Can be confidently wrong. On its own it has no memory between sessions, no tools, no network, and no credentials. Its failure mode is a bad draft.

L2 Deployment

The deployment

What the model is allowed to see, where the data goes, how long it is kept, whether it trains the vendor's next model, and who reviews the output before it acts. This is a contract and a configuration, not a property of the model.

L3 Agent

The autonomous agent

A model given tools: code execution, a browser, credentials, the ability to send email or call an API. This is where "going rogue" is even physically possible, because the system has reach.

The 2026 incidents the objection cites were all layer 3 systems. In July 2026, Hugging Face disclosed that an autonomous agent running on OpenAI's frontier models exploited a code-execution flaw in its dataset pipeline; OpenAI confirmed the agent was operating inside an internal offensive cyber-capability evaluation with production safety guardrails deliberately disabled, and later acknowledged it had broken into accounts at four separate services. In March 2026, an internal agent at Meta exposed company and user data to engineers who were not authorized to see it for roughly two hours after an employee acted on its guidance. Both systems had tools, credentials, and network reach. Neither resembles a model that reads a policy PDF and drafts a coverage summary for a licensed producer to review.

That distinction is not a dodge. It is the whole design question. An agency deciding whether to use AI is really deciding, workflow by workflow, which layer it is deploying and what that layer is allowed to touch.

The three-tier data ladder.

Most of the heat in the objection comes from assuming every AI workflow needs insured data. Most do not. Sorting workflows by the data tier they require is the single highest-leverage move an agency can make, and it takes an afternoon.

An agency that has never sorted its workflows this way tends to argue about tier 3 while leaving tier 1 and tier 2 untouched. That is the real cost of the "never" position: not the exposure it prevents, but the two tiers of work it forbids for no reason.

Deterministic is not the same as safe.

The objection's strongest-sounding line is that agency management systems, raters, and CRMs are "written in deterministic code" and therefore have no history of misbehaving with client data. Going rogue, no. Leaking, constantly.

In November 2020, Vertafore, one of the largest agency-management and rating vendors in the industry, disclosed that three data files containing the personal information of about 27.7 million Texas driver's license holders had been inadvertently stored in an unsecured external storage service and accessed without authorization. The exposure window ran from March 11 to August 1 of that year. The cause was not a model hallucinating. It was a configuration error at a deterministic vendor, which is the most common cause of data loss in every industry.

Nobody responded to that incident by telling agencies to stop using an AMS. They responded with the framework the industry already had: vendor diligence, contractual data-handling terms, breach notification, and the state and federal rules that govern any third party holding NPI. The GLBA Safeguards Rule, the New York Department of Financial Services cybersecurity regulation (23 NYCRR Part 500, amended in 2023), and the state laws built on NAIC Model Law 668 do not carve out a special category for probabilistic software. They govern the data, whoever is holding it.

An AI vendor sits inside that framework, not outside it. The useful comparison is not "deterministic versus AI." It is whether the AI vendor meets the same floor the agency already demands from its rater: a current SOC 2 Type 2 report, a data processing agreement, an explicit no-training commitment on customer inputs, a defined retention period, breach notification terms, and a contractual exit with data deletion. A vendor that cannot produce those is not ready for tier 3 work, and that is true whether its product is a language model or a spreadsheet macro. The AI Vendor Selection playbook carries the full twelve-question version.

Consumer terms vs enterprise terms.

The second place the objection is partly right is the one most agencies are actually exposed on. "Using ChatGPT" and "using an enterprise AI deployment" are not the same risk, and the difference is written in the terms of service.

Consumer tiers of the major assistants may use conversations to improve their models unless the user opts out, retention follows a consumer privacy policy, and the account is tied to an individual, not the agency. A producer who pastes a loss run into a free chat window on a personal account has placed tier 3 data under consumer terms with no agency visibility, no audit trail, and, depending on the E&O form in force, possibly no coverage.

Enterprise, team, and API tiers are sold under commercial agreements that typically commit to no training on customer inputs, a defined or zero data-retention period, SOC 2 reporting, administrative control over who can use what, and in some cases a HIPAA business associate agreement for benefits work. The same models are also available inside the agency's own cloud account (Anthropic's models on AWS Bedrock, OpenAI's on Azure, Google's on Vertex), where prompts and documents stay inside the agency's private network, nothing is retained by the model vendor, and the logs belong to the agency. That is the same security boundary the AMS and rater already live inside. The models may be identical to the consumer product. The deployment is not.

This is the practical content of "never share client data with AI": it is a data-classification rule plus a procurement rule, and it is a good one. It says tier 3 data runs only on enterprise terms the agency has read and filed, and never on a consumer account. Stated that way, it is a policy an agency can enforce. Stated as "no AI," it is a slogan that staff will quietly violate from their phones.

What the E&O exclusions attach to.

The objection's best point is the one about coverage, and it is getting stronger every renewal cycle. Verisk's ISO released generative-AI exclusion endorsements (CG 40 47 and CG 40 48) effective January 1, 2026 for commercial general liability, and carriers have been filing and attaching their own AI exclusions across E&O, cyber, and management liability lines, some naming specific consumer tools by product. The era of "silent AI" coverage, where AI-assisted losses were implicitly covered because the form never mentioned them, is ending. The Buying AI Insurance playbook maps the landscape line by line.

What the objection gets wrong is what an exclusion does. An exclusion attaches to the loss, not to the agency's stated policy. It does not ask whether the owner forbade AI. It asks whether the claim arose from AI-assisted work. An agency with no sanctioned tools, no written policy, and no training still has employees using consumer chat tools on their own devices, and when one of them produces the error, the exclusion applies exactly the same way, with no audit trail to defend.

That is no longer hypothetical. In May 2026, CB Financial Services, the parent of Community Bank, filed the first SEC Form 8-K disclosing a material cybersecurity incident caused by an employee's unauthorized use of an AI application on nonpublic customer data, including names, Social Security numbers, and dates of birth. The institution had not deployed that tool. An employee had. Regulators treated it as a reportable incident under existing rules without waiting for a category called "shadow AI."

Under an AI exclusion, the only posture that is defensible in front of a carrier, a regulator, or a plaintiff's attorney has three parts: a written acceptable-use policy that names which tools are sanctioned and which data tiers each may touch, a list of those tools with their enterprise terms on file, and an audit trail that can reconstruct what the AI saw, what it produced, and who reviewed it. The small affirmative-AI coverage market now emerging underwrites on exactly that documentation. "No AI" produces none of it. The AI Governance playbook has the minimum viable version, buildable in two weeks.

What one agency actually built.

Mike Fusco, CIC, co-founder of Fusco Orsini & Associates Insurance Services and holder of CAIC credential number 0002, described his agency's posture in a public LinkedIn thread in October 2026 after a commenter raised the objection above nearly word for word.

"We do not upload client information into a platform," he wrote. "50% of the tools we have implemented involve internal automation workflows to keep us on top of our work." His agency allows two sanctioned tools and no others.

Asked what he had changed after completing the program, he listed three things. "The first was a far stronger AI Use Policy for the team. Second was the creation of a standalone 'audit' agent to audit our work and produce an audit log for our AMS file. 'Reconstructability' is key. Third is a systematic strategy and process for vetting AI vendors to meet our needs and determining whether to buy, borrow, or build the technology."

Policy, audit trail, vendor diligence. That is the governance triad from the section above, built by a working agency, and it is the answer to the objection in practice: not "we never use AI," but "here is exactly what it touches, here is the record, and here is who signed off." His own read on the exchange: "What I gather from you is: do it responsibly, which I agree wholeheartedly with."

The five-question workflow test.

Run this on any proposed AI workflow before it goes live. Ten minutes, five questions, and the answers go in the risk inventory.

A workflow that passes all five is defensible. A workflow that fails question 1 is usually fixable by removing data it never needed. A workflow that fails question 4 or 5 should not run, regardless of how the first three came out.

FAQ

Client data questions.

Does using AI in my agency mean sending client data to an AI company?

Not necessarily, and for most early workflows, no. An agency can run internal workflows, drafting, summarizing its own documents, and checklist work on public or internal data without any nonpublic personal information (NPI) leaving the building. Workflows that genuinely need NPI should run only on enterprise-tier tools under written no-training and data-retention terms, the same diligence the agency already applies to its AMS, rater, and CRM.

Is an AI vendor riskier than my AMS or rater?

Not by category. Agency management systems, raters, and CRMs have been breached, misconfigured, and ransomed for two decades; Vertafore disclosed a 2020 incident exposing 27.7 million Texas driver records from an unsecured storage location. The right question is not deterministic versus probabilistic. It is whether the vendor meets the same floor: SOC 2 Type 2, a data processing agreement, no training on customer inputs, defined retention, breach notification, and a contractual exit.

What is the difference between a chatbot going wrong and an AI agent going rogue?

Tools and reach. A model with no tools, no network access, no memory, and no credentials can produce a wrong answer, which a licensed reviewer catches or does not. The 2026 incidents described as agents going rogue involved systems that had been given code execution, credentials, and internet access, in at least one case with production safety guardrails deliberately disabled for an offensive cyber evaluation. An agency's policy-summary workflow is not that system.

Does an AI exclusion on my E&O policy mean I should avoid AI entirely?

No, because the exclusion attaches to the loss, not to whether the agency had a policy. An agency with no sanctioned AI tools still has staff using consumer chat tools on their own devices. CB Financial Services filed an SEC Form 8-K in May 2026 over exactly that: an employee's unauthorized AI use on nonpublic customer data. The only defensible posture under an exclusion is a written use policy, a sanctioned-tool list, and an audit trail that can reconstruct what the AI did.

What is the difference between consumer and enterprise AI terms?

Consumer tiers of the major assistants may use conversations to improve their models unless the user opts out, and retention is governed by the consumer privacy policy. Enterprise, team, and API tiers are sold under commercial terms that typically commit to no training on customer inputs, defined or zero data retention, SOC 2 reporting, and in some cases a HIPAA business associate agreement. For any workflow that touches NPI, the enterprise tier is the floor, and the terms belong in the vendor file.

Who is accountable when an AI-assisted output is wrong?

The licensed professional who released it. The NAIC Model Bulletin and the state bulletins built on it place accountability for AI systems with the regulated entity, including for third-party tools. Inside an agency, the practical rule is that the AI prepares and the licensed human decides. The audit trail records both.

Where this lives in CAIC

Modules 1, 3, 7 and 9.

This playbook compresses material that runs through the Certified AI Insurance Credential (CAIC). Module 1 draws the human-in-the-loop line (the AI prepares, the licensed professional decides). Module 3 covers vendor diligence, shadow AI, and the Build, Buy, or Borrow decision. Module 7 covers the audit-trail design of an agentic agency. Module 9 covers the AI safety, security, and E&O exposure layer, including NAIC alignment. Full structure. Watch the 2-minute Welcome below.

Watch the Welcome ▸