Skip to main content
Gain AmericaGet in touch

AI in Insurance Underwriting & Claims: NAIC Model Bulletin Compliance Guide

How insurers deploy AI for underwriting and claims under the NAIC AI model bulletin and state DOI rules — governance, bias testing, and build-out talent.

By Gain America, Enterprise AI Advisory · Updated 2026-08-06

Insurers can deploy AI in underwriting and claims without triggering state regulators by standing up the written AIS governance program the NAIC Model Bulletin expects — documented risk management, bias and unfair-discrimination testing for any model touching rating or eligibility, human oversight of adverse decisions, and contractual accountability for third-party models — before the first system reaches production.

That sentence is the whole compliance strategy in miniature. The hard part is execution: the NAIC framework and its state-level adoptions are written as governance expectations, but they are satisfied by engineering artifacts — model inventories, testing pipelines, decision logs, and review workflows that actually exist in production. Insurance CIOs and chief actuaries who treat the bulletin as a legal memo end up with policies that describe systems nobody built. The carriers moving fastest treat it as a build specification.

What the NAIC model bulletin on AI actually requires

In December 2023 the NAIC adopted the Model Bulletin on the Use of Artificial Intelligence Systems by Insurers. It is guidance rather than a new model law, but the distinction matters less than it sounds: the bulletin tells insurers how state departments of insurance will apply existing law — unfair trade practices acts, unfair discrimination prohibitions, and market conduct examination authority — to AI-driven decisions. Adoption has been rapid. Wisconsin became the 24th state to adopt the bulletin in March 2025, and with roughly half the states now aligned, most of them with little or no modification, the bulletin is the de facto national baseline for insurance AI governance.

The core expectation is a written AIS program — an artificial intelligence systems program — proportionate to the insurer's use of AI and the risk of adverse consumer outcomes. In practice, examiners will look for:

  • Governance and accountability: board or senior-management oversight of AI use, named owners for each system, and documented approval paths for models that affect regulated insurance practices — marketing, underwriting, rating, claims, and fraud.
  • Risk management and internal controls: a complete inventory of AI systems (including vendor-supplied models), risk-tiering by consumer impact, and controls that scale with that tier.
  • Lifecycle documentation: how models were trained, on what data, with what validation, drift monitoring, and decommissioning criteria.
  • Consumer-outcome protection: processes to detect and mitigate unfair discrimination, plus meaningful human review of adverse decisions such as declinations, non-renewals, and claim denials.
  • Third-party oversight: due diligence, contractual audit rights, and ongoing monitoring for AI systems and data acquired from vendors.

None of this bans any particular use of AI. The bulletin is deliberately technology-neutral — it regulates outcomes and process, not architecture. That is good news for carriers with a disciplined build practice and bad news for anyone hoping a vendor contract would carry the compliance burden.

State rules that go further: Colorado SB 21-169 and the NY DFS circular

The NAIC bulletin is the floor. Two state regimes show where the ceiling is heading, and any multi-state carrier should design to them.

Colorado SB 21-169, enacted in 2021, prohibits insurers from using external consumer data and information sources (ECDIS), algorithms, or predictive models in ways that unfairly discriminate based on race, color, national or ethnic origin, religion, sex, sexual orientation, disability, gender identity, or gender expression. The Colorado Division of Insurance has implemented it in stages: first a governance and risk-management framework regulation for life insurers using ECDIS, then a quantitative testing regulation requiring life insurers to statistically test underwriting outcomes for unfairly discriminatory results by race and ethnicity — including estimating applicant race and ethnicity through approved inference methods where it is not collected directly — and to remediate and report what they find. Colorado is explicit about the direction of travel: the framework is expected to extend to other lines and practices over time.

Colorado's rule changed the question regulators ask. It is no longer "do you intend to discriminate?" — it is "show us the statistical evidence that your model's outcomes do not."

New York DFS Insurance Circular Letter No. 7 (2024), adopted in July 2024, applies to every insurer authorized to write business in New York that uses ECDIS or AI systems in underwriting or pricing. It requires a governance framework with board and senior-management oversight, and — critically — a documented, multi-step fairness analysis before an insurer may rely on an AI system or external data source: confirm the data and model do not use or proxy protected characteristics in a prohibited way, assess for disparate impact, and consider less discriminatory alternatives. The circular also states plainly that insurers may not rely on a vendor's assurances alone; the insurer using the model owns the analysis.

For a national carrier, the practical synthesis is straightforward: build one AIS program to the NAIC bulletin's structure, then add Colorado-grade quantitative outcome testing and New York-grade pre-deployment fairness assessment as standard gates for any model touching rating or eligibility. Meeting the strictest state by default is cheaper than maintaining fifty variants. Carriers that already run governed AI programs in banking will recognize the pattern — the supervisory logic closely parallels what we describe for FINRA- and SEC-regulated AI at banks and broker-dealers.

Underwriting copilots and claims agents: where the loss-ratio impact is

Governance exists to enable deployment, not prevent it. The use cases below are where carriers are getting measurable economics, roughly in ascending order of regulatory sensitivity.

Claims triage and FNOL agents. First notice of loss is the highest-volume, most document-heavy moment in the claims lifecycle. LLM-based agents now handle FNOL intake across channels, extract loss facts from unstructured reports, photos, and adjuster notes, classify severity, and route claims — fast-tracking clean low-severity claims toward straight-through processing while flagging complex or suspicious ones for senior adjusters. The economics come from three directions: lower loss-adjustment expense per claim, faster cycle times that reduce rental, storage, and litigation costs, and earlier identification of large-loss and fraud indicators, which reduces claims leakage. Because triage recommends routing rather than deciding coverage, it is a defensible first deployment — provided denial and reservation-of-rights decisions stay with licensed human adjusters.

Underwriting copilots. In commercial lines especially, underwriters spend a large share of their day assembling submissions: loss runs, schedules of values, broker emails, inspection reports. A retrieval-grounded copilot summarizes the submission against appetite, surfaces the carrier's own guidelines and comparable risks, drafts referral memos, and pre-fills rating inputs for human confirmation. The measurable impact is quote turnaround and submission-to-quote ratio — in a market where the first credible quote often wins the account, cutting days out of the desk time directly moves premium volume and, through better risk selection, the loss ratio. The architecture is a classic enterprise RAG build over underwriting guidelines and historical files, with citations so the underwriter can verify every claim the copilot makes.

Rating and eligibility models. Predictive models that feed rates or accept/decline decisions are the highest-sensitivity tier — this is precisely the territory Colorado and New York regulate. They belong in production only behind the full testing regime described below, with filed-rate discipline and actuarial sign-off.

Across all three, the deployment pattern that survives regulatory scrutiny is the same one that works in other regulated industries: AI drafts, retrieves, scores, and routes; humans decide anything adverse. Our guide to human-in-the-loop AI agents covers how to design review checkpoints that add control without destroying throughput, and these patterns generalize well beyond insurance to any regulated decision workflow.

Bias and unfair-discrimination testing for models touching rating and eligibility

Unfair discrimination is the oldest prohibition in insurance law, and AI did not change the rule — it changed the evidence regulators expect. A modern testing program for models in rating, eligibility, or claims settlement has four layers:

  1. Input review. Confirm no protected characteristics are used, and inventory features plausibly correlated with them — geography, credit-derived attributes, occupation, education, and third-party lifestyle data are the usual suspects.
  2. Proxy analysis. Test whether individual features or feature combinations effectively reconstruct protected class membership. A model can discriminate without ever seeing a protected attribute.
  3. Outcome testing. Compare approval rates, price distributions, claim settlement values, and SIU referral rates across demographic groups, using inference methods where protected attributes are not collected. This is what Colorado now requires quantitatively for life underwriting.
  4. Less-discriminatory-alternative search. Where disparities appear and are attributed to the model, document whether an alternative specification achieves comparable predictive power with less disparate impact — the analysis New York's circular expects before deployment.

The critical operational point: this is not a one-time filing exercise. Models drift, portfolios shift, and vendor data changes underneath you. Testing must run on a schedule, produce versioned artifacts an examiner can read, and gate redeployment. That makes it an engineering problem — the same pipeline discipline covered in our guide to evaluating AI agents in production applies directly, with fairness metrics added to the evaluation suite alongside accuracy.

A fairness test that lives in a data scientist's notebook does not exist as far as a market conduct examiner is concerned. It has to be a pipeline: scheduled, versioned, logged, and attached to the model it governs.

Third-party models: insurers own the outcomes of AI they buy

Most insurance AI is bought, not built — underwriting scores, claims severity models, fraud detection, telematics analytics. Both the NAIC bulletin and the NY DFS circular are unambiguous that purchasing a model does not transfer responsibility. If a vendor score produces unfairly discriminatory declines, the insurer that used it faces the regulator.

That converts vendor management into a technical discipline. Contracts should secure documentation of training data and methodology, rights to conduct or commission bias testing, notification of material model changes, and cooperation with regulatory inquiries. Where a vendor cannot or will not support outcome testing, the carrier needs to test around the black box — running its own applicant-level outcome analysis on the score's effects — or find another vendor. "The vendor said it was fine" is not an answer any DOI accepts.

Staffing the build: actuarial SMEs paired with contract LLM engineers

Everything above ends at the same bottleneck: who builds it. Compliant underwriting and claims AI needs two kinds of expertise that rarely live in the same person — deep insurance domain knowledge (actuarial science, underwriting authority structures, claims practice, state filing requirements) and current LLM engineering skill (retrieval architecture, agent orchestration, evaluation pipelines, observability). Carriers that try to hire the second group permanently run into the market reality we document in the enterprise AI talent gap: the engineers who have shipped governed production systems are scarce, expensive, and rarely looking for a carrier's org chart.

The staffing pattern that works is a paired-team model. The carrier supplies the domain spine — a chief actuary or appointed actuary for anything touching rates, underwriting and claims leaders as product owners, and compliance embedded from sprint one. Gain America supplies the build capacity: forward-deployed AI engineers who embed with the client team on a contract basis — LLM application engineers for copilots and agents, MLOps engineers for the testing and monitoring pipelines, and evaluation specialists who turn the fairness requirements above into running code. Because the engagement is C2C contract-based rather than permanent headcount, carriers scale the team up for the 6–12 month build and down to a maintenance posture afterward, keeping the institutional knowledge with the actuarial and compliance staff who remain. This is the same delivery model behind our broader financial services AI consulting practice, where the bar is identical: production systems that a regulator can examine and the business can trust.

The sequencing that consistently works: stand up the AIS program skeleton first — inventory, risk-tiering, and testing gates — then ship a low-sensitivity system (claims triage or an underwriting copilot) through those gates to prove the process, then graduate to rating-adjacent models once the governance muscle is real. Regulators reward carriers that can show their homework. The ones that get flagged are almost never the ones using the most AI — they are the ones that cannot explain the AI they use.

Frequently asked questions

What does the NAIC AI model bulletin require insurers to do?

The NAIC Model Bulletin on the Use of Artificial Intelligence Systems by Insurers, adopted in December 2023, expects every insurer using AI in regulated insurance practices to maintain a written AIS program covering governance, risk management, internal audit, bias mitigation, and oversight of third-party AI vendors. Roughly half of US states had adopted the bulletin with little or no modification by 2025, so it functions as the de facto national baseline.

Is the NAIC model bulletin legally binding?

The bulletin is guidance, not a new statute — but it interprets existing law. It tells insurers how state regulators will apply unfair trade practice, unfair discrimination, and market conduct statutes to AI-driven decisions. In states that have adopted it, examiners can request your AIS program documentation during market conduct exams, so the practical effect is binding.

What does Colorado SB 21-169 require for insurance AI?

Colorado SB 21-169, enacted in 2021, prohibits insurers from using external consumer data, algorithms, or predictive models in a way that unfairly discriminates based on protected characteristics. The Division of Insurance implemented it through a governance regulation for life insurers using ECDIS and a quantitative testing regulation that requires life insurers to statistically test underwriting outcomes for disparities by race and ethnicity and report results.

Are insurers responsible for AI models they buy from vendors?

Yes. Both the NAIC model bulletin and NY DFS Circular Letter No. 7 make clear that an insurer cannot outsource accountability. If a third-party underwriting score or claims model produces unfairly discriminatory or non-compliant outcomes, the insurer that used it answers to the regulator. Contracts should secure audit rights, documentation, and cooperation with regulatory inquiries.

What team do insurers need to build compliant underwriting and claims AI?

A working pattern pairs internal domain owners — actuaries, underwriters, claims leaders, compliance — with contract AI engineering talent: LLM application engineers, MLOps engineers, and evaluation specialists. The domain experts define the rules and fairness constraints; the engineers build governed pipelines, human-in-the-loop review, and audit logging. Most carriers staff the build phase with embedded contract engineers rather than permanent hires.

Build it with Gain America

Gain America staffs and deploys the engineers behind enterprise AI — from data center teams to forward deployed engineers.

Talk to our team