Status note: Version 5.0 of the Supplement was released on Aug 31, 2026, superseding draft 4.0, and it restructures the exhibits (see the changelog for the 4.0 to 5.0 changes). This guide is being reconciled to 5.0; the criteria below currently map to draft 4.0 question numbers and are under revision. Written comments on 5.0 are open through roughly Sep 30, with a verbal-comment meeting tentatively Oct 8 and version 7.0 as the adoption target at the Fall National Meeting in November 2026. When the guide moves, the changelog records what changed.
What this is
A readiness rubric for anyone building or operating AI in insurance: MGAs and program administrators running on a carrier’s paper, vendors selling into them, and the carriers who will be asked to describe both.
The NAIC Supplement is addressed to insurers. It is not addressed to you. That is the point. Exhibit B, question 4 asks the carrier to describe its oversight of AI systems used by professional service providers, and names MGAs explicitly. When a state examiner sends that question to your fronting carrier, the carrier answers it with whatever evidence you can produce. If you can’t produce it, the carrier will produce a version for you, on its terms, or it will stop asking and start restricting.
This rubric turns the Supplement’s four exhibits into five things any operator or vendor can be scored on. It is a self-assessment instrument. Nolte does not publish scores for named companies.
Who this matters for, and why
| If you are | The Supplement reaches you because | What you need from this guide |
|---|---|---|
| Carrier or fronting partner | Exhibits A to D are addressed to you. Exhibit B Q3 and Q4 ask you to describe the AI your vendors and delegated operators use. | One schema for what to ask, so every program and vendor on your book reports the same way. |
| MGA, program administrator, TPA, claims administrator | You are the “professional service provider” in Exhibit B Q4. Your carrier answers with your evidence or restricts what you can do. | The controls that make the evidence a query, and the reporting flow that sends it monthly. |
| AI or data vendor selling into insurance | Exhibit C ref 4 and Exhibit D col 5 name you. Exhibit B Q3b asks how your product was validated. A separate NAIC framework for third-party vendors is in draft. | What your buyers will ask for, and the contract schedule they will send. |
| Builder, founder, or engineering lead | The pattern that survives every regime is architectural, not procedural. Where you place probabilistic components decides your tier, your cost, and your exam. | The workflow map, the tiering scheme, and the nine controls. |
| Broker or embedded distributor | If a model touches quote, bind, or servicing on your side, you are in the chain. | The same, at the tier your systems actually occupy. |
The thread through all five: the examiner asks the carrier, the carrier asks everyone below it, and evidence that does not exist cannot flow.
What this is not
- Not a certification and not a compliance determination. The NAIC has said the Supplement itself is not one either. Nolte is not affiliated with or endorsed by the NAIC.
- Not a legal opinion. Bulletin adoption and AI-specific rules vary by state. Check the state tracker and your own counsel.
- Not a NIST or ISO crosswalk. Those exist elsewhere and they are useful. This rubric starts from the exam question, not the framework.
The five pillars
Weights reflect what an examiner reaches for first (Exhibit A), what a carrier will demand of the operators and vendors on its book (Exhibit B Q3, Q4), and where the detailed evidence burden lands (Exhibits C and D).
| # | Pillar | Weight | NAIC anchor |
|---|---|---|---|
| 1 | Inventory and classification | 20% | Exhibit A; Exhibit C ref 1 to 7 |
| 2 | Governance and accountability | 20% | Exhibit B narrative Q1, Q5; checklist 1 to 3 |
| 3 | Validation and observability | 25% | Exhibit B Q3; Exhibit C ref 8, 9 |
| 4 | Data provenance | 20% | Exhibit D; Exhibit B checklist 3d |
| 5 | Consumer outcome controls | 15% | Exhibit B checklist 3a, 3c, 3m, 3n; Exhibit C ref 12, 13 |
Each criterion scores 0, 25, 50, 75, or 100 on a full review. Pillar score is the mean of its criteria. Composite is the weighted sum. The interactive readiness check on this site uses a faster three-point version of the same rubric (Yes 100, Partly 50, No 0).
Pillar 1: Inventory and classification (20%)
Exhibit A asks for counts by operational area: models in use, models with direct consumer impact, models with material financial impact, models implemented in the past 12 months, and the use case for each. Exhibit C asks, per model, for name, version, type, implementation date, internal or third-party origin with vendor name, risk classification, known limitations, and autonomy level (automate, augment, support).
An operator that cannot answer Exhibit A in a day cannot answer anything downstream.
Criteria
- 1.1 A written inventory exists of every AI system in production, including vendor-embedded models and features inside platforms you rent.
- 1.2 Each entry is tagged to an Exhibit A operational area (the draft lists fourteen, including marketing, quotes and discounts, underwriting, ratemaking, claims, customer service, utilization management, fraud, investment, legal/compliance, producer services, reserves, catastrophe triage, reinsurance, plus other).
- 1.3 Each entry carries an autonomy level using the Supplement's definitions: support, augment, automate.
- 1.4 Each entry carries a risk classification with a written rationale. The Supplement leaves "high risk" to the company to define, so the definition itself is evidence.
- 1.5 Implementation date and version are recorded, and the inventory can produce "implemented in the past 12 months" on demand.
Common gap: the rented platform. If your PAS or rating vendor ships a model inside a feature, it is in your inventory whether or not you built it. Vendor name is a required field in Exhibit C and Exhibit D.
Pillar 2: Governance and accountability (20%)
Exhibit B (narrative) asks who maintains the framework, how it reaches the board, how it is integrated and remediated, and how effectiveness is assessed. The checklist version asks whether a written AI program exists, when it was adopted, how often it is reviewed, and whether the board or management were involved.
For a founder-led operator, “board reporting” is usually the carrier relationship plus your own leadership. Scale the ceremony, keep the substance.
Criteria
- 2.1 A written AI program exists, dated, with a review cadence.
- 2.2 A named role owns it.
- 2.3 The program addresses the checklist items 3a through 3n at least by reference: unfair trade practices, legal compliance, adverse consumer outcomes, privacy, suitability, ERM, SDLC integration, financial reporting impact, training, risk quantification, vendor procurement standards, complaint tracking, consumer disclosure.
- 2.4 AI risk is integrated into your SDLC (checklist 3h). If you build, this is where your delivery process becomes exam evidence.
- 2.5 The program has been reviewed at least once since adoption and the review is documented.
Common gap: a policy PDF written for the carrier’s onboarding questionnaire that nobody has reopened.
Pillar 3: Validation and observability (25%)
The heaviest pillar because it is the heaviest question. Exhibit B Q3 asks for validation and testing procedures on internally developed systems, on vendor-supplied systems, and the frequency, scope, and methodology of verification. Exhibit C ref 8 asks how model outputs are tested for drift, accuracy, unfair trade practices, unfair discrimination, and performance degradation, how the model was validated before deployment, and how it is monitored on an ongoing basis. Ref 9 asks for the last test date.
“Ongoing basis” is the operative phrase. A validation report from launch is a snapshot. An examiner asking for the last test date wants a system that produces one.
Criteria
- 3.1 Pre-deployment validation is documented per model, with the reference data source stated (internal, external, or both, per the Supplement's "Validation Method" definition).
- 3.2 Production monitoring exists per high-risk model: at minimum accuracy or outcome-rate tracking, drift detection, and a defined threshold that triggers review.
- 3.3 Every decision produced by an automate-level system is logged with inputs, output, model version, and timestamp, so any single outcome can be reconstructed.
- 3.4 Vendor-supplied models have a written validation procedure that does not depend solely on the vendor's own attestation.
- 3.5 A "last date of model testing" can be produced for every high-risk model without asking an engineer to go look.
Common gap: telemetry exists in the application but was never framed as governance evidence. This is usually a reporting problem, not a build problem.
Pillar 4: Data provenance (20%)
Exhibit D lists 25 data element types, from aerial imagery through telematics to non-traditional data, and asks for each: the AI system type using it, how it is used across operations, whether it is internally sourced, and if third-party, the vendor name. Checklist 3d covers privacy.
Criteria
- 4.1 Every data element feeding an AI model is mapped to an Exhibit D category.
- 4.2 Source is recorded per element: internal, or third-party with vendor named.
- 4.3 Training and test data lineage is documented for internally developed models.
- 4.4 Sensitive categories (age, gender, ethnicity, medical, criminal, income, geo-demographics) are either absent, or their use has a written justification and a proxy-discrimination check.
- 4.5 The data map is versioned and updated when a model or a vendor changes.
Common gap: geocoding and geo-demographics. Almost every rating engine uses them. Almost no one has written down that they do.
Pillar 5: Consumer outcome controls (15%)
Checklist 3a, 3c, 3m, and 3n: residual risk of unfair trade practices, evaluation of adverse consumer outcomes, complaint identification and tracking, and consumer disclosure. Exhibit C ref 12 asks how each model is reviewed for compliance with unfair trade practices and unfair claims settlement laws. Ref 13 asks whether any regulatory action has been taken against the company for use of the model.
Criteria
- 5.1 For each consumer-impacting model, a human review path exists for adverse decisions (declination, non-renewal, adverse rating, claim denial), and its trigger conditions are written.
- 5.2 Complaints are tagged to the AI system involved, if any, and the tag is queryable.
- 5.3 Consumer-facing disclosure of AI use exists where the state requires it, and the state tracker is the source of truth for where.
- 5.4 Each model has a documented compliance review against applicable unfair trade practices and claims settlement rules.
- 5.5 A regulatory-action log exists per model, even if empty.
Scoring bands
| Composite | Band | Read |
|---|---|---|
| 85 to 100 | Exam-ready | Your carrier can answer Exhibit B Q4 with your evidence, unedited. |
| 65 to 84 | Defensible | Gaps are documentation, not controls. Fixable in 30 days. |
| 40 to 64 | Exposed | At least one pillar has no system behind it. The carrier will notice before the examiner does. |
| Below 40 | Undocumented | You are running AI on someone else’s paper with nothing to show. |
Readiness check
The front door is the interactive readiness check: one question per criterion across the five pillars, scored on a three-point fast-track (Yes, Partly, No), with a live score and the control that closes each gap. It is the rubric above made answerable in a few minutes, not a separate scheme.
For a thirty-second gut check before you open it, ask:
- Can you list every AI system in production, including ones inside platforms you rent, today?
- Does each one have a written risk classification and a reason for it?
- Does a named person own your AI program, and has it been reviewed since it was written?
- Is AI risk a step in your software delivery process, not a separate document?
- For your highest-risk model, can you state the last date it was tested?
- Is that model monitored in production for drift and outcome rates, with a threshold that triggers review?
- Can you reconstruct any single automated decision from logs: inputs, output, model version, time?
- Is every data element feeding your models mapped to a source, with vendors named?
- Is there a human review path for adverse consumer decisions, with written triggers?
- If your fronting carrier forwarded Exhibit B question 4 to you tomorrow, could you answer it in a week?
If more than a couple are “no”, the 30-day readiness review is built for exactly this.
Versioning
- v0.1 (2026-08-24): initial draft against Supplement draft 4.0.
- v0.2 (2026-09-01): distinguished the full-review scoring scale from the interactive check’s three-point fast-track, and pointed the readiness-check section at the interactive tool. Kaio Pedreira review.
- Next revision: on release of Supplement draft 5.0, and to reconcile the Working Group’s late-August 2026 session.