I’ve reviewed six failed AI initiatives in insurance where the root cause was the same: no maturity model. Teams jumped straight to model deployment without first defining what “good” looks like. The result was predictable—misaligned stakeholders, inconsistent KPIs, and budgets that evaporated when executives couldn’t see ROI. After watching this pattern repeat, I built a lightweight maturity model specifically for insurance. It starts with a 4-level framework, uses a 24-question assessment, and is designed to be run in under 90 days with one FTE. Below is the exact playbook I’ve reused across four property & casualty insurers, one life & health carrier, and two MGAs. If you’re responsible for shaping AI strategy, this is the tactical guide you need.
Why a maturity model beats random pilots Insurance AI initiatives fail in three predictable stages:
Stage 1 (0-6 months): One-off pilots (e.g., “Let’s try generative AI on claims narratives”). Stage 2 (6-18 months): Tactical tooling (e.g., OCR for FNOL, RPA for bordereaux).
Stage 3 (18+ months): Re-platforming when executives realize they have 17 different AI tools and no way to measure them.
- A maturity model prevents this by forcing two things: a shared definition of “AI” and a shared vocabulary for progress. In practice, that means answering “What does Level 3 look like for underwriting?” before anyone writes a single line of Python. The model I use has four levels, which map to the insurance value chain:
- Level Scope
- Primary KPI Typical insurance use cases
Budget range 1. Ad-hoc
| Single process Cost reduction (FTE hours) | RPA for bordereaux, OCR for FNOL, chatbot for policy questions $20k–$50k | 2. Repeatable Departmental | Cycle time Automated triage rules, NLP for subrogation, predictive underwriting scoring | $150k–$450k 3. Managed |
|---|---|---|---|---|
| Cross-functional Loss ratio impact | Dynamic pricing models, claims leakage detection, fraud graph analytics $600k–$1.8M | 4. Optimizing Enterprise | Combined ratio AI-driven reinsurance buying, autonomous underwriting, real-time risk pricing | $3M+ |
| Notice the KPI progression: FTE hours → cycle time → loss ratio → combined ratio. This mirrors the insurance P&L. If you’re a CFO, you care about combined ratio. If you’re a claims manager, you care about cycle time. The model makes those priorities explicit. | Step 1: Select your assessor and assemble the team | You need one dedicated assessor—ideally a senior actuary, data product manager, or head of analytics. I’ve seen this role fail when it’s given to an external consultant who doesn’t understand underwriting, or to a junior analyst who can’t push back on inflated claims. The assessor’s job is to collect evidence, not to sugarcoat results. | Assemble a 5-person steering committee: Head of Claims (owns FNOL, triage, leakage) | Head of Underwriting (owns pricing, risk selection) CTO or Head of Engineering (owns data, tooling, integration) |
| CFO or FP&A lead (owns ROI, budget) Compliance Counsel (owns governance, model risk) | Each committee member must commit to one 90-minute session per week for 12 weeks. If they can’t, the model will stall. I’ve seen this happen at a $5B P&C carrier where the CFO delegated to a manager, who delegated to an analyst. The assessment stalled until the CFO showed up to the fifth session. | Step 2: Choose your framework and tailor it to insurance Two frameworks dominate: CMMI and ISO/IEC 33000. CMMI is more prescriptive; ISO is more flexible. For insurance, I combine both: | Use CMMI’s staged representation (Levels 1-5). Replace CMMI’s generic goals with insurance-specific practices (e.g., “Validate model risk” instead of “Institutionalize a defined process”). | Use ISO 33004 for the assessment method (documented, reproducible, consistent). Here’s the tailored model I use: |
| CMMI Level Insurance Practice | Evidence Required Time Estimate | 1. Ad-hoc AI projects exist but are uncoordinated | List of AI tools in use, no central registry 1 week | 2. Repeatable AI projects follow a documented process |
Playbook exists, training log, KPI dashboard 3 weeks
3. Managed AI projects are measured and controlled
Model risk framework, model inventory, audit trail 4 weeks
4. Optimizing AI projects drive measurable business outcomes
- Loss ratio impact, combined ratio delta, reinsurance savings 4 weeks
- The assessor scores each practice as “Not achieved,” “Partially achieved,” “Largely achieved,” or “Fully achieved.” The scores map to CMMI’s 0-100 scale. I’ve found that setting a target of 70+ for “Largely achieved” is realistic for most insurers. Step 3: Run the 24-question assessment
- The assessment is split into four tracks: Data & Infrastructure, Model Development, Operations & Governance, and Business Impact. Each track has six questions. The assessor must collect documentary evidence for each answer. Below is the exact question set I reuse:
- Data & Infrastructure Do you have a centralized data catalog that includes AI training datasets?
- Can you trace any AI output to its source dataset? Do you automate data quality checks for AI inputs?
Is your data infrastructure scalable to petabyte-scale AI models? Do you track data lineage for regulatory audits?
Is your data infrastructure PCI-DSS and SOC 2 compliant? Model Development
Do you use a standardized model development playbook? Do you validate model assumptions against historical claims?
- Do you bake model explainability into the development process? Do you test models against edge cases (e.g., catastrophe scenarios)?
- Do you document model limitations and failure modes? Do you version-control your model code?
- Operations & Governance Do you have a model risk framework approved by compliance?
Do you audit AI models annually? Do you run red-team exercises on AI systems?
| Do you have a documented incident response plan for AI failures? Do you train employees on AI ethics and bias? | Do you track AI-related incidents in a centralized log? Business Impact | Do AI projects have measurable KPIs tied to the P&L? Do you calculate ROI for AI projects before deployment? | Do you run A/B tests to validate AI improvements? Do you align AI roadmaps to underwriting and claims strategy? |
|---|---|---|---|
| Do you benchmark your AI maturity against peers? Do you publish AI results internally (e.g., loss ratio impact)? | The assessor scores each answer 0-4, then aggregates to a total score. The table below shows the mapping from total score to CMMI level: Total Score | CMMI Level Interpretation | 0-24 1. Ad-hoc |
| AI is a side project with no coordination 25-49 | 2. Repeatable AI projects follow a documented process but lack central coordination | 50-74 3. Managed | AI projects are measured, controlled, and centrally coordinated 75-100 |
| 4. Optimizing AI drives measurable business outcomes and is continuously improved | I once ran this assessment at a $2B regional P&C carrier. The total score was 22. The steering committee was shocked—until we dug into the evidence. The claims team had a chatbot (Level 1), underwriting had a pricing model (Level 2), but neither team shared data, metrics, or KPIs. The CFO immediately funded a centralized data lake. | Step 4: Build the evidence pack For each question, the assessor must collect three things: | A document (PDF, slide deck, wiki page) A screenshot or log entry proving the claim |
| A witness statement from a stakeholder (e.g., “I run the daily data quality check”) I use a simple Google Drive folder structure: | The folder must be read-only for the assessor and read-write for the steering committee. Any claim without evidence is automatically scored as “Not achieved.” I’ve seen assessors inflate scores by inventing evidence—don’t do it. If you can’t find the evidence, score it as “Not achieved.” | Step 5: Score the model and identify gaps | The assessor runs a 90-minute scoring session with the steering committee. Each member scores independently, then the group discusses outliers. The goal is to reach consensus on the final score for each practice. If consensus isn’t possible, the assessor has final say. |
I once had a CFO argue that a model risk framework was “Fully achieved” because “we have a policy.” The assessor pushed back—until the compliance counsel produced the actual framework document, which was three years old and hadn’t been updated. The score dropped to “Partially achieved.”
The output is a gap analysis: a list of practices that need improvement to reach the target level. The steering committee then prioritizes these gaps based on business impact and effort. Below is a template I reuse: Gap
Current Score Target Score
Business Impact Effort (FTE-weeks)
- Dependencies Centralized data catalog
- 0 4
- High (loss ratio impact) 8
- Data engineering team Model risk framework
- 2 4
- High (regulatory risk) 12
Legal, compliance AI KPI dashboard
- 1 4
- Medium (stakeholder alignment) 4
- Data team
- The steering committee then builds a 12-week roadmap. I’ve seen this roadmap fail when it’s treated as a wish list. To avoid this, each gap must have a single owner, a clear success criterion, and a budget. If any gap lacks these, it’s dropped from the roadmap.
- Step 6: Build the 90-day execution plan The plan must be concrete, measurable, and resourced. Below is a template I reuse for a $5B P&C carrier targeting Level 3:
- Week Action
Owner Success Criterion
- Budget 1-2
- Hire data catalog vendor (Collibra, Alation, or open-source) CTO
- Data catalog live with 80% of AI training datasets cataloged $50k
- 3-4 Build model inventory spreadsheet
- Model risk team Inventory includes 100% of models in production
- $10k 5-6
Run model explainability workshop Underwriting lead
- All underwriting models include SHAP/LIME explanations $5k
- 7-8 Build AI KPI dashboard (Power BI)
- FP&A Dashboard includes loss ratio, cycle time, combined ratio
- $15k 9-10
- Run red-team exercise on pricing models Compliance
- Red-team identifies 3 edge cases for pricing models $8k
11-12 Publish AI maturity report to board
| CFO Board approves $1.2M budget for next phase | Notice the budget is front-loaded: $88k in the first two months. This is intentional. The steering committee needs early wins to maintain momentum. If the budget is spread evenly, the committee loses interest. | I once ran this plan at a $1B life & health carrier. The data catalog vendor underestimated effort—it took 6 weeks instead of 2. The steering committee wanted to pivot to a cheaper vendor, but the assessor pushed back: “We committed to Collibra. Let’s finish.” They did, and the catalog became the foundation for their dynamic pricing model. |
|---|---|---|
| Step 7: Validate the model with a pilot Pick one high-impact use case that’s already in production. I use this criteria: | Already generating revenue or reducing losses Uses structured data (avoids NLP complexity) | Has a clear owner who’s bought into the maturity model Good pilot candidates: |
| Predictive underwriting scoring (loss ratio impact) Claims leakage detection (cycle time and loss ratio) | Fraud graph analytics (loss ratio) Bad pilot candidates: | Generative AI for policy drafting (no clear ROI) Computer vision for roof inspections (requires new data) |
| LLMs for call center transcripts (regulatory risk) The pilot must run for 90 days with the following milestones: | Baseline metrics (loss ratio, cycle time, combined ratio) Model deployment with version control | Explainability report (SHAP/LIME) Model risk validation (compliance sign-off) |
| ROI calculation (pre- vs. post-deployment) | I once ran a pilot on claims leakage detection at a $3B auto insurer. The model reduced leakage by 8%, saving $2.4M annually. The steering committee used this to justify the full maturity model rollout. Without the pilot, the CFO wouldn’t have funded the data catalog. | Step 8: Scale the model across the enterprise Scaling requires two things: a central AI enablement team and a federated model governance structure. |
The central team owns: Data catalog
Model registry AI KPI dashboard
Model risk framework The federated teams (claims, underwriting, actuarial) own:
- Use case ideation Data labeling
- Model explainability Business validation
- The central team must be staffed as follows for a $5B insurer: Role
Headcount Responsibilities
> 2024-AI-Maturity-Assessment > ├── 01-Data-Infrastructure > │ ├── Data-Catalog.pdf > │ ├── Data-Lineage-Screenshot.png > │ └── PCI-Compliance-Letter.pdf > ├── 02-Model-Development > │ ├── Model-Playbook.pdf > │ ├── Explainability-Report.pdf > │ └── Version-Control-Log.pdf > ├── 03-Operations-Governance > │ ├── Model-Risk-Framework.pdf > │ ├── Audit-Report.pdf > │ └── Incident-Log.csv > └── 04-Business-Impact > ├── ROI-Calculation.xlsx > ├── A-B-Test-Results.pdf > └── Peer-Benchmark-Report.pdf
AI Product Manager 1
Roadmap, prioritization, stakeholder alignment Data Engineer
2 Data catalog, data quality, infrastructure
ML Engineer 1
Model deployment, versioning, MLOps Model Risk Analyst
| 1 Model risk framework, audit, compliance | BI Analyst 1 | AI KPI dashboard, reporting The federated teams must embed one data scientist each. I’ve seen this fail when the central team tries to do everything—federated teams must own their data and models. The central team’s job is to enable, not to build. | Step 9: Automate the assessment After two or three assessments, you’ll notice patterns: | Certain questions always score low (e.g., “Do you track data lineage?”). Certain questions always score high (e.g., “Do you have a model inventory?”). | Certain evidence is reused (e.g., model risk framework, data catalog). Automate the repetitive parts with a simple Python script. Below is the core logic I reuse: |
|---|---|---|---|---|---|
| [Source: Insurtech Insights, AI Maturity Assessment Toolkit] | The script outputs an Excel report with three sheets: scores, gaps, and summary. The gaps sheet lists every question that scored below “Fully achieved” and flags missing evidence. I’ve used this script to cut assessment time from 2 weeks to 3 days. | Step 10: Iterate and sunset the model The maturity model isn’t a one-time exercise. It should be run annually, with the following triggers for mid-year updates: | New regulatory guidance (e.g., EU AI Act, NAIC AI guidance) Major AI failure (e.g., model bias, data breach) | Merger or acquisition New AI use case launched | Sunset the model when: The enterprise reaches Level 4. |
| The AI portfolio is stable and predictable. The central team shifts from enablement to optimization. | I’ve seen this happen at a $10B carrier after three years. The maturity model became redundant because the AI operating model was mature. The team pivoted to a continuous improvement loop: quarterly model audits, annual benchmarking, and real-time KPI dashboards. | Resource and timeline summary Here’s the realistic resource estimate for a $5B P&C insurer targeting Level 3: | Phase Duration | FTE Budget | Key Risk Assessment |
| 12 weeks 1 (assessor) + 5 (steering committee) | $30k (vendor tools, travel) Steering committee loses interest | Roadmap 4 weeks | 1 (assessor) + 3 (steering committee) $15k (consultant support) | Roadmap is too aspirational Pilot | 12 weeks 2 (data scientist) + 1 (ML engineer) |
$88k (vendor tools, infrastructure) Pilot fails to show ROI
Scale 24 weeks
6 (central team) + 3 (federated data scientists) $600k (salaries, tools, training)
| Central team becomes a bottleneck Automate | 8 weeks 1 (data engineer) + 0.5 (ML engineer) | $25k (tooling, cloud) Automation doesn’t save enough time | Total 52 weeks | $758k |
|---|---|---|---|---|
| >/td> | This budget is realistic for a $5B insurer. For a smaller carrier ($1B), scale down to $350k. For a larger carrier ($10B+), scale up to $1.2M. The biggest variable is the pilot—if it fails, the entire budget is at risk. That’s why the pilot must be a known winner, not an experiment. | What happens if you skip the maturity model | At a $3B regional insurer, the claims team built an NLP model to triage FNOL. It worked—until the underwriting team changed their risk selection rules. The model’s loss ratio predictions became meaningless. The CFO found out six months later when the combined ratio spiked. The fix cost $1.2M and six months of lost opportunity. The maturity model would have caught this gap in Week 4: “Do you align AI roadmaps to underwriting and claims strategy?” The answer was “No.” | At another insurer, the data science team deployed a dynamic pricing model without model risk validation. A regulator flagged it in an audit. The fix cost $400k and six months of remediation. The maturity model would have caught this gap in Week 6: “Do you validate model risk?” The answer was “No.” |
| The maturity model isn’t a compliance exercise. It’s a risk mitigation tool. The cost of skipping it is measured in lost revenue, regulatory fines, and blown budgets. The cost of running it is $758k for a $5B insurer—a rounding error compared to the risks it prevents. | Next steps If you’re accountable for AI strategy, here’s your 30-day action plan: | Appoint an assessor. Don’t delegate it. Assemble the 5-person steering committee. Get their commitment in writing. | Run the 24-question assessment. Collect evidence ruthlessly. Score the model. Identify the top three gaps. | Build a 90-day roadmap. Get board approval. Start with the pilot. Pick a known winner. Prove the model works. Then scale. The maturity model isn’t the destination—it’s the foundation. Without it, AI in insurance is just a collection of failed pilots and blown budgets. |
| Was this article helpful? Comments. | ||||
import pandas as pd
from openpyxl import load_workbook
# Load assessment data
scores = pd.read_excel("assessment_scores.xlsx", sheet_name="scores")
evidence = pd.read_excel("assessment_scores.xlsx", sheet_name="evidence")
# Map scores to CMMI levels
score_map = {
"Not achieved": 0,
"Partially achieved": 25,
"Largely achieved": 75,
"Fully achieved": 100
}
# Calculate total score
scores["numeric_score"] = scores["score"].map(score_map)
total_score = scores["numeric_score"].sum()
# Map to CMMI level
if total_score <= 24:
level = "1. Ad-hoc"
elif total_score <= 49:
level = "2. Repeatable"
elif total_score <= 74:
level = "3. Managed"
else:
level = "4. Optimizing"
# Generate gap analysis
gaps = scores[scores["score"] != "Fully achieved"]
gaps[["question", "score", "evidence_missing"]] = gaps.apply(
lambda row: [row["question"], row["score"], evidence[evidence["question"] == row["question"]]["document"].isna().all()],
axis=1, result_type="expand"
)
# Save output
with pd.ExcelWriter("ai_maturity_report.xlsx") as writer:
scores.to_excel(writer, sheet_name="scores", index=False)
gaps.to_excel(writer, sheet_name="gaps", index=False)
pd.DataFrame({"total_score": [total_score], "level": [level]}).to_excel(writer, sheet_name="summary", index=False)