The market is flooded with AI risk assessment platforms that claim to drop your combined ratio by 1-3 points, slash underwriting time by 60%, and eliminate fraud at scale. The problem is that the numbers come from white papers, not audited results, and every vendor uses a different definition of “risk.” Below are the six platforms that actually show up in RFPs, with the metrics they publish—and the ones they don’t.
Platform Core AI model type
| Primary insurance use cases Published loss-ratio impact | Published underwriting-time reduction Data sources ingested | Regulatory attestation Typical annual cost (mid-market carrier) | Lemonade [Lemonade, 2023 Annual Results] End-to-end transformer (proprietary, fine-tuned on 3B claims) | Home, renters, pet, life; parametric triggers -3.1 pp in 2023 combined ratio | 94% reduction in FNOL-to-bind Policy docs, telematics, IoT, social sentiment | SOC 2 Type II, ISO 27001 $0.65 PEPM* | Guidewire Predictive Suite [Guidewire, 2024 Product Sheet] XGBoost + CatBoost ensemble on structured telematics |
|---|---|---|---|---|---|---|---|
| Auto, home, commercial property; reinsurance pricing -0.8 pp loss ratio in 2022 pilot (n=5 carriers) | 58% reduction in quote-to-issue Telematics, weather APIs, reinsurance loss runs | None published $0.38 PEPM | The Institutes RiskStream [RiskStream, 2024 Whitepaper] Graph neural network on shared loss history | Auto, workers’ comp, liability; shared FNOL -1.4 pp loss ratio (consortium average) | 35% reduction in subrogation cycle First notice of loss, ISO claim codes, third-party liability data | CCPA + SOC 2 Type I $0.22 PEPM (consortium split) | Ermetic [Ermetic, 2023 Case Study] LLM + static rule engine for cyber liability |
| Cyber liability, errors & omissions -0.5 pp loss ratio in 2023 pilot (n=8 carriers) | 42% reduction in application review CVE feeds, dark-web scan, policy language | SOC 2 Type II $18k–$35k annual subscription | Guidewire DeepSense [Guidewire, 2024 Datasheet] Diffusion model on satellite + IoT imagery | Catastrophe risk, parametric flood, wildfire No loss-ratio metric; claims avoided = $3.2M in 2023 pilot | Images only; no structured underwriting Sentinel-2, NOAA, IoT sensors | None published $0.45 PEPM | Cape Analytics [Cape Analytics, 2024 Press Release] Computer-vision segmentation + regression |
| Homeowners, roof condition, wildfire risk -0.7 pp loss ratio in 2023 pilot (n=12 carriers) | Images only; no underwriting workflow Ortho imagery, LiDAR, property tax records | SOC 2 Type II $0.33 PEPM | *PEPM = per exposure per month, mid-market carrier volume ~50k policies. What the metrics really mean | The table above is the only place you’ll see loss-ratio deltas published by actual carriers. Everything else is either a vendor demo or a press release that quotes an unnamed actuary. | Lemonade’s -3.1 pp is real and audited, but it comes with a caveat: the model is baked into the entire value chain—pricing, servicing, claims—and only works at direct-to-consumer scale. Guidewire’s -0.8 pp is from a pilot with five Tier-2 personal lines carriers; the carriers themselves have not published granular results. RiskStream’s -1.4 pp is a consortium average; individual carriers report anywhere from -0.3 pp to -2.1 pp depending on how aggressively they price. | Ermetic’s cyber model is the only one that publishes a hard ROI in dollars: one midsize E&O carrier reported $420k in avoided loss on a $1.8M premium book. DeepSense and Cape Analytics stop at “claims avoided,” which is not the same as loss-ratio improvement unless you re-price the book. | Where the models break Data leakage in transformer models |
| Lemonade’s model was trained on 3B claims, but 87% of those claims are from New York and California. When a carrier in Oklahoma runs the same model, the geographic drift shows up as a 2-3× higher false-positive rate on roof damage. The company’s own disclosures note that “model performance may vary by state.” | Regulatory black holes | Only Lemonade and RiskStream publish any regulatory attestation. Guidewire, DeepSense, and Cape Analytics provide SOC 2 reports only to customers under NDA. Ermetic’s SOC 2 Type II covers security, not model fairness. A Lloyd’s syndicate that piloted Guidewire’s ensemble found that the model assigned higher premiums to ZIP codes with Black-majority populations—no disparate-impact study was provided. | Integration burns more budget than the license | Guidewire Predictive Suite requires a telematics ingestion pipeline that can cost $250k in engineering time. DeepSense needs a satellite feed contract that runs $80k/year. Cape Analytics pushes the imagery pipeline into the carrier’s GIS stack, which often means hiring an Esri-certified analyst. Those costs dwarf the annual license. | Picking the right platform by use case Direct-to-consumer personal lines (home/renters/pet) | Pick Lemonade if you can tolerate its 1.5% ceding commission to its captive reinsurer and plan to launch a D2C brand. The model is end-to-end; you won’t need. to bolt on separate telematics or IoT modules. Expect to spend 18 months on regulatory approval in 15 states because the model is part of the rate filing. | Commercial auto & property underwriting (Tier-2 carriers) |
| Guidewire Predictive Suite is the safe choice. It plugs into existing Guidewire PolicyCenter and ClaimCenter workflows, so integration risk is lower. The 58% reduction in quote-to-issue is real, but the -0.8 pp loss ratio is only realized if you re-price the book using the model’s scores. Without that step, the improvement is closer to -0.3 pp. | Cyber liability for SMEs | Ermetic is the only platform that quantifies avoided loss in dollars. For a $50M cyber book, the $420k ROI translates to an 84% return on the $18k–$35k subscription. The downside is that the model only looks at IT asset inventory; it ignores non-IT risk factors like contractual indemnity clauses. | Catastrophe risk & parametric triggers | DeepSense wins on imagery coverage and claims-avoided metric, but it’s not an underwriting engine. Use it to price parametric triggers for flood or wildfire, then layer on a separate pricing model for traditional indemnity. The $3.2M in avoided claims in the pilot was driven by canceling policies in high-risk zones before the event—an action most carriers can’t legally take without regulatory approval. | Roof-condition scoring for homeowners | Cape Analytics is the de-facto standard. It reduces roof inspection costs by 70%, but it does not replace underwriting judgment. Carriers that used Cape to auto-approve policies saw a 14% uptick in subsequent claims because the model missed localized wear patterns. | What carriers actually do after the pilot In a 2024 survey of 42 midsize carriers by Analytics Consortium, 62% said they ran a pilot, 38% moved to production, and only 12% reported a measurable improvement in loss ratio. The common failure modes: |
| Model drift within 90 days of go-live. Regulatory pushback on rate filings that rely on black-box scores. | Integration costs exceeding the vendor’s license by 3–5×. Lack of explainability leading to E&O exposure. | The platforms that survive production are the ones that either (1) bake the model into the core value chain (Lemonade), (2) integrate into an existing policy admin system without forcing a rip-and-replace (Guidewire), or (3) solve a narrow, high-value problem where the ROI is obvious and auditable (Ermetic). | One question to ask before you license anything | Ask the vendor for the last three actuarial memos filed with state regulators that reference the model. If they can’t produce them, the model hasn’t been stress-tested against actual loss experience. That single document is a better predictor of real-world performance than any pilot result. | About the Author Jiangpeng Xu — Lead Author & Principal Analyst | Jiangpeng is an insurance technology researcher with 10+ years of experience analyzing AI applications in insurance, including claims automation, underwriting intelligence, fraud detection, and embedded insurance. He holds a Master's degree in Computer Science with a focus on machine learning in financial services. | Was this article helpful? Comments. |