ClaimData-10K
A curated 10,000-claim gold-standard dataset. Fully expert-reviewed, adjudicated labels, and complete audit trails for model evaluation and red-teaming.
- 10,000 adjudicated claims
- 100% dual expert review
- Full provenance chain
ClaimDataPro builds verified claim scenarios, policy reasoning, damage classification, fraud indicators, adjuster workflows, reserve recommendations, settlement outcomes, and multimodal document & image data — engineered for training and evaluating advanced AI models.
Verified claim records
Claim categories covered
Expert-reviewed examples
Multimodal records
Benchmark tasks in ClaimBench
Production-grade datasets, benchmarks, and generation pipelines — purpose-built for teams training and evaluating insurance AI.
A curated 10,000-claim gold-standard dataset. Fully expert-reviewed, adjudicated labels, and complete audit trails for model evaluation and red-teaming.
100,000+ verified claims spanning property, casualty, auto, and specialty lines — built for pre-training and supervised fine-tuning at scale.
260+ benchmark tasks measuring damage classification, fraud detection, reserve estimation, settlement reasoning, and policy interpretation.
Purpose-built synthetic claim workflows matched to your book of business — rare perils, edge cases, and adversarial fraud patterns on demand.
Stream datasets, benchmarks, and sample packs through a versioned REST API with dataset diffing, license-scoped keys, and usage metering.
260+ held-out, contamination-controlled tasks. Every score below is reproducible against the public ClaimBench harness — refreshed quarterly.
Interactive scenario families spanning property, casualty, and adversarial fraud — each with verified labels and multimodal evidence.

Multimodal evidence, captured and labeled at scale
412K records
356K records
198K records
164K records
221K records
96K records
74K records
Every record passes a four-stage verification pipeline run by licensed adjusters, forensic specialists, and automated leakage controls.
Claims are sourced under license and stripped of PII through a deterministic de-identification pipeline with k-anonymity checks.
Licensed adjusters and forensic specialists label damage, causation, coverage, and fraud indicators under calibrated rubrics.
Every gold-standard record is reviewed by two independent experts; disagreements are resolved by a senior adjudication panel.
Records are stress-tested against ClaimBench task suites, checked for leakage, and versioned with full provenance metadata.
Inter-annotator agreement on gold-standard sets
PII de-identification coverage, independently audited
Train/test contamination events across all releases
Switch between live perils from ClaimData-10K — every record ships with documents, labels, provenance, and benchmark metadata.
CLM-2024-088213 · Homeowners HO-3 · FL
From foundation model labs to carrier data science teams — one verified data layer for the entire industry.
Domain pre-training corpora and evaluation sets for insurance-native reasoning.
Training data for underwriting, pricing, and claims triage products.
Benchmarks and synthetic data to validate internal models before deployment.
Adjuster workflow data and settlement outcomes for straight-through processing.
Adversarial fraud indicators, collusion networks, and staged-loss scenarios.
Custom data generation aligned to proprietary books of business and taxonomies.
From licensed source to your VPC — every stage is audited, sealed, and delivered under the governance standards procurement teams require.
Carrier partnerships, public CAT records, and domain-matched synthetic generation.
Deterministic PII scrubbing with differential-privacy guarantees on every release.
Cryptographic hashing and contamination screening before any record ships.
Direct delivery into your cloud, warehouse, or on-prem environment.
Independent audit reports available under NDA · security@claimdatapro.com
Start with a free sample pack, scale to full dataset access, or engage us for custom generation and on-prem delivery.
A 500-claim sample pack with schema docs, label rubric, and ClaimBench lite tasks.
Full ClaimData-10K access for teams validating models and running first evaluations.
Full dataset access for applied AI teams shipping insurance models to production.
Custom data generation, dedicated pipelines, and on-prem delivery for regulated environments.
Schema references, benchmark methodology, and API guides — everything your team needs to integrate ClaimDataPro into training and evaluation pipelines.

Tell us about your models and the claims data you need. We respond to every serious inquiry within one business day.