Which AI models understand insurance work?
We evaluate frontier models on two separate insurance tasks: price estimation for anonymized umbrella quote rows, and brokerage/agent task reasoning against benchmark reference answers for underwriting, eligibility, and coverage questions.
Price-estimation results compare predicted premiums and uncertainty ranges against actual quote outcomes. These metrics are separate from the brokerage/agent task leaderboard.
Price-estimation rows are scored against actual quote outcomes. Brokerage/agent task rows are scored against reference answers and judged separately, so their leaderboard should be read as answer-quality performance rather than premium-estimation performance.
General AI benchmarks rarely measure whether a model can reason through the details that matter in insurance: liability limits, carrier constraints, premium ranges, eligibility rules, and uncertainty. This benchmark focuses on those workflows.
Quote rows compare each model's estimated annual premium and range against the actual quote outcome. Coverage rewards calibrated ranges; MAPE rewards accurate point estimates; Winkler loss penalizes ranges that miss the actual quote or are too wide.
Brokerage/agent task rows compare model answers to benchmark reference answers with AI judging. The public task view reports aggregate judge scores, pairwise Elo, and win rate only.
For each shared scenario or question, every pair of model outputs is compared. Better outputs win the local battle, ties split credit, and Elo updates model strength within that benchmark section.
Public results are aggregate-only. The page does not expose raw prompts, row identifiers, model responses, judge reasoning, or any operational eval artifacts. The evals use anonymized data on no-retention and no-logging platforms, so customer data is never exposed even to model providers.
Hacker News
news.ycombinator.com