Back to News
RSS feedwww.askcooper.ai

Cooper’s Insurance Agent Benchmark Tests 17 AI Models on Insurance Documents

Summary

Cooper Labs introduces the first phase of its Insurance Agent Benchmark (IAB), designed to measure whether AI agents can complete insurance document tasks rather than answer isolated questions. The benchmark contains 166 documents from US commercial-lines workflows, including digital and scanned PDFs, spreadsheets, photographs, Outlook messages with attachments, long policies, damaged files, and documents with missing or unreadable values. Insurance professionals created and checked the ground truth, and answers that invent absent values are counted as failures. The study compares 17 pinned models in two settings: a model receiving the raw file in one call, and the same model using Cooper’s document harness for file-type routing, rendering, extraction, and chunked map-reduce on large files. Cooper increased the median accuracy by 9.4 percentage points across the models, although differences of roughly six points or less should be treated cautiously because each run was a single pass over 166 cases. Gemini 3.8 Flash achieved the highest point estimate, 85.6%, at an estimated $68 for the corpus; Gemini 3.7 Flash scored 83.3% at about $33 and tied for the highest reliability at 98.2%. The harness kept every model above 90% reliability, while model-alone runs failed to produce usable answers in 7% to 28% of cases. Models remained unreliable at abstaining when information was blank, redacted, or unreadable: the lowest reported hallucination rate was 10.6%. Prompt-injection resistance was comparatively strong, but long-policy clause retrieval and checkbox reading were uneven. The report measures document reading, not underwriting judgment, and is limited mainly to English-language US commercial-lines material. Later phases are planned for complete workflows and long-horizon browser tasks.