How to Discount AI Benchmark Scores for Real-World Finance Work
Summary
This analysis argues that public AI-agent benchmark scores should be treated like management-case figures: scrutinized and discounted for the assumptions behind them. It compares finance and professional-work leaderboards whose top scores rose sharply, including APEX-Agents increasing from 52.4% to 82.2% and AutomationBench from 11.6% to 51.3%, while warning that such results may not reflect production reliability. The author identifies four major sources of overstatement. Benchmarks often provide clean or pre-parsed documents, whereas real finance work involves scanned PDFs, inconsistent deal materials, complex spreadsheets, document-version discovery, OCR, layout parsing, and retrieval of footnotes or embedded charts. Single-run averages hide execution variance and task difficulty: reported multi-run gaps range from 10 to 12 points between Pass@1 and Pass^8, and from 17 to 31 points between Pass@4 and Pass^4. Small samples also produce wide uncertainty, and binary pass rates do not distinguish formatting errors from severe quantitative mistakes. Rubrics frequently check whether expected assertions appear without verifying citation support, contradictions, false figures, or whether a human analyst would approve the deliverable. Real work additionally requires clarification, live systems, permissions, version histories, side-effect monitoring, and time or token budgets, while some evaluations rely on weaker language models as judges. The article recommends evaluating representative workflows with raw documents, at least five runs per task, reliability metrics and confidence intervals, grounded citations, contradiction penalties, expert-validated judges, production-like integrations, fixed resource budgets, and a final sign-off test. Its central message is that benchmark scores measure useful reasoning under specified conditions, but not necessarily the speed, trustworthiness, or review burden of deployed finance automation.