FAB Introduces an Open Benchmark for Financial Due Diligence Agents
Summary
SecondState’s Finance Agents Benchmark (FAB) is an open-source project for testing how well LLM agents perform financial due diligence in a synthetic company data room. It combines a task dataset containing agent instructions, source documents, and rubrics with an execution harness that runs and grades agents. The current release covers 50 tasks, 160 documents, and 231 grading criteria for Meridian Industrial Supply LLC. Its documented setup requires Python 3.11 or later, uv, Docker or Podman, and model API credentials. In the reported evaluation, four models completed three trials on all 50 tasks, producing 600 answers; GPT-6 Luna served as the judge at maximum reasoning. A task counted as passed only when every criterion passed. DeepSeek V4.1 Flash had the highest task pass rate at 60.0%, followed by GPT-6 Sol at 58.7%, GPT-6 Luna at 50.7%, and GLM 5.3 Flash at 47.3%. Criterion pass rates ranged from 76.2% to 83.4%, while the number of tasks passed in all three trials ranged from 18 to 24. The project publishes a results report, trial-level data, model answers and traces, and a versioned Hugging Face data release. It cautions that the results cover one synthetic company and should be interpreted alongside the stated task assumptions and evaluation limitations. The code is MIT-licensed, while the data, tasks, rubrics, and results use CC BY 4.0; the project asks users to report the code and dataset revisions with results.