How Hex Used Evals to Reduce AI Errors Sevenfold
Summary
Hex describes how it used evaluations to develop Quick Edits, an AI feature that applies simple chart changes with a small model and hands more complicated requests to the Hex agent. The team began with several hundred cases based on real internal chart-edit requests and expanded the suite to more than 1,800 cases covering edits, handoffs, ambiguous requests, languages, chart states, and edit histories. Coding agents repeatedly tested changes and retained only those that improved results, a process Hex calls hill climbing. The wrong-edit rate fell from 21% to 3%, although Hex expects the rate in customer usage to be much lower because the evals were deliberately difficult. The team optimized wrong edits more heavily than incorrect handoffs because an unwanted chart change damages user trust, while a needless handoff mainly makes the correct result slower. Hex chose GPT-6 Luna despite an overall pass rate about one percentage point below GPT-5.6 Luna because it produced 20% fewer wrong edits at half the cost. More than half of the hill-climb gains came from the output validator rather than prompt changes. The validator was expanded to accept almost-correct formatting, round values, enforce chart-specific constraints, and reject invalid dates or property combinations. These changes raised pass rates by seven percentage points on Luna and 10 on Haiku, while reducing Haiku’s hard errors by 99%. The team also changed property labels and descriptions and removed a “no change” option that Haiku misused 30 times more often than it used correctly. To limit overfitting, nearly 40% of cases were held out, and changes were accepted only when holdout results improved. Ten attempts per case reduced run-to-run noise to 0.2 percentage points, while a full 18,000-attempt run cost under $10 on Luna and took a few minutes with high concurrency. Hex says this speed allowed it to evaluate prompts, validators, chart properties, handoff rules, and UX changes together, grounding the eval suite in the product experience it wanted to deliver.