Halv Reports 57.1% Lower Model Cost in SWE-rebench Checkpoint
Summary
Halv’s interim SWE-rebench report compares a routed coding workflow with vanilla Codex across 14 repository tasks, with three paired repetitions per task. Both arms used Codex 0.155.1 and an Astra medium coordinator, but Halv added an investigation stage, JEV model routing, optional delegation to a cheaper worker, and the Crux, RTK, and Headroom tools. Each arm passed 25 of 42 verifier runs, producing a 59.5% pass rate, while Halv recorded $144.67 in model cost versus $337.50 for vanilla Codex, a 57.1% reduction and $192.83 difference. Halv used 170,114,775 total input and output tokens, compared with 48,755,656 for vanilla, or 3.49 times as many tokens; the report attributes the lower recorded cost to a different mix of model usage and caching rather than token savings. The strongest task result was ArcadeDB-4455, where both arms passed all three repetitions and Halv recorded 81.7% lower cost. However, Halv cost more on four tasks and passed only one of three Perry-3982 repetitions versus three of three for vanilla. The report covers the first 14 of 111 planned tasks and excludes infrastructure-invalid attempts, a rolled-back policy experiment, diagnostics, and interrupted work. It therefore supports a cost claim for this combined workflow, not statistical equivalence, universal savings, or isolation of any individual component.