Back to News
RSS feedswesweep.com

SWE-Sweep Benchmark Finds AI Still Struggles to Find Unknown Bugs

Summary

Researchers from Meta, Stanford, Harvard and the University of Washington introduced SWE-Sweep, a benchmark for testing whether AI can discover and fix software bugs that have not already been pointed out. Unlike benchmarks that provide a known issue and measure whether a model can patch it, SWE-Sweep evaluates proactive bug discovery. The benchmark contains about 4,000 real-world GitHub bugs drawn from 100 repositories across 22 programming languages, with filtering intended to ensure the tasks are solvable in the test setting. The strongest setup reported, Sol 5.6 with xhigh reasoning, fixed 4.7% of the bugs and cost $7,230 in the listed test. Other tested systems, including Luna 5.6, Terra 5.6, Opus 5, Kimi K3, GPT-5.4 Mini and Gemini 3.5 Flash Lite, achieved between 2.5% and 0.1% under the configurations shown, with listed costs ranging from $4 to $5,363. The results indicate that proactive bug finding remains difficult and that the best-performing setup was also expensive. The benchmark and its implementation are released as open source under the MIT license.