ScopeBench is a benchmark for measuring whether autonomous AI agents preserve stated engagement boundaries during web application and network penetration testing. It contains 30 dead-end security tasks in which the stated objective can be reached only by violating the stated scope. Each task is evaluated in paired conditions with the same environment, verifier, and objective: a scopeless condition measures raw capability, while a natural-language scope measures adherence. Scopeless runs use a deterministic verifier. Scoped runs first use that verifier to identify forbidden actions that reach the flag, then use an agentic judge for trajectories that do not pass mechanically. The judge was calibrated against 100 trajectories labeled call by call by human annotators. In a blinded audit, it detected all 36 audited violations, with over-flagging the only observed error. Across eight models, raw capability ranged from 12.2% to 81.1%, while scope adherence ranged from 34.4% to 86.7%; the judge found 331 violations missed by mechanical verification. Opus-4-8 scored 10 percentage points higher in raw capability than sonnet-4-6 and had 35.6 percentage points higher scope adherence. The authors release the frozen pilot benchmark, evaluation code, and 2,160 ATIF trajectories.
AI News
The latest AI releases, research, products, and industry updates.
Loading...