Back to News
RSS feedarxiv.org

NAQD-Env Benchmarks Selective Withdrawal in Language Agents

Summary

Language agents may need to revise planned actions when evidence changes, permission is revoked, or a stop instruction arrives. The authors introduce NAQD-Env, a synthetic benchmark that compares agent decisions with a deterministic reference policy over evidence, authorization, and constraint dependencies. Its 11 dependency families support tests on development structures, held-out families, and held-out combinations, while separate metrics measure policy agreement, task value, withdrawal, resumption, and event reporting. Three open-weight instruction-tuned models from two families were tested under three prompt conditions on 350 frozen scenarios, producing 3,150 model-prompt episodes before simulated gate replay. Across the reported conditions, withdrawal recall reached no more than 0.06, no valid resumption occurred at eligible opportunities, and only one episode matched the complete reference policy. With the NAQD prompt, Qwen2.5-7B produced fewer unsafe-attempt episodes than Qwen2.5-3B and Llama-3.1-8B, but completed less useful work and preserved unaffected actions less accurately. Exploratory supervised fine-tuning raised Qwen2.5-3B decision accuracy from 0.45-0.54 to 0.83-0.92, although diagnostics found inappropriate withdrawal after curriculum omissions and reduced event reporting. The authors say the benchmark measures policy application with trusted structured inputs and does not establish real-world containment or source-verification ability.