Back to News
RSS feedarxiv.org

Why Time-Budgeted AI Agents Still Struggle to Use Time Productively

Summary

This study examines whether small language-model agents can both respect explicit wall-clock budgets and use the available time to improve their task performance. The researchers evaluate Qwen3.6-27B on five MLE-Bench Lite competitions and Qwen3-4B on Zork I through Jericho, where additional computation can affect results. When the budget is mentioned only in the prompt, the agents do not reliably control their runtime. The paper attributes this to missing timing feedback, poor estimates of action duration, and the absence of a learned mapping between remaining time and strategy. The researchers test harness mechanisms that expose timing information or enforce deadlines, as well as reinforcement learning with budget-aware rewards. Timing information substantially improves Qwen3.6-27B’s budget adherence without measurable performance loss, while enforcement hooks make adherence tighter. GRPO training produces near-perfect adherence on Zork I and generalizes to budgets excluded from training, but it does not outperform the untrained timing-aware harness on MLE-Bench. More importantly, learning to stop on time does not make agents use extra time productively: policies often repeat actions, and training across multiple budgets tends to collapse toward the shortest-budget strategy. The results identify time adherence and productive time allocation as separate challenges for budget-conditioned agents.