Curating Always-Loaded Context for LLM Agents Under Capacity Constraints
Summary
This paper studies how to curate fixed context files, such as AGENTS.md, that LLM agents load at the beginning of every session. Because every retained token is charged again in later rounds, growing these files can reduce performance. The authors formulate curation as a capacitated assortment problem in which instructions consume a finite attention budget and retained instructions impose a per-session setup cost. They prove an upper bound on the optimal file size, independent of the number of candidate instructions, and show that appending every instruction with positive standalone value can be arbitrarily worse than choosing an optimal subset. A token budget limits the damage caused by underestimating token prices. The paper then analyzes learning from session feedback: harms and benefits from loaded instructions are observable, but missing instructions produce evidence only when their absence causes harm. Consequently, deleting instructions that agents ignore can inevitably remove useful ones, and the authors characterize how much evidence should precede adding an instruction. They also bound regret when human reviewers can inspect only a limited number of edits per period. Experiments with irrelevant rules drawn from real context files show that such rules reduce language-model compliance.