RSS feedarxiv.org
Belief-Calibrated Optimization Gives LLM Agents a Persistent World Model
Summary
Researchers propose Belief-Calibrated Optimization, which stores and revises an LLM optimizer’s beliefs about how the environment responds to edits. The approach improves training pass rates across five agent benchmarks and often transfers better after the target model is swapped. An ablation indicates that the document’s information content drives the benefit.