Back to News
RSS feedarxiv.org

Belief-Calibrated Optimization Gives LLM Agents a Persistent World Model

Summary

Researchers propose Belief-Calibrated Optimization, which stores and revises an LLM optimizer’s beliefs about how the environment responds to edits. The approach improves training pass rates across five agent benchmarks and often transfers better after the target model is swapped. An ablation indicates that the document’s information content drives the benefit.