Back to News
RSS feedarxiv.org

Fine-Tuning Outperforms RAG in Most Tested Language-Model World-Modeling Settings

Summary

World models simulate how environments change so agents can plan the consequences of their actions. This study systematically compares fine-tuned language models with retrieval-augmented generation (RAG) world models across five embodied, web-navigation, and social environments. Fine-tuning leads to higher agent rewards in 15 of 20 tested settings, although RAG is more data-efficient when additional experience is limited. Both approaches improve when they receive more diverse exploration data, while fine-tuned models benefit disproportionately from scaling the amount of collected experience. For RAG systems, the authors use counterfactual intervention to estimate retrieval-stage errors and find that retrievers consistently select suboptimal transitions from the experience buffer. A hierarchical query-reformulation strategy performs better than the traditional retrieval pipeline. The resulting hybrid system combines a parameterized representation of core environment dynamics with retrieval from an actively maintained memory store, and consistently outperforms the other tested methods across multiple environments and models.