Back to News
RSS feedarxiv.org

The Price of Thought: Test-Time Reasoning Does Not Reliably Improve LLM Trading Returns

Summary

This study examines whether increasing inference-time reasoning improves the economic value of large language models used for trading. Researchers compare representative DeepSeek, GPT, and Gemini models while varying reasoning effort and holding the available information, prompts, output formats, and portfolio construction constant. The evaluation covers a full year of U.S. equities under numerical inputs, identifiable news, and masked news, using more than 800,000 asset predictions and repeated generations. Across all three model families, additional reasoning does not reliably improve net portfolio returns after trading costs. For DeepSeek, tested from no reasoning through maximum reasoning, performance changes nonmonotonically rather than improving steadily with more thought. Repeated generations also produce unstable treatment effects and portfolio selections even when aggregate scores are similar. The authors conclude that reasoning can alter financial decisions without reliably increasing their economic value, so each task should be validated before deployment.