Back to News
RSS feedarxiv.org

Why Tool-Equipped Language Models Make Unsupported Claims

Summary

Researchers study why tool-equipped language models make unsupported final claims despite instructions against guessing. In a Qwen3-32B setup, missing evidence repaired all 33 observed claims, while uninformative responses repaired none. An automatic checker corrected wrong claims without changing correct answers. Gemma 4 used the tool in every trial and produced no unsupported claim. The findings are limited to two synthetic task families and fixed model setups.