Back to News
RSS feedarxiv.org

Audio LLMs Can Detect When Their Transcriptions Are Unreliable

Summary

Audio large language models can misinterpret degraded recordings because an incorrect transcription leads them to answer the wrong query. This paper studies whether an audio LLM can recognize when its own transcription is unreliable. Prompting the model to judge its future transcription performs poorly: in most cases, it predicts that the transcription will be reliable. Existing signals, including speech-quality predictors, generation uncertainty, and transcript-conditioned word-error-rate estimation, provide limited help. The authors find that reliability is strongly encoded in representations from the model's audio encoder. They use those representations from a frozen encoder to build a lightweight predictor that classifies reliability before generation. When a query is predicted to be unreliable, the system can ask the user to clarify; reliable queries continue through the original audio LLM unchanged. The predictor reaches 81.10% in-domain macro-F1 and 78.09% cross-domain macro-F1, exceeding the strongest baselines by 10.33 and 11.93 points. Reliability labels can also transfer across audio LLM families, with transfer performance closely related to how their model-specific reliability boundaries align.