Adapting ASR Models to Noisy Police Audio with Pseudo-Labeling
Summary
Pretrained automatic speech recognition systems perform poorly on noisy Broadcast Police Communication (BPC), limiting analysis of police decision-making. This study evaluates unsupervised pseudo-labeling as a way to adapt the foundation ASR models Whisper and Qwen3-ASR without costly human transcripts. Using BPC corpora from Baltimore and Chicago, the authors find that internal confidence signals, including log-probabilities and STAR scores, cannot reliably separate high-quality from poor pseudo-labels. They therefore introduce an external LLM-as-a-judge filter that uses contextual plausibility to remove likely incorrect transcripts. The filter is more aggressive than the internal metrics and significantly lowers the word error rate of the pseudo-labeled training sets in both cities. However, its results remain substantially behind those from an oracle filter, so the approach does not eliminate the filtering problem. The paper also proposes cross-model pseudo-labeling, in which one ASR model is fine-tuned using pseudo-labels generated by the other, and identifies this as a promising direction for future work.