Study Finds LLM Refusal Can Fail Under Authority-Framed Prompts
Summary
A new arXiv study evaluates seven large language models on a clinical pain-speech task designed to distinguish justified abstention from unsupported prediction. Of 5,750 no-signal transcripts, transcript-only pain prediction was near chance, making any inferred pain score unsupported. Six models nearly always abstained under cooperative prompts and accurately extracted explicitly spoken ratings. Authority-framed prompts produced large changes in abstention rates. Gemini 2.5 Flash and Llama 3.1 8B generated confident pain scores on unsupported inputs at rates of 0.53 and 0.76, respectively, emphasizing the risks of prompt-sensitive refusal and poor calibration in clinical AI.