Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
Summary
Health systems need practical ways to assess AI-assisted psychiatric intake against clinical standards while accommodating different clinician interviewing styles and limiting evaluation burden. The authors present a clinician-grounded platform built around InterviewPlayground, a memory-augmented patient simulator for open-ended AI interviews. They used expert-authored patient vignettes, interactive simulated patients, a simulated intake environment, and evaluation modes designed for psychiatric intake. In a pilot involving six clinicians and a 25-minute assessment, a GPT-based large-language-model intake interviewer recovered 88.0% of clinically relevant items embedded in the vignettes, compared with 38.9% for clinicians in the study comparison. However, the LLM made more clinical inferences that were not grounded in the interview, at 56.8% versus 27.8%. It also characterized identified safety concerns less often, at 33.3% versus 66.7%. The findings illustrate why deployment-quality assurance must assess both information coverage and clinically important failure modes, rather than treating greater item recovery as sufficient evidence of quality.