Health systems need practical ways to assess AI-assisted psychiatric intake against clinical standards while accommodating different clinician interviewing styles and limiting evaluation burden. The authors present a clinician-grounded platform built around InterviewPlayground, a memory-augmented patient simulator for open-ended AI interviews. They used expert-authored patient vignettes, interactive simulated patients, a simulated intake environment, and evaluation modes designed for psychiatric intake. In a pilot involving six clinicians and a 25-minute assessment, a GPT-based large-language-model intake interviewer recovered 88.0% of clinically relevant items embedded in the vignettes, compared with 38.9% for clinicians in the study comparison. However, the LLM made more clinical inferences that were not grounded in the interview, at 56.8% versus 27.8%. It also characterized identified safety concerns less often, at 33.3% versus 66.7%. The findings illustrate why deployment-quality assurance must assess both information coverage and clinically important failure modes, rather than treating greater item recovery as sufficient evidence of quality.
AI News
The latest AI releases, research, products, and industry updates.
Loading...