Back to News
RSS feedhugoj0s3.dev

Building a Configurable Skill Interview with Real-Time AI Voice Agents

Summary

The article describes a sample application that evaluates a participant's knowledge through a real-time voice interview. A configurable AI interviewer asks one question at a time, adapts difficulty to the participant's answers, supports interruptions and thinking pauses, and ends after a planned duration plus optional extra time. A second AI agent reads the transcript and produces a score, level label, summary, strengths, and improvement areas; the application clamps the score and obtains the label from the skill configuration rather than trusting the model. Skills are defined entirely in JSON, including interview instructions, agent tone and voice, model effort, timing, report rules, and point requirements, with startup validation for invalid files. The architecture separates session and skill management from realtime audio processing. Interfaces cover speech-to-text and text-to-speech, the interview agent, browser audio transport, and the conversation runner, allowing OpenAI and Azure Speech integrations to be replaced. Participant audio is streamed through a WebSocket, Azure Speech emits partial and final transcriptions, and the runner cancels the agent's speech when the participant starts talking. Agent replies are streamed, split into sentences, synthesized incrementally, and played while later text is still being generated. The server, rather than the model, controls interview timing because prompt-based time tracking proved unreliable. The sample keeps sessions, transcripts, and reports in memory, stores no audio, and uses polling while reports are evaluated. The author notes that production deployments would need durable background jobs and a push-based result update mechanism. The project requires .NET 9, an OpenAI API key, an Azure Speech resource, and a microphone, and its code is available on GitHub.