Back to News
RSS feedgithub.com

gsplat-talkinghead Brings Lip-Synced Gaussian-Splat Avatars to AI Voice Agents

Summary

gsplat-talkinghead is an open-source React component library that adds lip-synced Gaussian-Splat avatars to AI voice agents. It runs Gaussian-Splat rendering and the wav2arkit neural lip-sync pipeline in the browser, with no separate rendering infrastructure. Provider wrappers support OpenAI Realtime, OpenAI GPT-Live, Qwen Realtime, ElevenLabs Conversational AI, Vapi, and LiveKit Agents. The library handles WebRTC or provider SDK connections, resamples audio to 16 kHz, runs an ONNX model through onnxruntime-web, and maps the output to ARKit facial blendshapes. Full neural lip sync is available when a provider exposes a remote media stream or raw PCM; ElevenLabs instead uses a coarse volume-based fallback because its SDK exposes only audio-level data. Developers supply backend callbacks that mint ephemeral keys, conversation tokens, participant tokens, or SDP answers, keeping provider secrets out of browser code. Built-in avatars include Jack, Jane, John, and Sasha, while custom Gaussian-Splat bundles can be hosted separately. Shared options cover emotions, backgrounds, session timeouts, end phrases, and session callbacks. The project also exposes adapter hooks for composing the avatar into custom layouts and supports browser-side tools, including an emotion-setting tool for compatible providers.