Griffin by Tavus: A New Standard for Live AI Avatars
Tavus has introduced Griffin, a real-time AI avatar generation system capable of creating natural non-verbal behavior during conversations.
The technology processes audio and video streams in real time, synchronizing facial expressions, gestures, and gaze with the interlocutor's speech without using pre-recorded clips.
Griffin-Lite operates at 720p resolution and 25 frames per second. The model generates video in 320-millisecond fragments, using only three diffusion steps to decode each block of eight frames.
The average latency of the video generator on an NVIDIA H100 is 0.43 seconds after receiving the audio signal. The full median reaction time of the system in conversational tests ranges from 1.9 to 2.2 seconds.
The infrastructure supports Baseten, Daily, and Cerebrium. Exact hardware requirements for full model functionality are not disclosed.
The streaming avatar market is actively developing thanks to competitors offering similar solutions.
Google Gemini 3.8 Live supports native speech-to-speech dialogue, can see the camera and screen, and works with 97 languages via the Enterprise API.
Runway Characters allows creating characters from a single image with WebRTC support. The cost is approximately $0.20 per minute of conversation.
Vidu S2-Avatar offers real-time motion control and environment changes, accepting photos of objects for integration into the stream.
Griffin's key distinction lies in the tight coupling of conversation context understanding with the generation of non-verbal signals.
The model independently controls gestures, gaze, and emotions, adjusting them in each subsequent 320-millisecond video fragment.
Tavus claims Griffin passed the video Turing test: 48% of participants mistakenly identified the avatar as a real person in a blind experiment.
The experiment was conducted on a small sample of 54 people, who were initially told they were meeting another research participant.
The high level of realism raises concerns about potential fraud, which explains the limited access to the technology at this stage.










