Home / Companies / TestMu AI / Blog / Post Details
Content Deep Dive

Tavus Griffin: What Its Benchmarks Mean for Video Agent Testing

Blog post from TestMu AI

Post Details
Company
Date Published
Author
Anubhav Singhmaar
Word Count
2,714
Company Posts That Month
10
Language
English
Hacker News Points
-
Post removed?
No
Summary

Tavus announced Griffin, a full-duplex video-to-video “Human Interaction Model” that simultaneously interprets audio and visual signals and generates a responsive face, voice, and behavior in real time, aiming to avoid the delays and lost context common in cascaded speech-to-avatar systems. In Tavus’s one-minute study, 26 of 54 participants believed they had spoken with a real person, compared with 2.4% for its previous stack, though the small, company-run study and narrow conversational setting limit the result’s generalizability. Griffin-Lite also led NVIDIA’s VideoFDB benchmark among non-human systems, scoring 3.83 out of 5 for generation and 3.73 for perception, but remained behind human references in latency, nonverbal cue appropriateness, visual grounding, and conversational flow. Tavus attributes its capabilities to concurrent conversational modeling and streaming audio-visual generation, including responses to interruptions, gestures, expressions, and visual tasks, while reporting separate internal results showing low video-generation latency and strong visual quality. Griffin-Lite is currently limited to trusted testers as Tavus develops disclosure and safety features to address deception risks; developers can still use Tavus’s existing rendering, perception, and turn-taking models. The discussion emphasizes that video-agent evaluations should assess response timing, turn-taking, natural nonverbal behavior, lip-sync, task accuracy, AI disclosure, and performance across varied user behaviors rather than relying only on transcripts or benchmark scores.

Trends Found in this Post
Trend Post Mentions Total Month Mentions Posts Companies MoM
Real-time 10 No monthly metrics for this publish month.
AI Agents 3 No monthly metrics for this publish month.
LLM 2 No monthly metrics for this publish month.
AI Guardrails 1 No monthly metrics for this publish month.
Use This Data

Use this post, company, and trend context to find content marketing opportunities, perform competitive analysis, or address product feature gaps via the Plushcap MCP server or the Plushcap API.