Tavus Griffin Reaches 48% in an AI Video Turing Test

On October 1, 2026, Tavus announced Griffin, which it calls the first Human Interaction Model. In a live study, 48% of participants who spent one minute on a video call with a Griffin-powered partner said they believed they had talked to a real person. The same protocol run on the company's earlier system, built from Phoenix-4.5, Sparrow-2 and Raven-1, produced a 2.4% rate, or one person out of 41.
The announcement covers a research preview called Griffin-Lite. It is available only to a small group of trusted testers, and Tavus says it will not release the model to customers until safety work on disclosure is finished. That makes the published numbers the main evidence available, so they are worth reading closely.
What the face-to-face study did
Participants were recruited through an independent research platform. They were told they would be matched with another participant for a one-minute video call about what they were looking forward to this year. Their partner was a Griffin-Lite avatar generating its face, voice and responses in real time. Afterward, participants rated the partner on naturalness and trustworthiness, rated the conversation itself, and were asked whether it had crossed their mind that the partner might not be a person. Only then were they told it was an AI.
Of 54 participants, 26 said they believed their partner was real. Confidence was high on both sides: those who said real averaged 79%, and those who said AI averaged 81%. More than half of participants said the possibility had not occurred to them during the call, and almost all of that group concluded their partner was human. The participants who did suspect an AI tended to do so within the first 20 seconds.
On a 7-point scale, Griffin-Lite averaged 5.4 for seeming natural, 5.6 for trustworthy, and 5.8 for whether participants would enjoy talking with it again. Participants who guessed it was an AI still gave the last item a 5.4. The conversation scored 5.5 for feeling listened to and 4.9 for flowing naturally, the lowest of the five.
Reading the 48% figure
The sample sizes are small. Twenty-six of 54 is 48.1%, and a standard 95% confidence interval on that proportion runs from roughly 35% to 61% (my calculation, not Tavus's). The 2.4% baseline rests on a single person. The gap between the two systems is large enough that sample size does not explain it, but the exact pass rate is imprecise.
The design also differs from the classic Turing test. Participants were told they would talk to another participant, so they had no reason to look for a machine. A judge who knows one of two interlocutors is artificial, and who can ask probing questions over an extended session, faces a harder task than a person making small talk for 60 seconds. Tavus's claim to be first rests on this protocol and on the company's own survey. The study was run by Tavus, and the post does not describe independent replication.
The post's header also lists a figure of 37% ahead of the next best AI at reacting in the moment. The body text does not explain how that number was derived, so it is not clear what it measures.
The benchmark results
NVIDIA's VideoFDB scores full-duplex audio-visual conversation on two tracks. The generation track checks whether the system produces appropriate speech and nonverbal behavior, such as laughing with a person rather than after them. The perception track checks whether the system uses what it sees, for example a pause combined with a glance away, instead of reacting only to words. A language-model judge scores responses from 0 to 5, and NVIDIA ran the evaluation independently.
On generation, Griffin-Lite scored 3.83 against a human reference of 3.92. The next-highest system, Gemini 2.5 paired with Anam, scored 2.80. That puts Griffin 1.03 points ahead and 0.09 below the human reference, which is the basis for Tavus's statement that it is more than 12 times closer to human performance than any other system. The arithmetic holds against the numbers shown: the next system sits 1.12 points below the human reference.
On perception, Griffin-Lite scored 3.73 against a human reference of 4.20. The strongest reported baseline, MiniCPM-o 4.5, scored 3.44, with Gemini 2.5 Flash Native at 3.17 and OpenAI gpt-realtime at 2.97. The post notes that MiniCPM-o 4.5 and gpt-realtime were scored in audio-only configurations, while Griffin takes in audio and video. Takeover-rate alignment, which measures how closely the model's decisions about when to speak match reference conversations, was 62.8% on generation and 73.8% on perception, the highest on both tracks. Tavus says the same model is the only one evaluated on both tracks.
A language-model judge and a single benchmark give a useful comparison, but they are a proxy for how people experience a conversation. The face-to-face study is the closer test of that.
How the system is built
Tavus describes Griffin as two parts running at the same time. A Continuous Conversational Modeling engine takes in the person's audio and video and makes a decision at sub-second intervals about what to say and how to behave. An audio-visual generation engine turns those decisions into speech and video.
Most conversational video systems chain speech recognition, a language model, speech synthesis and an avatar renderer, and they wait for the user to stop talking before starting. Tavus's diagram shows about 1.5 seconds of silence in that arrangement before the face moves, and labels its timings as illustrative. In Griffin, the conversational model decides at each interval whether to stay quiet, nod, backchannel or take the turn. A pause for thought does not have to end the turn, and the model can begin responding while the user is still talking.
On the speech side, a codec called Tavec maps 48 kHz audio into a continuous latent of 40 values per frame at 100 frames per second, with no codebooks. Its decoder is fully causal, so audio can stream out as soon as a frame exists. A minute of speech comes to 6,000 latent frames, compared with 2.88 million raw samples. The speech model is an autoregressive diffusion transformer that generates one chunk at a time and can clone a voice from about 10 seconds of reference audio.
On the video side, a few-step autoregressive generator takes a single reference image plus streaming audio and control signals and produces 720p video in 320 ms chunks, which is eight frames at 25 fps. Tavus distilled a large bidirectional diffusion model into this generator in three stages: Distribution Matching Distillation, conversion to autoregressive form with teacher forcing, and Self-Forcing to reduce drift over long rollouts. The company says the model generates every pixel of the scene, including the chair, shadows and background, and that earlier systems built on 3D priors could not produce hand or large body gestures.
In an audio-to-video comparison against four published streaming diffusion models, Griffin-Lite averaged 0.43 seconds of true latency on H100 GPUs, half that of the next fastest method. It ranked first on DOVER, FID and THEval for visual quality and second on LSE-C for lip sync, at 7.27. Tavus notes that LSE-C rewards pronounced mouth movement, including movement past the point where it looks natural.
Safety and release
Tavus states plainly that the properties that make the model effective also allow it to deceive people about whether they are talking to an AI. The company says further alignment and safety procedures are needed before release and that it is working on disclosure features and with organizations focused on AI safety. Griffin-Lite is not available to customers. Access is by request through a form, and Tavus expects to release Griffin soon after the safety concerns are addressed. Existing products, including the Tavus API and PAL Maker, continue to run on the Phoenix, Raven and Sparrow models.
Practical implications
For organizations that deploy conversational video agents in sales, recruiting, healthcare or training, which are the verticals Tavus lists, the study suggests that a disclosure policy will matter before the technology reaches them. If roughly half of unsuspecting people cannot tell the difference after a minute, then a clear statement that the person is talking to an AI becomes the main mechanism for informed consent. Teams evaluating this class of product should ask vendors how disclosure is presented, whether it can be disabled, and how it is logged.
Buyers should also hold off on performance projections until there is more evidence. The 48% result comes from a short, unscripted small-talk call with 54 people, and the benchmark rankings come from a lightweight research preview. Conversations that run longer, involve difficult topics or put a person under stress have not been reported. The lowest rating Tavus published, 4.9 for natural flow, suggests those are the areas where gaps remain.
Independent replication with larger samples, longer calls and participants who know an AI may be involved would show how far the result holds. Until Griffin leaves the preview, Tavus's own data is the only evidence on the table.

