Introduction
Tavus sells an API for real-time AI video agents. On the developer side it also offers async avatar video. Tavus calls the agents PALs. A PAL joins a live WebRTC call, renders a face, listens, watches your camera or screen share, and answers out loud. You can drop one into your own app or send it into a Google Meet, Zoom or Microsoft Teams call from a calendar invite.
Most avatar tools render a talking head from a script. Tavus is built for the live call: the agent takes turns with a person in real time, and you supply the LLM, tools and knowledge behind it. This review covers what the product does, what the pricing page says it costs, and where the limits are. scored.tools does not test tools hands-on, so everything below comes from Tavus's own pages and docs plus the third-party sources named in the text.
Our score: 7.2 / 10.
| Criterion | Score |
|---|---|
| Output quality | 6 |
| Ease of use | 8 |
| Pricing value | 7 |
| API and integration | 8 |
| Problem fit | 7 |
| Overall | 7.2 |
Key Features
Conversational Video Interface (CVI)
CVI is the main product. It runs live WebRTC video calls between a person and a PAL. According to the docs, one API covers the face, voice, perception, turn-taking and the call transport. That saves you from wiring together separate speech, avatar and video vendors and then fixing the latency between them.
Phoenix, Raven and Sparrow models
Tavus splits the agent into three in-house models. Phoenix renders the face. Raven handles perception, meaning it reads the user's camera, tone of voice and screen share, so the agent can react to what it sees. Sparrow handles turn-taking, which decides when the agent should speak and when it should let the person finish. Turn-taking is where many voice agents feel off, so having a dedicated model for it is a sensible design choice.
On quality, the outside evidence is mixed. An Orca Router write-up on Tavus's Griffin launch cites an independent NVIDIA benchmark (VideoFDB) and a study in which 48% of participants mistook the agent for a human. Two caveats apply. The 48% study was run by Tavus itself on 54 people, and Griffin is a research preview open to selected testers, not part of the CVI product customers buy today. Tavus's own figure for the shipping Phoenix-4.5 stack in the same kind of test was 1 in 41. A third-party review cited in our scoring evidence describes the voice as robotic. That gap is why output quality gets the lowest score here, a 6.
Bring your own LLM, tools and knowledge
For builders, this is the feature that matters most. Per the docs, you can plug in your own LLM, define tool calls, attach MCP connectors, load a knowledge base, and set memory stores, objectives and guardrails. The agent's reasoning stays in your stack. Tavus provides the face, the ears and the call. If you already have a working text or voice agent, you can put a face on it without rebuilding the logic.
Meeting bots for Meet, Zoom and Teams
A PAL can accept a calendar invite and join a Google Meet, Zoom or Microsoft Teams call. That covers use cases like interview practice, onboarding sessions and sales demos, where the other person stays in a tool they already use and doesn't have to open your app.
Custom faces
Tavus trains custom faces from a short video or a single image. Paid plans include a monthly allowance of custom faces, and extra faces are billed separately (details below). The free plan uses stock faces only.
Developer tooling
The integration surface is wide: a React component library, an embeddable widget, LiveKit and Pipecat integrations, and webhooks. The docs also ship an OpenAPI spec, an MCP server and a CLI. Our API score of 8 reflects that depth. The one gap our evidence flags is that SDKs in several languages are not shown, so outside JavaScript and React you should expect to work against the REST API directly.
There is also a no-code PAL Maker and quickstarts for web, mobile and the three meeting platforms, which supports the 8 for ease of use.
Pricing Breakdown
Tavus publishes two sets of plans on its pricing page: developer/API plans for building products, and consumer PALs plans for people who just want to talk to an agent. This review focuses on the developer plans.
| Plan | Price | Live call (CVI) minutes | Video generation | Concurrent calls | Custom faces |
|---|---|---|---|---|---|
| Free | $0/mo | 25/mo | 5 min/mo | 1 | Stock only |
| Starter | $59/mo | 100/mo, then $0.37/min | 10 min/mo, then $1/min | 3 | 3/mo, $65 each after |
| Growth | $397/mo | 1,250/mo, then $0.32/min | 100 min/mo, then $0.90/min | 10 | 7/mo, $40 each after |
| Enterprise | Custom | Volume discounts | Volume discounts | Custom | Unlimited |
The consumer PALs plans are separate: Free with 15 minutes of voice and video calls, Plus at $20/month with 150 minutes, and Max at $50/month with 500 minutes. All three include unlimited messaging.
What the per-minute rate means in practice
Straight arithmetic from the published rates: a 30-minute call costs $11.10 in overage on Starter and $9.60 on Growth. Starter's 100 included minutes cover about three such calls a month. Growth's 1,250 minutes cover about 41. Past that, 1,000 extra minutes cost $370 on Starter or $320 on Growth.
That is fine for a sales demo, a paid tutoring session or a clinical intake call, where one conversation is worth far more than ten dollars. It is expensive for a free support widget on a busy site. The pay-as-you-go model is also harder to budget for than a flat fee, which is part of why pricing value scores a 7 and not higher.
Pros & Cons
Pros
- One API covers face, voice, perception, turn-taking and the call itself.
- You can swap in your own LLM and tools, so the agent's brain stays yours.
- Pricing is published per minute, and there's a free tier to build against.
- The docs go deep, with an OpenAPI spec, an MCP server and a CLI.
- PALs can join Meet, Zoom and Teams calls from a calendar invite.
Cons
- Per-minute cost climbs fast for long or high-volume calls.
- Self-serve plans cap concurrency at 1, 3 or 10 live calls.
- Realistic faces bring disclosure and consent duties that you have to design for.
- Tavus renamed personas and replicas to PALs and faces, so older tutorials don't match the current docs.
- Voice quality draws mixed outside reviews.
The concurrency cap is the limit most likely to catch teams out. Ten simultaneous calls on a $397 plan means a launch spike, a webinar or a busy support hour hits the ceiling quickly, and raising it means an Enterprise conversation. Plan for that before you ship anything public.
Who Is It For
Tavus fits a developer who already has an agent working in text or voice and wants a live face on it for conversations where each minute carries real value. The customer cases cited in onyxranked.com's 2026 review point the same way: Final Round AI (interview practice, reported at 100K+ users) and CareFlick. Interview prep, coaching, tutoring, onboarding and sales qualification are the obvious fits.
It is a weaker fit if you need marketing or training videos rendered from a script. HeyGen and Synthesia are built around async avatar video, and Tavus's 5 to 100 included generation minutes are a side feature next to CVI. If you need a live agent but your users don't need to see a face, a voice-only stack like ElevenLabs Conversational AI avoids the video rendering cost and the extra consent questions that come with a realistic human face.
Regulated industries can use Tavus, but the disclosure work falls on you. Tavus's own Griffin preview, which nearly half of its test callers took for a human, shows where realism is heading. That is a selling point for engagement and a liability if users aren't told up front.
Verdict
Tavus is the most complete API we've reviewed for putting a live AI face on a video call while keeping your own LLM, tools and knowledge behind it. The three-model design (Phoenix, Raven, Sparrow), the Meet, Zoom and Teams bots, and the depth of the docs all hold up on paper. The weak points are cost at scale, low concurrency on self-serve plans, and outside reports that the voice can sound robotic.
Recommendation: start on the free tier and build against the 25 included minutes before you pay. Move to Starter at $59 to train a custom face and test with real users. Commit to Growth or Enterprise only once you know each call is worth more than the $0.32 to $0.37 a minute it costs. For high-value, low-volume conversations, Tavus earns its 7.2. For high-volume, low-margin traffic, the numbers don't work yet.