Introduction
Resemble AI sells two products on one pricing page. The first is voice generation: text-to-speech, speech-to-speech and voice cloning, built on Chatterbox, its open-source model under the MIT licence. The second is security: deepfake detection for audio, images and video, plus audio watermarking and speaker identity search. The 2026 pricing page leads with security, and the subscription tiers are priced for security teams. If all you want is a cloned voice reading your scripts, you pay per second on the free Flex plan and never need the $350 tier.
We checked Resemble's pricing page on 29 September 2026, and the detection and seat prices below come from it. The voice generation rates come from Resemble's May 2026 pricing and its 8 May 2026 changelog. They are harder to find on the current page, so check them in the app before you set a budget.
Pricing Tiers Table
| Plan | Monthly billing | Annual billing (per month) | Seats included | Built for |
|---|---|---|---|---|
| Flex | $0 + usage | $0 + usage | 1 (extra seats $20/month each) | Developers, small TTS projects, low-volume detection |
| Team | $350 | $280 ($3,360/year) | 5 | Small trust and safety or security teams |
| Business | $1,000 | $800 ($9,600/year) | 20 | Companies that need SSO and org-wide meeting protection |
| Enterprise | Custom | Custom | Custom | On-premises deployment, SLAs, custom model training |
Voice generation rates
| Item | Price |
|---|---|
| Text-to-speech | $0.0005 per second of generated audio ($1.80 per hour) |
| Speech-to-speech | $0.0005 per second of output audio |
| First voice clone | Free, no card required |
| Further clones | $2 each per the May 2026 changelog; add-on listing shows Rapid clones at $2/month and Professional clones at $5/month per voice |
| Extra seat on Flex | $20/month per user |
Detection and security rates
| Service | Flex | Team and Business |
|---|---|---|
| Audio deepfake detection | $0.035/second | $0.015/second |
| Image detection | $0.035/image | $0.015/image |
| Video detection | $0.07/second | $0.03/second |
| Intelligence (audio analysis) | $0.025/second | $0.015/second |
| Identity search | $0.0005/call | $0.0005/call |
| Watermark encode | $0.0005/call | $0.0005/call |
| Watermark decode | $0.0002/call | $0.0002/call |
| Agent detection | $1 per 1,000 sessions | $1 per 1,000 sessions |
What Each Tier Gets You
Flex: $0 a month, pay per use
Flex has no subscription fee. You buy credits, and they never expire. You get one seat, full access to the audio, image and video detection APIs, the TTS and speech-to-speech APIs, and one free voice clone. The two cloning paths are Rapid Clone (about 10 seconds of audio, ready in under a minute) and Professional Clone (10 to 25+ minutes of audio, roughly 40 minutes of training). Chatterbox covers 23 languages.
For voice generation, Flex is the only plan most people need. Ten hours of TTS costs $18. The paid tiers lower detection and intelligence rates, and the pricing page lists no TTS discount for them.
Team: $350 a month, or $280 billed annually
Team gives you five seats and cuts detection costs by more than half: audio detection drops from $0.035 to $0.015 per second, video from $0.07 to $0.03. It also adds meetings intelligence (deepfake checks on live calls, with a one-click Microsoft Teams integration), large file uploads and batch uploads.
The break-even point is easy to work out. Team saves $0.02 per second of audio detection, so the $350 monthly fee pays for itself at 17,500 seconds, about 4.9 hours of screened audio a month. On annual billing ($280) the line drops to about 3.9 hours. Below that, Flex plus a few $20 seats is cheaper. Five seats on Flex cost $80 a month.
Business: $1,000 a month, or $800 billed annually
Business uses the same per-unit rates as Team. The extra $650 a month buys 15 more seats (20 in total), SSO, org-wide calendar integration for Meetings, security notifications for meetings and configurable deployment options. Resemble's voice creation page also says Voice Cloning API access requires Business or higher, which matters to developers (see Hidden Costs).
If you need SSO, the choice is made for you. If you only need seats, 15 extra seats at $20 each come to $300, well below the $650 gap, but Resemble does not say whether Team accepts add-on seats. Ask before you sign.
Enterprise: custom
Enterprise adds volume pricing, SLAs, custom model training, on-premises deployment and a dedicated support contact. Resemble has SOC 2 Type II compliance, and on-premises installs run on Docker or Kubernetes. Banks, telcos and government buyers who cannot send call audio to a third-party cloud will end up here. A third-party pricing tracker, checkthat.ai, puts the usual move to Enterprise at around $500 a month of Flex spend. Resemble does not publish that number.
Hidden Costs
- The cloning API sits behind Business. You can clone voices in the web app on Flex. The voice creation product page says Voice Cloning API access needs Business or higher. If your product creates clones for your users programmatically, plan for $800 to $1,000 a month before usage, or get the answer in writing from sales.
- Clone pricing is inconsistent. The 8 May 2026 changelog says clones after the first cost $2 each. Resemble's add-on listing from the same month shows $2 a month for a Rapid clone and $5 a month for a Professional clone. A one-off fee and a monthly one add up very differently. Ten Professional voices at $5 a month cost $600 a year.
- Detection is billed by the second, with no cap. On Flex, audio detection costs $126 per hour of audio. Video detection costs $252 per hour on Flex and $108 on Team or Business. A contact centre that screens every inbound call will spend far more on usage than on the plan fee.
- Seats on Flex cost $20 each. A 10-person team on Flex pays $180 a month in seats before generating a second of audio.
- Annual billing locks you in. The 20% annual discount saves $840 a year on Team and $2,400 on Business, and you commit to a full year of a product whose detection accuracy you have mostly seen in vendor claims.
- Self-hosting is free, but the GPUs are not. Chatterbox is MIT-licensed, so you can run TTS on your own hardware at no licence cost. You then pay for GPUs and the engineering time to run them.
How It Compares to Competitors
| Tool | Entry paid plan | Mid tier | Voice cloning | Deepfake detection |
|---|---|---|---|---|
| Resemble AI | Flex, $0 + $0.0005/second TTS | Team, $350/month | First clone free, then $2 | Yes, audio, image and video |
| ElevenLabs | Starter, $6/month (30,000 credits) | Creator $22, Pro $99/month | Instant from Starter, Professional from Creator | No detection product |
| Murf AI | Creator, $29/month or $19 annual (24 hours/year) | Business, $99/month or $66 annual (96 hours/year) | Enterprise only | No detection product |
| Synthesia | Starter, $29/month or $18 annual (10 video minutes/month) | Creator, $89/month or $64 annual (30 minutes/month) | Creator and up | No detection product |
On raw TTS cost, Resemble is cheap. An hour of audio costs $1.80. ElevenLabs Creator at $22 gives you 121,000 credits, which is roughly two hours of speech on its multilingual model at about one credit per character. Murf's API charges $0.03 per 1,000 characters, and an hour of narration runs to about 54,000 characters, so around $1.60 an hour. Resemble and the Murf API are close on price. ElevenLabs costs more per hour and gives you a larger voice library, a better studio editor and more tooling for dubbing and agents.
Synthesia is a video avatar tool, so you only compare it with Resemble if you need a talking presenter on screen. Its voice cloning starts on the $89 Creator plan.
Detection is where Resemble stands apart. None of the three competitors sells a deepfake detection API, so the Team and Business tiers compete with dedicated security vendors. Those vendors mostly hide their prices behind sales calls, so Resemble's published per-second rates are a point in its favour.
Which Plan Should You Pick
- Developer adding TTS or speech-to-speech to an app: Flex. At $1.80 per hour of audio you can ship and measure real usage before paying anything fixed. Check whether your cloning flow needs the API, because that pushes you to Business.
- Creator making voiceovers: Flex, or skip Resemble. ElevenLabs Creator at $22 or Murf Creator at $19 annual give you a better editing workflow for script-to-audio work.
- Security team screening under about 4 hours of audio a month: Flex with extra seats at $20 each.
- Security team screening more than 5 hours a month, or needing Teams meeting protection: Team on annual billing at $280 a month.
- Organisation that requires SSO or has more than 5 analysts: Business at $800 a month annual. Ask whether Team takes add-on seats before you accept the jump.
- Regulated industry, call audio that cannot leave your network: Enterprise, and negotiate on your projected per-second volume.
Verdict
We score Resemble AI 6 out of 10, with 6 for pricing value. Resemble's pricing is more open than its reputation suggests. Every detection rate is public, Flex costs nothing to start, and the TTS rate of $0.0005 per second is among the lowest we have seen from a hosted API. API integration is its strongest area (8 out of 10), with SDKs covering detection, watermarking and identity.
The weak spots are proof and consistency. The detection accuracy claims come from Resemble itself, and we found no independent benchmark of the detection product. Clone pricing reads differently depending on which Resemble page you trust, and the cloning API needs the $1,000 Business tier. Start on Flex, run your real audio through it for a month, and only move to Team once your usage shows you will pass about 4 hours of screened audio a month.