Resemble AI Pricing 2026: Plans, Rates and Hidden Costs

Resemble AI's 2026 pricing: free pay-as-you-go Flex, Team at $350/month, Business at $1,000/month. What each tier costs in practice and who should pay.

Introduction

Resemble AI sells two products on one pricing page. The first is voice generation: text-to-speech, speech-to-speech and voice cloning, built on Chatterbox, its open-source model under the MIT licence. The second is security: deepfake detection for audio, images and video, plus audio watermarking and speaker identity search. The 2026 pricing page leads with security, and the subscription tiers are priced for security teams. If all you want is a cloned voice reading your scripts, you pay per second on the free Flex plan and never need the $350 tier.

We checked Resemble's pricing page on 29 September 2026, and the detection and seat prices below come from it. The voice generation rates come from Resemble's May 2026 pricing and its 8 May 2026 changelog. They are harder to find on the current page, so check them in the app before you set a budget.

Pricing Tiers Table

PlanMonthly billingAnnual billing (per month)Seats includedBuilt for
Flex$0 + usage$0 + usage1 (extra seats $20/month each)Developers, small TTS projects, low-volume detection
Team$350$280 ($3,360/year)5Small trust and safety or security teams
Business$1,000$800 ($9,600/year)20Companies that need SSO and org-wide meeting protection
EnterpriseCustomCustomCustomOn-premises deployment, SLAs, custom model training

Voice generation rates

ItemPrice
Text-to-speech$0.0005 per second of generated audio ($1.80 per hour)
Speech-to-speech$0.0005 per second of output audio
First voice cloneFree, no card required
Further clones$2 each per the May 2026 changelog; add-on listing shows Rapid clones at $2/month and Professional clones at $5/month per voice
Extra seat on Flex$20/month per user

Detection and security rates

ServiceFlexTeam and Business
Audio deepfake detection$0.035/second$0.015/second
Image detection$0.035/image$0.015/image
Video detection$0.07/second$0.03/second
Intelligence (audio analysis)$0.025/second$0.015/second
Identity search$0.0005/call$0.0005/call
Watermark encode$0.0005/call$0.0005/call
Watermark decode$0.0002/call$0.0002/call
Agent detection$1 per 1,000 sessions$1 per 1,000 sessions

What Each Tier Gets You

Flex: $0 a month, pay per use

Flex has no subscription fee. You buy credits, and they never expire. You get one seat, full access to the audio, image and video detection APIs, the TTS and speech-to-speech APIs, and one free voice clone. The two cloning paths are Rapid Clone (about 10 seconds of audio, ready in under a minute) and Professional Clone (10 to 25+ minutes of audio, roughly 40 minutes of training). Chatterbox covers 23 languages.

For voice generation, Flex is the only plan most people need. Ten hours of TTS costs $18. The paid tiers lower detection and intelligence rates, and the pricing page lists no TTS discount for them.

Team: $350 a month, or $280 billed annually

Team gives you five seats and cuts detection costs by more than half: audio detection drops from $0.035 to $0.015 per second, video from $0.07 to $0.03. It also adds meetings intelligence (deepfake checks on live calls, with a one-click Microsoft Teams integration), large file uploads and batch uploads.

The break-even point is easy to work out. Team saves $0.02 per second of audio detection, so the $350 monthly fee pays for itself at 17,500 seconds, about 4.9 hours of screened audio a month. On annual billing ($280) the line drops to about 3.9 hours. Below that, Flex plus a few $20 seats is cheaper. Five seats on Flex cost $80 a month.

Business: $1,000 a month, or $800 billed annually

Business uses the same per-unit rates as Team. The extra $650 a month buys 15 more seats (20 in total), SSO, org-wide calendar integration for Meetings, security notifications for meetings and configurable deployment options. Resemble's voice creation page also says Voice Cloning API access requires Business or higher, which matters to developers (see Hidden Costs).

If you need SSO, the choice is made for you. If you only need seats, 15 extra seats at $20 each come to $300, well below the $650 gap, but Resemble does not say whether Team accepts add-on seats. Ask before you sign.

Enterprise: custom

Enterprise adds volume pricing, SLAs, custom model training, on-premises deployment and a dedicated support contact. Resemble has SOC 2 Type II compliance, and on-premises installs run on Docker or Kubernetes. Banks, telcos and government buyers who cannot send call audio to a third-party cloud will end up here. A third-party pricing tracker, checkthat.ai, puts the usual move to Enterprise at around $500 a month of Flex spend. Resemble does not publish that number.

Hidden Costs

  • The cloning API sits behind Business. You can clone voices in the web app on Flex. The voice creation product page says Voice Cloning API access needs Business or higher. If your product creates clones for your users programmatically, plan for $800 to $1,000 a month before usage, or get the answer in writing from sales.
  • Clone pricing is inconsistent. The 8 May 2026 changelog says clones after the first cost $2 each. Resemble's add-on listing from the same month shows $2 a month for a Rapid clone and $5 a month for a Professional clone. A one-off fee and a monthly one add up very differently. Ten Professional voices at $5 a month cost $600 a year.
  • Detection is billed by the second, with no cap. On Flex, audio detection costs $126 per hour of audio. Video detection costs $252 per hour on Flex and $108 on Team or Business. A contact centre that screens every inbound call will spend far more on usage than on the plan fee.
  • Seats on Flex cost $20 each. A 10-person team on Flex pays $180 a month in seats before generating a second of audio.
  • Annual billing locks you in. The 20% annual discount saves $840 a year on Team and $2,400 on Business, and you commit to a full year of a product whose detection accuracy you have mostly seen in vendor claims.
  • Self-hosting is free, but the GPUs are not. Chatterbox is MIT-licensed, so you can run TTS on your own hardware at no licence cost. You then pay for GPUs and the engineering time to run them.

How It Compares to Competitors

ToolEntry paid planMid tierVoice cloningDeepfake detection
Resemble AIFlex, $0 + $0.0005/second TTSTeam, $350/monthFirst clone free, then $2Yes, audio, image and video
ElevenLabsStarter, $6/month (30,000 credits)Creator $22, Pro $99/monthInstant from Starter, Professional from CreatorNo detection product
Murf AICreator, $29/month or $19 annual (24 hours/year)Business, $99/month or $66 annual (96 hours/year)Enterprise onlyNo detection product
SynthesiaStarter, $29/month or $18 annual (10 video minutes/month)Creator, $89/month or $64 annual (30 minutes/month)Creator and upNo detection product

On raw TTS cost, Resemble is cheap. An hour of audio costs $1.80. ElevenLabs Creator at $22 gives you 121,000 credits, which is roughly two hours of speech on its multilingual model at about one credit per character. Murf's API charges $0.03 per 1,000 characters, and an hour of narration runs to about 54,000 characters, so around $1.60 an hour. Resemble and the Murf API are close on price. ElevenLabs costs more per hour and gives you a larger voice library, a better studio editor and more tooling for dubbing and agents.

Synthesia is a video avatar tool, so you only compare it with Resemble if you need a talking presenter on screen. Its voice cloning starts on the $89 Creator plan.

Detection is where Resemble stands apart. None of the three competitors sells a deepfake detection API, so the Team and Business tiers compete with dedicated security vendors. Those vendors mostly hide their prices behind sales calls, so Resemble's published per-second rates are a point in its favour.

Which Plan Should You Pick

  • Developer adding TTS or speech-to-speech to an app: Flex. At $1.80 per hour of audio you can ship and measure real usage before paying anything fixed. Check whether your cloning flow needs the API, because that pushes you to Business.
  • Creator making voiceovers: Flex, or skip Resemble. ElevenLabs Creator at $22 or Murf Creator at $19 annual give you a better editing workflow for script-to-audio work.
  • Security team screening under about 4 hours of audio a month: Flex with extra seats at $20 each.
  • Security team screening more than 5 hours a month, or needing Teams meeting protection: Team on annual billing at $280 a month.
  • Organisation that requires SSO or has more than 5 analysts: Business at $800 a month annual. Ask whether Team takes add-on seats before you accept the jump.
  • Regulated industry, call audio that cannot leave your network: Enterprise, and negotiate on your projected per-second volume.

Verdict

We score Resemble AI 6 out of 10, with 6 for pricing value. Resemble's pricing is more open than its reputation suggests. Every detection rate is public, Flex costs nothing to start, and the TTS rate of $0.0005 per second is among the lowest we have seen from a hosted API. API integration is its strongest area (8 out of 10), with SDKs covering detection, watermarking and identity.

The weak spots are proof and consistency. The detection accuracy claims come from Resemble itself, and we found no independent benchmark of the detection product. Clone pricing reads differently depending on which Resemble page you trust, and the cloning API needs the $1,000 Business tier. Start on Flex, run your real audio through it for a month, and only move to Team once your usage shows you will pass about 4 hours of screened audio a month.

Stay sharp on AI tools

Weekly picks, new reviews, and deals. No spam.