Description
Chatterbox Turbo, developed by Resemble AI, is an open-source text-to-speech (TTS) model designed for speed and expressiveness. With 350 million parameters, it achieves a latency of just 75 milliseconds, making it up to six times faster than real-time on a GPU. This model supports zero-shot voice cloning, allowing users to replicate any voice using only five seconds of reference audio, without the need for additional training or fine-tuning.
Chatterbox Turbo is unique in its inclusion of paralinguistic prompting, which enables the model to perform natural vocal reactions such as sighs, gasps, and coughs in the cloned voice. This feature enhances the expressiveness of the generated audio, making it suitable for applications in voice assistants, interactive media, and real-time agent loops.
The model is built with production readiness in mind, offering a simple installation process and comprehensive documentation. It is available under the MIT license, ensuring flexibility for developers. Chatterbox Turbo also includes PerTh watermarking, a psychoacoustic technique that embeds data into audio outputs, ensuring authenticity and traceability without compromising audio quality.
Chatterbox Turbo's performance has been tested against proprietary models like ElevenLabs Turbo v2.5 and Cartesia Sonic 3, demonstrating superior results in zero-shot scenarios. The model is accessible via GitHub and Hugging Face, providing developers with the tools and resources needed to integrate it into their projects quickly.
Chatterbox Turbo Highlights
Open-source TTS model
350M parameters
75ms latency
Zero-shot voice cloning
Paralinguistic prompting
PerTh watermarking
MIT license
Streaming-ready inference
Getting Started with Chatterbox Turbo
Install: Use pip for installation
Configure: Set up with reference scripts
Use: Clone voices and generate speech
Optimize: Adjust emotion and paralinguistic tags
Chatterbox Turbo's Use Cases
- Voice Assistants
- Interactive Media
- Voice Cloning
- Emotion Control
- Secure Audio








