The catalog
* FAMILY · coqui/xtts

XTTS Cross-Language Voice Cloning

XTTS clones a voice from a six-second audio clip and speaks it in a different language, carrying one speaker's identity across language boundaries.

1 variants$0.00154 per processing secondcoqui provider
* ABOUT XTTS

Working with XTTS

XTTS is a voice generation model built around cross-language cloning: a clip of roughly six seconds is enough to capture a voice, which the model can then speak in other languages. That short sample requirement is what makes it practical.

When to choose XTTS

Choose XTTS when one person's voice has to appear in several languages and you only have a short sample. RVC converts a spoken performance into a trained voice; Mureka Vocal Clone targets singing.

* FREQUENTLY ASKED

About XTTS

How much audio does XTTS need to clone a voice?
About six seconds. That short a sample is unusual and is the main practical advantage of the model.
Can the cloned voice speak another language?
Yes — cross-language cloning is the point: the captured voice is carried into languages the original clip was not recorded in.
XTTS or RVC?
XTTS clones from a short sample and generates speech; RVC converts a spoken performance you record into an already-trained target voice.