Back to Moonshine

Text to Speech

docs/using/text-to-speech.md

0.1.214.9 KB
Original Source

Text to Speech

Voice interfaces often need to talk back, and Moonshine's TextToSpeech is designed to make that easy, across multiple languages. It's also self-contained, so you can use it independently from the transcription and agent modules.

Getting Started

You configure a synthesizer with chainable setters, call load() to fetch and open the voice, and then pass text into say() to speak it on the default audio device:

python
from moonshine_voice import TextToSpeech

tts = TextToSpeech().language("fr")
tts.load()
tts.say("Bonjour, mon ami")
tts.wait()  # block until playback finishes

load() blocks, since the first call may have to download a voice. Pass on_progress() a handler to drive a progress bar:

python
tts = TextToSpeech().language("fr").on_progress(lambda fraction, file: print(f"{fraction:.0%}"))
tts.load()

say() returns immediately and queues the text for background synthesis and playback. Calling say() multiple times queues each utterance in order, and the next utterance is pre-synthesized while the current one plays. You can also pass a list of strings, cancel everything with stop(), or poll with is_talking():

python
tts.say(["One.", "Two.", "Three."])
tts.stop()  # cancel remaining utterances and halt playback

If you're on a machine without an audio output, or want to do further processing, you can retrieve the audio samples using the synthesize() method:

<!-- doc-test: run -->
python
from moonshine_voice import TextToSpeech

tts = TextToSpeech().language("en-us")
tts.load()
audio_data, sample_rate = tts.synthesize("Howdy, partner")

As you can see, text to speech supports multiple languages. To see which are available, run the list_tts_languages() function:

<!-- doc-test: run -->
python
from moonshine_voice import list_tts_languages
list_tts_languages()

['ar-msa', 'de-de', 'en-gb', 'en-us', 'es-ar', 'es-es', 'es-mx', 'fr-fr', 'hi-in', 'it-it', 'ja-jp', 'ko-kr', 'nl-nl', 'pt-br', 'pt-pt', 'ru-ru', 'tr-tr', 'uk-ua', 'vi-vn', 'zh-hans']

For each language, you can list which voices are available:

<!-- doc-test: run -->
python
from moonshine_voice import list_tts_voices

list_tts_voices("ru")

{'present': [], 'downloadable': ['piper_ru_RU-denis-medium', 'piper_ru_RU-dmitri-medium', 'piper_ru_RU-irina-medium', 'piper_ru_RU-ruslan-medium']}

If a voice is marked as downloadable that means if you pass it to voice() then Moonshine will download it to a cache automatically, and it will be available on your machine with no internet access required for subsequent calls.

Voice Cloning

The integrated ZipVoice model can imitate someone's voice, given a short audio clip. Pass the clip to clone_from(), either as a path to a .wav file or as a (pcm, sample_rate) pair of mono float samples. You can also pass transcript, the text spoken in the clip; when omitted, Moonshine auto-transcribes the clip with its ASR model before cloning (this takes a few extra seconds on first use):

python
from moonshine_voice import TextToSpeech
import importlib.resources;

clone_path = importlib.resources.files("moonshine_voice.assets").joinpath("clone-test.wav")
clone_transcript = "Ever tried. Ever failed. No matter. Try Again. Fail again. Fail better."

tts = TextToSpeech().language("en-us").cloning()
tts.load()
tts.clone_from(clone_path, transcript=clone_transcript)
tts.say("Ask not what your country can do for you, but what you can do for your country")
tts.wait()

cloning() tells load() to fetch ZipVoice and its clone-ASR assets up front, so clone_from() only swaps the reference clip. Call cloning() before load() — without it, clone_from() / start_cloning() raise a clear error. Catalog voices and cloning are mutually exclusive: voice() clears cloning, and cloning() clears the catalog voice.

To clone from someone speaking into the microphone rather than from a file, start_cloning() hands back a VoiceClone that listens until it has heard enough usable speech:

python
clone = tts.start_cloning()
clone.on_ready(lambda: print("Got it, you can stop talking."))
clone.from_microphone()
tts.clone_from(clone)

Picking the clip out of the recording runs Moonshine's built-in voice-activity detector, which is compiled into the library, so nothing is downloaded for this step. from_microphone() blocks until the clip is ready or 20 seconds have passed; on_progress() reports how long it has been recording and how much speech it has found so far.

You can also try cloning from the command line. Since you won't always have easy access to a clean transcript of the speech you want to clone from, you can leave it out and have Moonshine automatically generate one, in both the API and command line.

<!-- doc-test: parse-only -->
bash
curl -O -L 'https://github.com/moonshine-ai/moonshine/raw/refs/heads/main/language-bindings/python/src/moonshine_voice/assets/clone-test.wav'

python3 -m moonshine_voice.tts \
  --clone clone-test.wav \
  --text "I am so excited about Moonshine Voice's text to speech"

Voice Samples

To help you choose a voice, here are sample clips of each one saying "Welcome to Moonshine Voice text to speech". Each entry is the voice name you can pass to voice(); click the ▶ next to it to hear it.

ZipVoice

These voices were created using the zero-shot voice cloning capabilities of ZipVoice, a high-quality flow-matching TTS model from the k2-fsa team. It takes significantly longer to generate than Kokoro or PiperTTS, but offers voice cloning and more realistic speech.

zipvoice_american_female zipvoice_american_male zipvoice_australian_male
zipvoice_canadian_female zipvoice_canadian_male zipvoice_english_female
zipvoice_english_male zipvoice_indian_female zipvoice_indian_male
zipvoice_irish_female zipvoice_irish_male zipvoice_new_zealand_female
zipvoice_northern_irish_female zipvoice_south_african_female zipvoice_south_african_male

Kokoro

These voices come from the excellent Kokoro project, an 82-million-parameter open-weight TTS model that delivers quality comparable to much larger models.

American FemaleAmerican MaleBritish FemaleBritish Male
kokoro_af_alloy kokoro_am_adam kokoro_bf_alice kokoro_bm_daniel
kokoro_af_aoede kokoro_am_echo kokoro_bf_emma kokoro_bm_fable
kokoro_af_bella kokoro_am_eric kokoro_bf_isabella kokoro_bm_george
kokoro_af_heart kokoro_am_fenrir kokoro_bf_lily kokoro_bm_lewis
kokoro_af_jessica kokoro_am_liam
kokoro_af_kore kokoro_am_michael
kokoro_af_nicole kokoro_am_onyx
kokoro_af_nova kokoro_am_puck
kokoro_af_river kokoro_am_santa
kokoro_af_sarah
kokoro_af_sky

Piper TTS

The Piper project provides over a hundred lightweight voices across all of the languages Moonshine supports, from many contributors — too many to sample here. You can listen to every Piper voice on the Piper voice samples page, and use any of them with Moonshine through the piper_ voice names returned by list_tts_voices().

Converting Graphemes to Phonemes

As you may notice from the voice names, Moonshine Voice uses models from the fantastic Kokoro and PiperTTS projects. You can find full details on all the model and data sources we use for text to speech at core/moonshine-tts/data/README.md.

Given that there are other great TTS projects out there, why does the world need yet another implementation? Moonshine tries to run on as many platforms as possible and supports commercial applications, and both Kokoro and Piper use espeak-ng to convert text strings into phonemes, representations of the noises associated with the sentence, in the International Pronunciation Alphabet. Espeak-ng is licensed under the GPL, and while I am a fan of free software, the terms do make it hard to incorporate into applications that don't also release their source code under a similar license.

In the cloud this isn't as much of an issue, as many uses of espeak-ng can be implemented by calling out to an external executable, so the dependency isn't as problematic. This isn't an option on many edge operating systems unfortunately, as the only way to include code on iOS or Android is to link it into the application, which requires open sourcing the calling code.

To allow wider usage, we developed our own "grapheme to phoneme" module that performs a similar role, but has been written from scratch. You'll find the implementation in core/moonshine-tts and it's released under the same MIT License as the rest of this code base.

Every language requires a different process to convert its written form into speech, and often it varies by dialect too. This is why espeak-ng is so widely used, it has had years of work put into it to encode linguistic knowledge into a complex set of rules, many of which are heuristics that require a lot of testing to get right. The Moonshine Voice G2P engine is still new, and will need similar tuning to handle all of the variations across languages, but I'm hoping the initial implementation is a good start and will benefit from community feedback and contributions over time. Here are the current results for intelligibility across languages, using scripts/tts_g2p_intelligibility.py:

LanguageMoonshine CERReference CER
ar_msa20.8%15.3%
de_de18.3%9.2%
en_us12.6%9.8%
es_ar7.9%10.6%
es_es4.2%4.5%
es_mx3.2%2.6%
fr_fr14.8%9.4%
hi_in26.5%15.9%
it_it24.2%11.4%
ja_jp38.1%16.8%
ko_kr25.0%18.6%
nl_nl15.9%3.3%
pt_br19.7%4.9%
pt_pt43.8%24.6%
ru_ru16.9%5.0%
tr_tr8.9%7.9%
uk_ua27.7%15.6%
vi_vn79.0%36.5%
zh_hans37.8%32.6%

If you want access to just the grapheme to phoneme capability, without the speech synthesis, you can call it directly:

<!-- doc-test: run -->
python
from moonshine_voice import GraphemeToPhonemizer

g2p = GraphemeToPhonemizer("en-us")
g2p.to_ipa("Hello world")

'həlˈoʊ wˈɝld'