Back to Moonshine

Moonshine Voice

docs/index.md

0.1.22.1 KB
Original Source

Moonshine Voice

Voice Interfaces for Everyone

Moonshine Voice is an open source AI toolkit for developers building real-time voice agents and applications.

  • Everything runs on-device, so it's fast, private, and you don't need an account, credit card, API keys, or a GPU.
  • The framework and models are optimized for live streaming applications, offering low latency responses by doing a lot of the work while the user is still talking.
  • All speech to text models are based on our cutting edge research and trained from scratch, so we can offer higher accuracy than Whisper Large V3 at the top end, down to tiny 1MB models for constrained deployments.
  • It's easy to integrate across platforms, with the same library running on Python, iOS, Android, MacOS, Linux, Windows, Raspberry Pis, IoT devices, microcontrollers, DSPs, and wearables.
  • Batteries are included. Its high-level APIs offer complete solutions for common tasks like transcription, text to speech, voice cloning, speaker identification (diarization), command recognition, and conversational agents, so you can build your voice application with a single library.
  • It supports multiple languages, including English, Spanish, Mandarin, Japanese, Korean, Vietnamese, Ukrainian, and Arabic for STT, and English, Spanish, Arabic, German, French, Hindi, Italian, Japanese, Korean, Dutch, Portuguese, Russian, Turkish, Ukrainian, Vietnamese, and Mandarin for TTS.

Join our community on Discord to get live support.