1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
-
Updated
Oct 7, 2026 - Python
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
Clone a voice in 5 seconds to generate arbitrary speech in real-time
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Generate audiobooks from e-books, voice cloning & 1158+ languages!
Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音,一键全自动视频搬运AI字幕组
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
Open-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI 视频翻译配音工具。
A 100M-parameter multilingual TTS model for real-time CPU inference, voice cloning, and 48 kHz stereo generation
Beta: The truly free, and open-source cross-platform CapCut replacement (supports MCPs).
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
A simple, high-quality voice conversion tool focused on ease of use and performance.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
MARS5 speech model (TTS) from CAMB.AI
GPT-SoVITS ONNX Inference Engine & Model Converter
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Zig backend, Swift frontend macOS app with chat, music, voice, video generation.
To associate your repository with the voice-cloning topic, visit your repo's landing page and select "manage topics."