voice-cloning
Tracked open-source repos tagged voice-cloning, sorted by stars.
Related topics
Topics that frequently appear alongside voice-cloning on the same repo.
Recent risers
Repos created in the last 90 days, tagged voice-cloning.
- #1
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 61,476+138Star change over the last 7 days - #2
Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 60,117-4Star change over the last 7 days - #3
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 45,985+18Star change over the last 7 days - #4
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
★ 36,595+345Star change over the last 7 days - #5
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
★ 23,426+359Star change over the last 7 days - #6
Generate audiobooks from e-books, voice cloning & 1158+ languages!
★ 20,097+34Star change over the last 7 days - #7
Netflix-level subtitle cutting, translation, alignment, and even dubbing - one-click fully automated AI video subtitle team | Netflix级字幕切割、翻译、对齐、甚至加上配音,一键全自动视频搬运AI字幕组
★ 18,339+49Star change over the last 7 days - #8
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
★ 14,835+2,880Star change over the last 7 days - #9
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
★ 12,730+54Star change over the last 7 days - #10
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
★ 12,676+5Star change over the last 7 days - #11
YuE: Open Full-song Music Generation Foundation Model, something similar to Suno.ai but open
★ 6,416+8Star change over the last 7 days - #12
Open-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI 视频翻译配音工具。
★ 5,404+25Star change over the last 7 days - #13
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
★ 4,058+18Star change over the last 7 days - #14
A simple, high-quality voice conversion tool focused on ease of use and performance.
★ 3,674+14Star change over the last 7 days - #15★ 2,818+1Star change over the last 7 days
- #16
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214+115Star change over the last 7 days - #17★ 1,760+3Star change over the last 7 days
- #18
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
★ 1,573+2Star change over the last 7 days - #19
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
★ 1,555+4Star change over the last 7 days - #20
A Python/Pytorch app for easily synthesising human voices
★ 1,440+0Star change over the last 7 days - #21
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.
★ 1,427+1Star change over the last 7 days - #22
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #23
AI Podcast Generator for bilingual episodes, Multi Languages, Alternative to NotebookLLM;真人对话AI播客生成器,多语言,多音色
★ 1,289-1Star change over the last 7 days - #24
A webui for different audio related Neural Networks
★ 1,246+0Star change over the last 7 days - #25
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
★ 1,228+4Star change over the last 7 days - #26
A ComfyUI custom node integration for local multi-engine multi-language Text-to-Speech and Voice Conversion. Supports: RVC, Echo-TTS, Qwen3-TTS, Cozy Voice 3, Step Audio EditX, IndexTTS-2, Chatterbox (classic and multilingual), F5-TTS, Higgs Audio 2, 3, and VibeVoice with unlimited text length, SRT timing, Character support, and many audio tools
★ 1,187+7Star change over the last 7 days - #27
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
★ 1,115+245Star change over the last 7 days - #28
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B, or Audacity multi-track. Built on Qwen3-TTS.
★ 999+17Star change over the last 7 days - #29
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
★ 972+1Star change over the last 7 days - #30
A lightweight, offline Android Text-to-Speech (TTS) engine enabling seamless system-wide voice cloning and high-fidelity text reading. / 运行在安卓本地的轻量级文字转语音 (TTS) 引擎,支持离线发音人提取、零门槛音色克隆与双擎系统级全局听书。
★ 803+2Star change over the last 7 days