text-to-speech
Tracked open-source repos tagged text-to-speech, sorted by stars.
- #31★ 5,831+0Star change over the last 7 days
- #32
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
★ 5,568+0Star change over the last 7 days - #33
Open-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI 视频翻译配音工具。
★ 5,404+0Star change over the last 7 days - #34
Dockerized OpenAI-compatible wrapper for Kokoro-82M text-to-speech w/multiplatform CPU, AMD, NVIDIA GPU PyTorch; multi-speaker, auto-stitching, caption timestamps, SSML, optional readalong web UI
★ 5,399+0Star change over the last 7 days - #35
DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code
★ 4,855+0Star change over the last 7 days - #36★ 4,624+0Star change over the last 7 days
- #37
A nearly-live implementation of OpenAI's Whisper.
★ 4,246+0Star change over the last 7 days - #38
Foundational model for human-like, expressive TTS
★ 4,203+0Star change over the last 7 days - #39
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
★ 4,058+0Star change over the last 7 days - #40
Converts text to speech in realtime
★ 4,021+0Star change over the last 7 days - #41
:stuck_out_tongue_closed_eyes: TensorFlowTTS: Real-Time State-of-the-art Speech Synthesis for Tensorflow 2 (supported including English, French, Korean, Chinese, German and Easy to adapt for other languages)
★ 3,996+0Star change over the last 7 days - #42
A simple, high-quality voice conversion tool focused on ease of use and performance.
★ 3,674+0Star change over the last 7 days - #43
A single Gradio + React WebUI with extensions for ACE-Step, OmniVoice, Kimi Audio, Piper TTS, GPT-SoVITS, CosyVoice, XTTSv2, DIA, Kokoro, OpenVoice, ParlerTTS, Stable Audio, MMS, StyleTTS2, MAGNet, AudioGen, MusicGen, Tortoise, RVC, Vocos, Demucs, SeamlessM4T, and Bark!
★ 3,250+0Star change over the last 7 days - #44
The official Python SDK for the ElevenLabs API.
★ 3,083+0Star change over the last 7 days - #45
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
★ 3,061+0Star change over the last 7 days - #46
aeneas is a Python/C library and a set of tools to automagically synchronize audio and text (aka forced alignment)
★ 2,862+0Star change over the last 7 days - #47★ 2,818+0Star change over the last 7 days
- #48★ 2,628+0Star change over the last 7 days
- #49
🚀 一键部署(含离线整合包)!基于 ChatTTS ,支持流式输出、音色抽卡、长音频生成和分角色朗读。简单易用,无需复杂安装。
★ 2,593+0Star change over the last 7 days - #50
MARY TTS -- an open-source, multilingual text-to-speech synthesis system written in pure java
★ 2,583+0Star change over the last 7 days - #51★ 2,530+0Star change over the last 7 days
- #52
Vietnamese TTS with instant voice cloning • On-device • Real-time CPU inference • 24kHz audio quality • Chuyển văn bản thành giọng nói tiếng Việt • Text to speech tiếng Việt • TTS tiếng Việt
★ 2,468+0Star change over the last 7 days - #53
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
★ 2,379+58Star change over the last 7 days - #54
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
★ 2,367+0Star change over the last 7 days - #55
PyTorch implementation of VALL-E(Zero-Shot Text-To-Speech), Reproduced Demo https://lifeiteng.github.io/valle/index.html
★ 2,215+0Star change over the last 7 days - #56
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214+0Star change over the last 7 days - #57
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2,207+0Star change over the last 7 days - #58
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2,075+0Star change over the last 7 days - #59★ 2,048+0Star change over the last 7 days
- #60
AI-native video production toolkit for Claude Code
★ 2,036+0Star change over the last 7 days