text-to-speech
Tracked open-source repos tagged text-to-speech, sorted by stars.
Related topics
Topics that frequently appear alongside text-to-speech on the same repo.
Recent risers
Repos created in the last 90 days, tagged text-to-speech.
- #1
A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents
★ 2,362 - #2
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214 - #3
Open-source, local-first video editor where creators and AI agents edit the same real timeline.
★ 792 - #4
无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, on-screen translations, and text-to-speech (TTS).
★ 786 - #5
Lightning-fast, free, local first voice dictation for macOS with on-device transcription
★ 541
- #1
利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.
★ 119,977+1,509Star change over the last 7 days - #2
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
★ 75,515+337Star change over the last 7 days - #3
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 61,476+138Star change over the last 7 days - #4
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
★ 55,719+1,707Star change over the last 7 days - #5
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
★ 45,985+18Star change over the last 7 days - #6★ 39,814+9Star change over the last 7 days
- #7★ 37,419+61Star change over the last 7 days
- #8
🚀Clone a voice in 5 seconds to generate arbitrary speech in real-time
★ 36,912+0Star change over the last 7 days - #9
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
★ 36,595+345Star change over the last 7 days - #10★ 23,686+109Star change over the last 7 days
- #11
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
★ 23,426+359Star change over the last 7 days - #12★ 19,387+2Star change over the last 7 days
- #13
Translate the video from one language to another and embed dubbing & subtitles.
★ 18,871+38Star change over the last 7 days - #14★ 17,476+1Star change over the last 7 days
- #15
VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
★ 14,835+2,880Star change over the last 7 days - #16
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
★ 14,575+95Star change over the last 7 days - #17
Lightning-Fast, On-Device, Multilingual TTS — running natively via ONNX.
★ 13,757+20Star change over the last 7 days - #18
Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.
★ 12,730+54Star change over the last 7 days - #19
Use Microsoft Edge's online text-to-speech service from Python WITHOUT needing Microsoft Edge or Windows or an API key
★ 11,847+29Star change over the last 7 days - #20
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10,276+1Star change over the last 7 days - #21★ 9,949+3Star change over the last 7 days
- #22
EmotiVoice 😊: a Multi-Voice and Prompt-Controlled TTS Engine
★ 8,522-1Star change over the last 7 days - #23
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
★ 7,826+23Star change over the last 7 days - #24
High-quality multi-lingual text-to-speech library by MyShell.ai. Support English, Spanish, French, Chinese, Japanese and Korean.
★ 7,616+7Star change over the last 7 days - #25
eSpeak NG is an open source speech synthesizer that supports more than hundred languages and accents.
★ 6,802+13Star change over the last 7 days - #26
On-device Speech AI for Apple Silicon
★ 6,353+8Star change over the last 7 days - #27
StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models
★ 6,342+3Star change over the last 7 days - #28
This repository contains a hand-curated resources for Prompt Engineering with a focus on Generative Pre-trained Transformer (GPT), ChatGPT, PaLM etc
★ 6,306+7Star change over the last 7 days - #29
Silero Models: pre-trained text-to-speech models made embarrassingly simple
★ 6,084-1Star change over the last 7 days - #30
Turn your PC, Mac, or Linux box into an AI server. LLM inference, chat UI, voice, agents, workflows, RAG, and image generation.
★ 5,931+1,108Star change over the last 7 days