tts
Tracked open-source repos tagged tts, sorted by stars.
- #121
AI-powered multi-voice audiobook generator — LLM script annotation, voice cloning, voice design, LoRA training, per-line style control, and export to MP3, chaptered M4B, or Audacity multi-track. Built on Qwen3-TTS.
★ 995+16Star change over the last 7 days - #122
🧸 Lobe Vidol - Making Virtual Idols Accessible for EveryOne
★ 986+1Star change over the last 7 days - #123
A powerful 3B-parameter, LLM-based Reinforcement Learning audio edit model excels at editing emotion, speaking style, and paralinguistics, and features robust zero-shot text-to-speech
★ 972+2Star change over the last 7 days - #124★ 970+2Star change over the last 7 days
- #125★ 955-1Star change over the last 7 days
- #126
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
★ 953+3Star change over the last 7 days - #127
Make Azure natural TTS voices accessible to any SAPI 5-compatible application.
★ 933+12Star change over the last 7 days - #128
An Open-Sourced LLM-empowered Foundation TTS System
★ 921+2Star change over the last 7 days - #129
🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v12、PaddleOCR(PPOCRv5)、Whisper等主流模型
★ 908+3Star change over the last 7 days - #130
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
★ 902+0Star change over the last 7 days - #131
Webui for using XTTS and for finetuning it
★ 895+0Star change over the last 7 days - #132★ 885+0Star change over the last 7 days
- #133
Transform PDFs into AI podcasts for engaging on-the-go audio content.
★ 875+1Star change over the last 7 days - #134
A user-friendly toolkit for voice recgonition/transcription/conversion etc. | 简单易用的语音工具箱
★ 873+0Star change over the last 7 days - #135★ 867+1Star change over the last 7 days
- #136
The simplest and lowest-cost AI integration solution. If you like this project, please give it a Star~ | 最简单、最低成本的AI接入方案。喜欢本项目的话点个 Star 吧~
★ 856+2Star change over the last 7 days - #137
SumatraPDF fork: Chinese EPUB/MOBI, smart PDF dark mode, OCR, TTS, offline dictionary.
★ 825+12Star change over the last 7 days - #138
Voxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn ML framework
★ 819+2Star change over the last 7 days - #139
🔥🔥 Kokoro in Rust. https://huggingface.co/hexgrad/Kokoro-82M Insanely fast, realtime TTS with high quality you ever have.
★ 816+0Star change over the last 7 days - #140
Convert any git repository into an engaging podcast
★ 815+0Star change over the last 7 days - #141
Terminal eBook Reader with Audiobook-Quality Text-to-Speech — Supports EPUB, PDF, DOCX, HTML, RTF, TXT, and MD.
★ 806+2Star change over the last 7 days - #142★ 805+1Star change over the last 7 days
- #143
A lightweight, offline Android Text-to-Speech (TTS) engine enabling seamless system-wide voice cloning and high-fidelity text reading. / 运行在安卓本地的轻量级文字转语音 (TTS) 引擎,支持离线发音人提取、零门槛音色克隆与双擎系统级全局听书。
★ 803+4Star change over the last 7 days - #144
Speech to Text to Speech. Song now playing. Sends text as OSC messages to VRChat to display on avatar. (STTTS) (Speech to TTS) (VRC STT System) (VTuber TTS)
★ 802+3Star change over the last 7 days - #145
AI VTuber with LLM, ASR, TTS, OCR, CV and more technologies to live stream or play Minecraft with you.
★ 802+4Star change over the last 7 days - #146
Confucius4-TTS: a Multilingual and Cross-Lingual Zero-Shot TTS Engine
★ 793+9Star change over the last 7 days - #147
A modular Swift SDK for audio processing with MLX on Apple Silicon
★ 772+6Star change over the last 7 days - #148★ 763+1Star change over the last 7 days
- #149
无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, on-screen translations, and text-to-speech (TTS).
★ 757+71Star change over the last 7 days - #150
Real-time voice assistant — WebRTC streaming, faster-whisper ASR, local LLM, Vui Nano (300M) TTS. OpenAI Realtime API compatible. Voice cloning, barge-in, ~9× realtime on a 4090. Apache 2.0.
★ 756+9Star change over the last 7 days