text-to-speech
Tracked open-source repos tagged text-to-speech, sorted by stars.
- #61★ 1,836+0Star change over the last 7 days
- #62
The Self-Coding System for Your App — Alan AI SDK for Android
★ 1,807+0Star change over the last 7 days - #63★ 1,760+0Star change over the last 7 days
- #64
The Self-Coding System for Your App — Alan AI SDK for Flutter
★ 1,756+0Star change over the last 7 days - #65
Automatically translates the text of a video based on a subtitle file, and then uses AI voice services to create a new dubbed & translated audio track where the speech is synced using the subtitle's timings.
★ 1,746+0Star change over the last 7 days - #66
An awesome browser extension that reads aloud webpage content with one click
★ 1,737+0Star change over the last 7 days - #67
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1,707+0Star change over the last 7 days - #68
The Self-Coding System for Your App — Alan AI SDK for Ionic
★ 1,649+0Star change over the last 7 days - #69
Unofficial Parallel WaveGAN (+ MelGAN & Multi-band MelGAN & HiFi-GAN & StyleMelGAN) with Pytorch
★ 1,646+0Star change over the last 7 days - #70
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
★ 1,622+0Star change over the last 7 days - #71
Topic → 4K narrated video for coding agents. v4.0: all TTS via the ttsCN engine component (11 platforms incl. MiniMax voice clone, native word-level subtitle sync), manifest-based Asset Engine, Remotion composition, cost-gated AI generation
★ 1,599+0Star change over the last 7 days - #72
Official MiniMax Model Context Protocol (MCP) server that enables interaction with powerful Text to Speech, image generation and video generation APIs.
★ 1,573+0Star change over the last 7 days - #73
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
★ 1,557+0Star change over the last 7 days - #74
A comprehensive ComfyUI integration for Microsoft's VibeVoice text-to-speech model, enabling high-quality single and multi-speaker voice synthesis directly within your ComfyUI workflows.
★ 1,555+0Star change over the last 7 days - #75★ 1,548+0Star change over the last 7 days
- #76★ 1,542+0Star change over the last 7 days
- #77
Talking Head (3D): A JavaScript class for real-time lip-sync using full-body 3D avatars.
★ 1,522+0Star change over the last 7 days - #78
A Python/Pytorch app for easily synthesising human voices
★ 1,440+0Star change over the last 7 days - #79★ 1,437+0Star change over the last 7 days
- #80
Self-host the powerful Chatterbox TTS model. This server offers a user-friendly Web UI, flexible API endpoints (incl. OpenAI compatible), predefined voices, voice cloning, and large audiobook-scale text processing. Runs accelerated on NVIDIA (CUDA), AMD (ROCm), and CPU.
★ 1,427+0Star change over the last 7 days - #81
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
★ 1,418+0Star change over the last 7 days - #82
Synchronized Translation for Videos. Video dubbing
★ 1,413+0Star change over the last 7 days - #83
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #84
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
★ 1,398+0Star change over the last 7 days - #85
[ICASSP 2024] 🍵 Matcha-TTS: A fast TTS architecture with conditional flow matching
★ 1,353+0Star change over the last 7 days - #86
[ICLR 2025] SOTA discrete acoustic codec models with 40/75 tokens per second for audio language modeling
★ 1,317+0Star change over the last 7 days - #87
Offline inference engine for art, real-time voice conversations, LLM powered chatbots and automated workflows
★ 1,314+0Star change over the last 7 days - #88
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+0Star change over the last 7 days - #89
A webui for different audio related Neural Networks
★ 1,246+0Star change over the last 7 days - #90
A Flow Matching-based Text-to-Speech Model with Emoji-driven Style Control
★ 1,228+0Star change over the last 7 days