speech-to-text
Tracked open-source repos tagged speech-to-text, sorted by stars.
- #31
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
★ 5,568+0Star change over the last 7 days - #32
Open-source AI video localization and dubbing for YouTube/Bilibili: speech recognition, subtitle translation, voice cloning, audio mixing and rendering. 开源 AI 视频翻译配音工具。
★ 5,404+0Star change over the last 7 days - #33★ 4,780+0Star change over the last 7 days
- #34
JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.
★ 4,679+0Star change over the last 7 days - #35★ 4,624+0Star change over the last 7 days
- #36
On-device subtitle generation that connects directly to DaVinci Resolve, Premiere, and After Effects.
★ 4,122+0Star change over the last 7 days - #37
Lightweight and powerful real-time audio/speech translation tool based on Windows LiveCaptions.
★ 3,637+0Star change over the last 7 days - #38
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
★ 3,403+0Star change over the last 7 days - #39
OpenAI Whisper ASR Webservice API
★ 3,328+0Star change over the last 7 days - #40
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
★ 3,171+0Star change over the last 7 days - #41
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
★ 3,147+0Star change over the last 7 days - #42
Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative
★ 3,101+0Star change over the last 7 days - #43
Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!
★ 3,068+0Star change over the last 7 days - #44★ 2,864+0Star change over the last 7 days
- #45
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2,842+0Star change over the last 7 days - #46
Open-source meeting transcription API for Google Meet, Microsoft Teams & Zoom. Auto-join bots, real-time WebSocket transcripts, MCP server for AI agents. Self-host or use hosted SaaS.
★ 2,741+0Star change over the last 7 days - #47
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
★ 2,723+0Star change over the last 7 days - #48
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for native performance, just 10MB. Completely undetectable in video calls, screen shares, and recordings.
★ 2,619+0Star change over the last 7 days - #49
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
★ 2,606+0Star change over the last 7 days - #50★ 2,537+0Star change over the last 7 days
- #51
🔊 Awesome list for Whisper — an open-source AI-powered speech recognition system developed by OpenAI
★ 2,375+0Star change over the last 7 days - #52
Record your screen, ship a demo. Free and open-source, GPU-accelerated, no watermarks, no subscriptions. Windows, macOS, Linux. Actively maintained.
★ 2,287+0Star change over the last 7 days - #53★ 2,278+0Star change over the last 7 days
- #54
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214+0Star change over the last 7 days - #55
ASR/STT subtitle generator. Uses Qwen3-ASR, local LLM, Whisper, TEN-VAD. Noise-robust for JAV
★ 2,196+0Star change over the last 7 days - #56
🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks
★ 2,173+0Star change over the last 7 days - #57★ 2,170+0Star change over the last 7 days
- #58★ 2,046+0Star change over the last 7 days
- #59★ 1,892+0Star change over the last 7 days
- #60
ggml speech-to-text inference for 16+ model families
★ 1,872+0Star change over the last 7 days