speech-to-text
Tracked open-source repos tagged speech-to-text, sorted by stars.
- #121
Collection of resources on the applications of Large Language Models (LLMs) in Audio AI.
★ 738+0Star change over the last 7 days - #122
Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。
★ 727+0Star change over the last 7 days - #123
Rapida is an open-source, end-to-end voice AI orchestration platform for building real-time conversational voice agents with audio streaming, STT, TTS, VAD, multi-channel integration, agent state management, and observability.
★ 723+4Star change over the last 7 days - #124★ 720+0Star change over the last 7 days
- #125★ 718-1Star change over the last 7 days
- #126
语音api示例
★ 707-1Star change over the last 7 days - #127
speech to text benchmark framework
★ 698+1Star change over the last 7 days - #128
A free, local desktop app to extract subtitles (SRT) from video and translate them into any language — unlimited use, no signup, no cloud.
★ 688+15Star change over the last 7 days - #129
Speech Recognition for React Native Expo projects
★ 678+4Star change over the last 7 days - #130
Transcribe and translate voice into LRC file using Whisper and LLMs (GPT, Claude, et,al). 使用whisper和LLM(GPT,Claude等)来转录、翻译你的音频为字幕文件。
★ 678+0Star change over the last 7 days - #131★ 671+0Star change over the last 7 days
- #132
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
★ 638+0Star change over the last 7 days - #133
Creating a software for automatic monitoring in online proctoring
★ 634-1Star change over the last 7 days - #134
Fast, private, local-first voice app for Apple Silicon Macs — dictation, file/media transcription, meeting recording, Transforms, and a public automation CLI. Free and open-source.
★ 630+14Star change over the last 7 days - #135★ 620+1Star change over the last 7 days
- #136
Real-time transcription using faster-whisper
★ 614+0Star change over the last 7 days - #137
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
★ 614+11Star change over the last 7 days - #138★ 582+3Star change over the last 7 days
- #139
A Python library for solving reCAPTCHA v2 and v3 with Playwright
★ 579+3Star change over the last 7 days - #140
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 578+0Star change over the last 7 days - #141
A modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChat
★ 563+0Star change over the last 7 days - #142
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
★ 557+1Star change over the last 7 days - #143
한국어 음성인식 STT API 리스트. 각 성능 벤치마크.
★ 547+0Star change over the last 7 days - #144
Browser-only Canva-style presentation studio, powered by local Web AI.
★ 547+22Star change over the last 7 days - #145
Lightning-fast, free, local first voice dictation for macOS with on-device transcription
★ 544+11Star change over the last 7 days - #146
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
★ 529+1Star change over the last 7 days - #147
Phonetisaurus G2P
★ 521-1Star change over the last 7 days - #148
This tool uses AI to evaluate your pronunciation.
★ 518+1Star change over the last 7 days - #149
open source audio and video transcription software
★ 515+0Star change over the last 7 days - #150
Open-source AI voice typing for macOS, Windows, and Linux. Press a hotkey, speak naturally, get polished text in any app.
★ 512—Star change over the last 7 days