speech-recognition
Tracked open-source repos tagged speech-recognition, sorted by stars.
- #121★ 719+0Star change over the last 7 days
- #122★ 718+0Star change over the last 7 days
- #123
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
★ 714+0Star change over the last 7 days - #124★ 712+2Star change over the last 7 days
- #125
语音api示例
★ 708+0Star change over the last 7 days - #126
speech to text benchmark framework
★ 697+0Star change over the last 7 days - #127
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
★ 688+0Star change over the last 7 days - #128
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processing. Code included. Star the repository to support the advancement of speech technology!
★ 684+0Star change over the last 7 days - #129
A free audio dataset of spoken digits. An audio version of MNIST.
★ 678+0Star change over the last 7 days - #130
Speech Recognition for React Native Expo projects
★ 677+5Star change over the last 7 days - #131★ 671+0Star change over the last 7 days
- #132
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
★ 670+12Star change over the last 7 days - #133
Command line interface for the built-in speech recognition and transcription capabilities in macOS.
★ 669+0Star change over the last 7 days - #134
ChatGPT at home! A better alternative to commercial smart home assistants, built on the Raspberry Pi using LiteLLM and LangGraph.
★ 648+1Star change over the last 7 days - #135
:speech_balloon: /so.nus/ STT (speech to text) for Node with offline hotword detection
★ 638+0Star change over the last 7 days - #136
Android Input Method Editor (IME) based on Whisper
★ 632+4Star change over the last 7 days - #137
Real-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, livestreamers, and watching foreign content. Windows 实时音频翻译,ASR 语音识别后 LLM 流式翻译显示,适合 VTuber、主播和外语视频观看。
★ 625+34Star change over the last 7 days - #138
Real-time transcription using faster-whisper
★ 614+0Star change over the last 7 days - #139
C++ ggml runtime hub for multilingual ASR and TTS models: Cohere Transcribe, Parakeet TDT, Voxtral, Canary 1B v2, etc, plus universal forced alignment, and more
★ 609+21Star change over the last 7 days - #140
Connectionist Temporal Classification (CTC) decoder with dictionary and language model.
★ 580+1Star change over the last 7 days - #141
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 578+1Star change over the last 7 days - #142
The Self-Coding System for Your App — Alan AI SDK for React Native
★ 575+0Star change over the last 7 days - #143
A modular node-programming language, program creator, animation system, toolkit, router, and debugger made for VRChat
★ 563-1Star change over the last 7 days - #144
🗣 An overlay that gets your user’s voice permission and input as text in a customizable UI
★ 557+1Star change over the last 7 days - #145
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 553+3Star change over the last 7 days - #146
한국어 음성인식 STT API 리스트. 각 성능 벤치마크.
★ 547+1Star change over the last 7 days - #147
High-quality and streaming Speech-to-Speech interactive agent in a single file. 只用一个文件实现的流式全双工语音交互原型智能体!
★ 541+0Star change over the last 7 days - #148
Lightning-fast, free, local first voice dictation for macOS with on-device transcription
★ 540+7Star change over the last 7 days - #149
Automatic Speech Recognition(ASR), Text-To-Speech(TTS) engine. 中英语音识别、多角色语音合成,支持多语言,准确率高
★ 529+0Star change over the last 7 days - #150
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
★ 529+1Star change over the last 7 days