speech-recognition
Tracked open-source repos tagged speech-recognition, sorted by stars.
- #61
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #62
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
★ 1,398+8Star change over the last 7 days - #63
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
★ 1,371+13Star change over the last 7 days - #64
Whisper command line client compatible with original OpenAI client based on CTranslate2.
★ 1,346+2Star change over the last 7 days - #65
Personal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any time with API integration, Todo list generator, Opens any website with just a voice command, Plays Music, Wikipedia searching, Dictionary with Intelligent Sensing i.e. auto spell checking, Weather Reporting i.e. temp, wind speed, humidity, YouTube searching, Google Map searching, Youtube Downloading, etc.
★ 1,341+5Star change over the last 7 days - #66★ 1,338+42Star change over the last 7 days
- #67
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+0Star change over the last 7 days - #68
Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.
★ 1,274+1Star change over the last 7 days - #69
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
★ 1,262+1Star change over the last 7 days - #70
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
★ 1,222-1Star change over the last 7 days - #71★ 1,212-1Star change over the last 7 days
- #72
💁 Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. 🛫☑️
★ 1,179+0Star change over the last 7 days - #73
A practical lab for building, testing, and evaluating apps with Apple's Foundation Models framework.
★ 1,178+1Star change over the last 7 days - #74
A private logbook with a staff of personal AI assistants. Agents read what you record and propose what to do next — you approve the changes. End-to-end encrypted sync between your own devices — servers only ever see ciphertext. Local AI optional.
★ 1,168+3Star change over the last 7 days - #75
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
★ 1,161+4Star change over the last 7 days - #76
Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.
★ 1,153+0Star change over the last 7 days - #77★ 1,149-1Star change over the last 7 days
- #78
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
★ 1,132+1Star change over the last 7 days - #79
The Self-Coding System for Your App — Alan AI SDK for Cordova
★ 1,132+0Star change over the last 7 days - #80
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,118+130Star change over the last 7 days - #81
AI Vtuber for Streaming on Youtube/Twitch
★ 1,115+1Star change over the last 7 days - #82
The open-source iOS app that's making quality voice transcription more accessible on mobile devices.
★ 1,102+2Star change over the last 7 days - #83
💬📝 A small dictation app using OpenAI's Whisper speech recognition model.
★ 1,100+2Star change over the last 7 days - #84
Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.
★ 1,094+1Star change over the last 7 days - #85
Offline speech recognition for Android with Vosk library.
★ 1,056+0Star change over the last 7 days - #86★ 1,040+1Star change over the last 7 days
- #87
:zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords
★ 1,010+0Star change over the last 7 days - #88
An editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI models
★ 1,009+2Star change over the last 7 days - #89★ 984+4Star change over the last 7 days
- #90
Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS
★ 960+1Star change over the last 7 days