asr
Tracked open-source repos tagged asr, sorted by stars.
- #31
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.
★ 1,777-1Star change over the last 7 days - #32
百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断
★ 1,759+2Star change over the last 7 days - #33
Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.
★ 1,622+8Star change over the last 7 days - #34
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
★ 1,570+8Star change over the last 7 days - #35
🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.
★ 1,520+8Star change over the last 7 days - #36
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
★ 1,512+10Star change over the last 7 days - #37
🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.
★ 1,418+0Star change over the last 7 days - #38
Synchronized Translation for Videos. Video dubbing
★ 1,413+2Star change over the last 7 days - #39
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
★ 1,371+13Star change over the last 7 days - #40★ 1,347+6Star change over the last 7 days
- #41
实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s.
★ 1,303-1Star change over the last 7 days - #42
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+0Star change over the last 7 days - #43
WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries
★ 1,262+1Star change over the last 7 days - #44
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
★ 1,222-1Star change over the last 7 days - #45
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
★ 1,161+4Star change over the last 7 days - #46
First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
★ 1,136+3Star change over the last 7 days - #47
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
★ 1,132+1Star change over the last 7 days - #48
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,118+143Star change over the last 7 days - #49★ 1,057+8Star change over the last 7 days
- #50
Offline speech recognition for Android with Vosk library.
★ 1,056+0Star change over the last 7 days - #51★ 1,053+15Star change over the last 7 days
- #52★ 1,040+1Star change over the last 7 days
- #53★ 984+4Star change over the last 7 days
- #54★ 939+0Star change over the last 7 days
- #55
VocoType 是一款运行在本地端侧的隐私安全语音输入工具,通过快捷键即可将语音实时转换为文字并自动输入到当前应用。支持语音转文字MCP、AI 优化文本、自定义替换词典、录音视频转文字等功能,让语音输入更高效、更安全。
★ 928+9Star change over the last 7 days - #56
This project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model.
★ 914+0Star change over the last 7 days - #57
🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v12、PaddleOCR(PPOCRv5)、Whisper等主流模型
★ 908+3Star change over the last 7 days - #58
A user-friendly toolkit for voice recgonition/transcription/conversion etc. | 简单易用的语音工具箱
★ 873+0Star change over the last 7 days - #59
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
★ 870-1Star change over the last 7 days - #60★ 854+3Star change over the last 7 days