asr
Tracked open-source repos tagged asr, sorted by stars.
Related topics
Topics that frequently appear alongside asr on the same repo.
Recent risers
Repos created in the last 90 days, tagged asr.
- #1
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214 - #2
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
★ 501 - #3
早耳 - Real-time multilingual speech-to-text on CPU only. Live subtitles, browser dashboard, speaker labels, translation. No GPU, no cloud.
★ 327
- #1★ 23,864+61Star change over the last 7 days
- #2
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
★ 20,136+64Star change over the last 7 days - #3
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech)
★ 18,379+21Star change over the last 7 days - #4
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
★ 15,101+17Star change over the last 7 days - #5
Speech-to-text, text-to-speech, speaker diarization, speech enhancement, source separation, and VAD using next-gen Kaldi with onnxruntime without Internet connection. Support embedded systems, Android, iOS, HarmonyOS, Raspberry Pi, RISC-V, RK NPU, Axera NPU, Ascend NPU, x86_64 servers, websocket server/client, support 12 programming languages
★ 14,575+95Star change over the last 7 days - #6
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
★ 12,676+5Star change over the last 7 days - #7
A PyTorch-based Speech Toolkit
★ 11,802+10Star change over the last 7 days - #8
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
★ 9,209+37Star change over the last 7 days - #9
This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!
★ 8,143+14Star change over the last 7 days - #10
🤖 wukong-robot 是一个简单、灵活、优雅的中文语音对话机器人/智能音箱项目,支持ChatGPT多轮对话能力,还可能是首个支持脑机交互的开源智能音箱项目。
★ 7,126+2Star change over the last 7 days - #11
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
★ 6,210+13Star change over the last 7 days - #12
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
★ 5,634+1Star change over the last 7 days - #13
html5 js 录音 mp3 wav ogg webm amr g711a g711u 格式,支持pc和Android、iOS部分Web浏览器、Hybrid App(提供Android iOS App源码)、微信,提供ASR语音识别转文字 H5版语音通话聊天示例 DTMF编码解码
★ 5,631+3Star change over the last 7 days - #14★ 5,229+1Star change over the last 7 days
- #15
The media player for language learning, with dual subtitles, AI-generated subtitles, real-time translation, and more!
★ 4,070+12Star change over the last 7 days - #16
Streamer-Sales 销冠 —— 卖货主播 LLM 大模型🛒🎁,一个能够根据给定的商品特点从激发用户购买意愿角度出发进行商品解说的卖货主播大模型。🚀⭐内含详细的数据生成流程❗ 📦另外还集成了 LMDeploy 加速推理🚀、RAG检索增强生成 📚、TTS文字转语音🔊、数字人生成 🦸、 Agent 使用网络查询实时信息🌐、ASR 语音转文字🎙️、Vue 生态搭建前端🍍、FastAPI 搭建后端🗝️、Docker-compose 打包部署🐋
★ 3,764+1Star change over the last 7 days - #17
Hold a key, speak, release — AI-polished text appears at your cursor in any app. Open-source voice input for macOS & Windows. (按住快捷键说话,松开即得润色后的文字)
★ 3,403+44Star change over the last 7 days - #18
OpenAI Whisper ASR Webservice API
★ 3,328+2Star change over the last 7 days - #19
Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.
★ 3,171+2Star change over the last 7 days - #20
[AutoArk] GPA (General Purpose Audio) can do ASR, TTS and voice conversion with one tiny model!
★ 3,061+164Star change over the last 7 days - #21
faster_whisper GUI with PySide6
★ 2,995+4Star change over the last 7 days - #22★ 2,864+0Star change over the last 7 days
- #23
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2,842+1Star change over the last 7 days - #24
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
★ 2,723+13Star change over the last 7 days - #25
🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
★ 2,606+1Star change over the last 7 days - #26
快速提取音视频内容,整理成一份结构化的markdown笔记
★ 2,483+7Star change over the last 7 days - #27
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
★ 2,214+115Star change over the last 7 days - #28
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
★ 1,975+2Star change over the last 7 days - #29
ggml speech-to-text inference for 16+ model families
★ 1,872+19Star change over the last 7 days - #30
Next-gen AI+IoT framework for T2/T3/T5AI/ESP32/and more – Fast IoT and AI Agent hardware integration
★ 1,816+5Star change over the last 7 days