asr
Tracked open-source repos tagged asr, sorted by stars.
- #61★ 854+3Star change over the last 7 days
- #62
Voxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn ML framework
★ 819+2Star change over the last 7 days - #63
🎙️ An intelligent voice productivity assistant that turns speech into clean text, useful actions, and structured knowledge. It helps users capture ideas, communicate naturally, automate repetitive tasks, and stay productive across different apps and workflows.
★ 809+6Star change over the last 7 days - #64
AI VTuber with LLM, ASR, TTS, OCR, CV and more technologies to live stream or play Minecraft with you.
★ 803+5Star change over the last 7 days - #65
说点啥(BiBi Keyboard):一个基于 Kotlin 的 Android 平台的 LLM 与 ASR 语音输入法键盘应用 An LLM ASR voice input method keyboard application for the Android platform based on Kotlin
★ 774+21Star change over the last 7 days - #66★ 767+0Star change over the last 7 days
- #67
基于PaddlePaddle实现的语音识别,中文语音识别。项目完善,识别效果好。支持Windows,Linux下训练和预测,支持Nvidia Jetson开发板预测。
★ 763+1Star change over the last 7 days - #68
Running speech to text model (whisper.cpp) in Unity3d on your local machine.
★ 751+1Star change over the last 7 days - #69
Pytorch实现的流式与非流式的自动语音识别框架,同时兼容在线和离线识别,目前支持Conformer、Squeezeformer、DeepSpeech2模型,支持多种数据增强方法。
★ 727+0Star change over the last 7 days - #70
Open-Source Toolkit for End-to-End Speech Recognition leveraging PyTorch-Lightning and Hydra.
★ 714+0Star change over the last 7 days - #71
Offline Speech Recognition with OpenAI Whisper and TensorFlow Lite for Android
★ 688+0Star change over the last 7 days - #72
INTERSPEECH 2023-2024 Papers: A complete collection of influential and exciting research papers from the INTERSPEECH 2023-24 conference. Explore the latest advances in speech and language processing. Code included. Star the repository to support the advancement of speech technology!
★ 684+0Star change over the last 7 days - #73★ 677+0Star change over the last 7 days
- #74
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
★ 673+15Star change over the last 7 days - #75★ 671+0Star change over the last 7 days
- #76
Real-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, livestreamers, and watching foreign content. Windows 实时音频翻译,ASR 语音识别后 LLM 流式翻译显示,适合 VTuber、主播和外语视频观看。
★ 634+39Star change over the last 7 days - #77
🤗 Optimum Intel: Accelerate inference with Intel optimization tools
★ 618+1Star change over the last 7 days - #78
语音识别理论、论文和PPT
★ 618+0Star change over the last 7 days - #79
📣 商用级开源语音自动识别程序库,开箱即用,全平台支持,中英文混合识别。A Cross-platform implementation of ASR inference. It's based on ONNXRuntime and FunASR. We provide a set of easier APIs to call ASR models.
★ 610+0Star change over the last 7 days - #80
Open Source Voice Agent Platform
★ 592+2Star change over the last 7 days - #81
A cross-platform real-time subtitle display software. 一个跨平台的实时字幕显示软件。
★ 579+4Star change over the last 7 days - #82
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 578+1Star change over the last 7 days - #83
SpeechIO Leaderboard: a large, robust, comprehensive, benchmarking platform for Automatic Speech Recognition.
★ 553+3Star change over the last 7 days - #84
A speech recognition library running in the browser thanks to a WebAssembly build of Vosk
★ 529+1Star change over the last 7 days - #85
ICASSP 2023-2024 Papers: A complete collection of influential and exciting research papers from the ICASSP 2023-24 conferences. Explore the latest advancements in acoustics, speech and signal processing. Code included. Star the repository to support the advancement of audio and signal processing!
★ 526+1Star change over the last 7 days - #86
Pocket-sized AI chatbot built using a RPI Zero 2w / 5
★ 513+10Star change over the last 7 days - #87
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
★ 502+8Star change over the last 7 days - #88
Clip any video into a narration recap with claude code skill|用 claude code skill 把任何视频剪辑成中文解说视频,支持剪映导出
★ 501—Star change over the last 7 days