Skip to main content
buildradar
Sign in
Topic · asr

asr

Tracked open-source repos tagged asr, sorted by stars.

86 repos
  • sherpa-ncnn@k2-fsa

    Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

    1,777-1Star change over the last 7 days
  • bailing@wwbin2017

    百聆 是一个类似GPT-4o的语音对话机器人,通过ASR+LLM+TTS实现,集成DeepSeek R1等优秀大模型,接入openClaw,真正的个人语音助手,时延低至800ms,Mac等低配置也可运行,支持打断

    1,759+2Star change over the last 7 days
  • dsnote@mkiol

    Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.

    1,622+8Star change over the last 7 days
  • video-analyzer@byjlw

    Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition

    1,570+8Star change over the last 7 days
  • amical@amicalhq

    🎙️ AI Dictation App - Open Source and Local-first ⚡ Type 3x faster, no keyboard needed. 🆓 Powered by open source models, works offline, fast and accurate.

    1,520+8Star change over the last 7 days
  • Fun-ASR@QwenAudio

    Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

    1,512+10Star change over the last 7 days
  • Speech-AI-Forge@lenML

    🍦 Speech-AI-Forge is a project developed around TTS generation model, implementing an API Server and a Gradio-based WebUI.

    1,418+0Star change over the last 7 days
  • SoniTranslate@R3gm

    Synchronized Translation for Videos. Video dubbing

    1,413+2Star change over the last 7 days
  • CrisperWhisper@nyrahealth

    Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

    1,371+13Star change over the last 7 days
  • voicemode@mbailey

    Natural voice conversations with Claude Code

    1,347+6Star change over the last 7 days
  • VideoChat@Henry-23

    实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s.

    1,303-1Star change over the last 7 days
  • StreamSpeech@ictnlp

    StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

    1,288+0Star change over the last 7 days
  • vosk-server@alphacep

    WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries

    1,262+1Star change over the last 7 days
  • Whisper-Finetune@yeyupiaoling

    Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment

    1,222-1Star change over the last 7 days
  • speech-swift@soniqo

    AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

    1,161+4Star change over the last 7 days
  • Mega-ASR@xzf-thu

    First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**

    1,136+3Star change over the last 7 days
  • conformer@sooftware

    [Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)

    1,132+1Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    1,118+143Star change over the last 7 days
  • violin@shang-zhu

    Open-source Video Translation Skill

    1,057+8Star change over the last 7 days
  • Offline speech recognition for Android with Vosk library.

    1,056+0Star change over the last 7 days
  • murmure@Kieirra

    Fully local, private and cross platform Speech-to-Text with LLM Post-processing

    1,053+15Star change over the last 7 days
  • pykaldi@pykaldi

    A Python wrapper for Kaldi

    1,040+1Star change over the last 7 days
  • sherpa@k2-fsa

    Speech-to-text server framework with next-gen Kaldi

    984+4Star change over the last 7 days
  • espresso@freewym

    Espresso: A Fast End-to-End Neural Speech Recognition Toolkit

    939+0Star change over the last 7 days
  • vocotype-cli@233stone

    VocoType 是一款运行在本地端侧的隐私安全语音输入工具,通过快捷键即可将语音实时转换为文字并自动输入到当前应用。支持语音转文字MCP、AI 优化文本、自定义替换词典、录音视频转文字等功能,让语音输入更高效、更安全。

    928+9Star change over the last 7 days
  • whisper.api@innovatorved

    This project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model.

    914+0Star change over the last 7 days
  • SmartJavaAI@geekwenjie

    🔥🔥🔥Java免费离线AI算法工具箱,支持人脸识别,活体检测,表情识别、目标检测、实例分割、行人检测、OCR文字识别、车牌识别、表格识别、ASR+TTS、机器翻译等功能,Maven引用即可使用。支持PyTorch、Tensorflow,已集成 Mtcnn、InsightFace、SeetaFace6、YOLOv8~v12、PaddleOCR(PPOCRv5)、Whisper等主流模型

    908+3Star change over the last 7 days
  • Easy-Voice-Toolkit@Spr-Aachen

    A user-friendly toolkit for voice recgonition/transcription/conversion etc. | 简单易用的语音工具箱

    873+0Star change over the last 7 days
  • PPASR@yeyupiaoling

    基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型

    870-1Star change over the last 7 days
  • GLM-ASR@zai-org

    GLM-ASR-Nano: A robust, open-source speech recognition model with 1.5B parameters

    854+3Star change over the last 7 days
← Back to topics