Skip to main content
buildradar
Sign in
Topic · speech-recognition

speech-recognition

Tracked open-source repos tagged speech-recognition, sorted by stars.

156 repos
  • whisper-jax@sanchit-gandhi

    JAX implementation of OpenAI's Whisper model for up to 70x speed-up on TPU.

    4,679-1Star change over the last 7 days
  • pocketsphinx@cmusphinx

    A small speech recognizer

    4,337+2Star change over the last 7 days
  • distil-whisper@huggingface

    Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.

    4,114+2Star change over the last 7 days
  • OpenAI Whisper ASR Webservice API

    3,328+2Star change over the last 7 days
  • Whisper & Faster-Whisper standalone executables for those who don't want to bother with Python.

    3,171+2Star change over the last 7 days
  • willow@HeyWillow

    Open source, local, and self-hosted Amazon Echo/Google Home competitive Voice Assistant alternative

    3,101+1Star change over the last 7 days
  • whishper@pluja

    Transcribe any audio to text, translate and edit subtitles 100% locally with a web UI. Powered by whisper models!

    3,068+2Star change over the last 7 days
  • lingvo@tensorflow

    Lingvo

    2,864+0Star change over the last 7 days
  • whisper-timestamped@linto-ai

    Multilingual Automatic Speech Recognition with word-level timestamps and confidence

    2,842+1Star change over the last 7 days
  • STT@coqui-ai

    🐸STT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.

    2,606+1Star change over the last 7 days
  • qwen-audio-agent@QwenAudio

    A realtime voice runtime that keeps Agents talking, working, and present. Real-time Voice Runtime for AI Agents

    2,362+68Star change over the last 7 days
  • 🎙Speech recognition using the tensorflow deep learning framework, sequence-to-sequence neural networks

    2,173+0Star change over the last 7 days
  • parlor@fikrikarim

    On-device, real-time multimodal AI with features similar to GPT-Live

    2,048+7Star change over the last 7 days
  • FireRedASR@FireRedTeam

    Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.

    1,975+2Star change over the last 7 days
  • masr@nobody132

    中文语音识别; Mandarin Automatic Speech Recognition;

    1,968+0Star change over the last 7 days
  • julius@julius-speech

    Open-Source Large Vocabulary Continuous Speech Recognition Engine

    1,935+1Star change over the last 7 days
  • A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.

    1,894+1Star change over the last 7 days
  • alan-sdk-ios@alan-ai

    The Self-Coding System for Your App — Alan AI SDK for iOS

    1,878-3Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Android

    1,807-1Star change over the last 7 days
  • whisper-turbo@FL33TW00D

    Cross-Platform, GPU Accelerated Whisper 🏎️

    1,794+0Star change over the last 7 days
  • sherpa-ncnn@k2-fsa

    Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

    1,777-1Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Flutter

    1,756-1Star change over the last 7 days
  • alan-sdk-ionic@alan-ai

    The Self-Coding System for Your App — Alan AI SDK for Ionic

    1,649-1Star change over the last 7 days
  • dsnote@mkiol

    Speech Note Linux app. Note taking, reading and translating with offline Speech to Text, Text to Speech and Machine translation.

    1,622+8Star change over the last 7 days
  • OBS plugin for local speech recognition and captioning using AI

    1,597+8Star change over the last 7 days
  • amica@semperai

    Amica is an open source interface for interactive communication with 3D characters with voice synthesis and speech recognition.

    1,591-2Star change over the last 7 days
  • Custom nodes that extend the capabilities of Comfyui

    1,525+1Star change over the last 7 days
  • SALMONN@bytedance

    SALMONN family: A suite of advanced multi-modal LLMs

    1,521+3Star change over the last 7 days
  • Fun-ASR@QwenAudio

    Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

    1,512+10Star change over the last 7 days
  • SpeechT5@microsoft

    Unified-Modal Speech-Text Pre-Training for Spoken Language Processing

    1,449+0Star change over the last 7 days
← Back to topics