Skip to main content
buildradar
Sign in
Topic · speech-recognition

speech-recognition

Tracked open-source repos tagged speech-recognition, sorted by stars.

156 repos
  • 💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies

    1,399+0Star change over the last 7 days
  • mlx-tune@ARahim3

    Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.

    1,398+8Star change over the last 7 days
  • CrisperWhisper@nyrahealth

    Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

    1,371+13Star change over the last 7 days
  • whisper-ctranslate2@Softcatala

    Whisper command line client compatible with original OpenAI client based on CTranslate2.

    1,346+2Star change over the last 7 days
  • J.A.R.V.I.S@GauravSingh9356

    Personal Assistant built using python libraries. It does almost anything which includes sending emails, Optical Text Recognition, Dynamic News Reporting at any time with API integration, Todo list generator, Opens any website with just a voice command, Plays Music, Wikipedia searching, Dictionary with Intelligent Sensing i.e. auto spell checking, Weather Reporting i.e. temp, wind speed, humidity, YouTube searching, Google Map searching, Youtube Downloading, etc.

    1,341+5Star change over the last 7 days
  • voxtype@peteonrails

    Voice-to-text with push-to-talk for Wayland compositors

    1,338+42Star change over the last 7 days
  • StreamSpeech@ictnlp

    StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

    1,288+0Star change over the last 7 days
  • pruna@PrunaAI

    Pruna is a model optimization framework built for developers, enabling you to deliver faster, more efficient models with minimal overhead.

    1,274+1Star change over the last 7 days
  • vosk-server@alphacep

    WebSocket, gRPC and WebRTC speech recognition server based on Vosk and Kaldi libraries

    1,262+1Star change over the last 7 days
  • Whisper-Finetune@yeyupiaoling

    Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment

    1,222-1Star change over the last 7 days
  • quillman@modal-labs

    A voice chat app

    1,212-1Star change over the last 7 days
  • Treasure-of-Transformers@ashishpatel26

    💁 Awesome Treasure of Transformers Models for Natural Language processing contains papers, videos, blogs, official repo along with colab Notebooks. 🛫☑️

    1,179+0Star change over the last 7 days
  • A practical lab for building, testing, and evaluating apps with Apple's Foundation Models framework.

    1,178+1Star change over the last 7 days
  • lotti@matthiasn

    A private logbook with a staff of personal AI assistants. Agents read what you record and propose what to do next — you approve the changes. End-to-end encrypted sync between your own devices — servers only ever see ciphertext. Local AI optional.

    1,168+3Star change over the last 7 days
  • speech-swift@soniqo

    AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

    1,161+4Star change over the last 7 days
  • Ирина - русский голосовой ассистент для работы оффлайн. Поддерживает скиллы через плагины.

    1,153+0Star change over the last 7 days
  • lhotse@lhotse-speech

    Tools for handling multimodal data in machine learning projects.

    1,149-1Star change over the last 7 days
  • conformer@sooftware

    [Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)

    1,132+1Star change over the last 7 days
  • The Self-Coding System for Your App — Alan AI SDK for Cordova

    1,132+0Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    1,118+130Star change over the last 7 days
  • AI-Waifu-Vtuber@ardha27

    AI Vtuber for Streaming on Youtube/Twitch

    1,115+1Star change over the last 7 days
  • Whisperboard@Saik0s

    The open-source iOS app that's making quality voice transcription more accessible on mobile devices.

    1,102+2Star change over the last 7 days
  • whisper-writer@savbell

    💬📝 A small dictation app using OpenAI's Whisper speech recognition model.

    1,100+2Star change over the last 7 days
  • Real-time full-duplex speech recognition server, based on the Kaldi toolkit and the GStreamer framwork.

    1,094+1Star change over the last 7 days
  • Offline speech recognition for Android with Vosk library.

    1,056+0Star change over the last 7 days
  • pykaldi@pykaldi

    A Python wrapper for Kaldi

    1,040+1Star change over the last 7 days
  • TensorFlowASR@TensorSpeech

    :zap: TensorFlowASR: Almost State-of-the-art Automatic Speech Recognition in Tensorflow 2. Supported languages that can use characters or subwords

    1,010+0Star change over the last 7 days
  • StoryToolkitAI@octimot

    An editing tool that uses AI to transcribe, understand content and search for anything in your footage, integrated with ChatGPT and other AI models

    1,009+2Star change over the last 7 days
  • sherpa@k2-fsa

    Speech-to-text server framework with next-gen Kaldi

    984+4Star change over the last 7 days
  • VoiceStreamAI@alesaccoia

    Near-Realtime audio transcription using self-hosted Whisper and WebSocket in Python/JS

    960+1Star change over the last 7 days
← Back to topics