Skip to main content
buildradar
Sign in
Topic · voice-activity-detection

voice-activity-detection

Tracked open-source repos tagged voice-activity-detection, sorted by stars.

Repos
23
Total stars
83,655
Avg. stars
3,637
Share
0.00%

Topics that frequently appear alongside voice-activity-detection on the same repo.

Recent risers

Repos created in the last 90 days, tagged voice-activity-detection.

No new repos tagged with this topic in the last 90 days.

  • FunASR@modelscope

    Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

    20,136+64Star change over the last 7 days
  • pyannote-audio@pyannote

    Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding

    10,500+16Star change over the last 7 days
  • NoiseTorch@noisetorch

    Real-time microphone noise suppression on Linux.

    10,315+2Star change over the last 7 days
  • silero-vad@snakers4

    Silero VAD: pre-trained enterprise-grade Voice Activity Detector

    10,112+29Star change over the last 7 days
  • ffsubsync@smacke

    Automagically synchronize subtitles with video.

    7,864+7Star change over the last 7 days
  • FluidAudio@FluidInference

    Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

    2,723+13Star change over the last 7 days
  • pluely@iamsrikanthnani

    The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for native performance, just 10MB. Completely undetectable in video calls, screen shares, and recordings.

    2,619+12Star change over the last 7 days
  • ten-vad@TEN-framework

    Voice Activity Detector (VAD) : low-latency, high-performance and lightweight

    2,256+7Star change over the last 7 days
  • voice_datasets@jim-schwoebel

    🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).

    2,223+1Star change over the last 7 days
  • vad@ricky0123

    Voice activity detector (VAD) for the browser with a simple API

    2,046+1Star change over the last 7 days
  • diart@juanmc2005

    A python package to build AI-powered real-time audio applications

    2,024+1Star change over the last 7 days
  • sherpa-ncnn@k2-fsa

    Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.

    1,777-1Star change over the last 7 days
  • 💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies

    1,399+0Star change over the last 7 days
  • speech-swift@soniqo

    AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML

    1,161+4Star change over the last 7 days
  • Python AI assistant 🧠

    1,014+0Star change over the last 7 days
  • whisper.net@sandrohanea

    Whisper.net. Speech to text made simple using Whisper Models

    941+1Star change over the last 7 days
  • inaSpeechSegmenter@ina-foss

    CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.

    910+1Star change over the last 7 days
  • auditok@amsehili

    An voice activity detection and audio segmentation tool

    860+1Star change over the last 7 days
  • FireRedASR2S@FireRedTeam

    A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.

    672+13Star change over the last 7 days
  • WhisperS2T@shashikg

    An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine

    578+1Star change over the last 7 days
  • FireRedVAD@FireRedTeam

    A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD

    521+6Star change over the last 7 days
  • subaligner@baxtree

    Automatically synchronize and translate subtitles, or create new ones by transcribing, using pre-trained DNNs, Forced Alignments and Transformers. https://subaligner.readthedocs.io/

    509+0Star change over the last 7 days
  • android-vad@gkonovalov

    Android Voice Activity Detection (VAD) library. Supports WebRTC VAD GMM, Silero VAD DNN, Yamnet VAD DNN models.

    506+1Star change over the last 7 days
← Back to topics