voice-activity-detection
Tracked open-source repos tagged voice-activity-detection, sorted by stars.
Related topics
Topics that frequently appear alongside voice-activity-detection on the same repo.
Recent risers
Repos created in the last 90 days, tagged voice-activity-detection.
No new repos tagged with this topic in the last 90 days.
- #1
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
★ 20,136+64Star change over the last 7 days - #2
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
★ 10,500+16Star change over the last 7 days - #3
Real-time microphone noise suppression on Linux.
★ 10,315+2Star change over the last 7 days - #4
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
★ 10,112+29Star change over the last 7 days - #5★ 7,864+7Star change over the last 7 days
- #6
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
★ 2,723+13Star change over the last 7 days - #7
The Open Source Alternative to Cluely - A lightning-fast, privacy-first AI assistant that works seamlessly during meetings, interviews, and conversations without anyone knowing. Built with Tauri for native performance, just 10MB. Completely undetectable in video calls, screen shares, and recordings.
★ 2,619+12Star change over the last 7 days - #8★ 2,256+7Star change over the last 7 days
- #9
🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).
★ 2,223+1Star change over the last 7 days - #10★ 2,046+1Star change over the last 7 days
- #11★ 2,024+1Star change over the last 7 days
- #12
Real-time speech recognition and voice activity detection (VAD) using next-gen Kaldi with ncnn without Internet connection. Support iOS, Android, Linux, macOS, Windows, Raspberry Pi, VisionFive2, LicheePi4A etc.
★ 1,777-1Star change over the last 7 days - #13
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #14
AI speech toolkit for Apple Silicon — ASR, TTS, speech-to-speech, VAD, and diarization powered by MLX and CoreML
★ 1,161+4Star change over the last 7 days - #15
Python AI assistant 🧠
★ 1,014+0Star change over the last 7 days - #16
Whisper.net. Speech to text made simple using Whisper Models
★ 941+1Star change over the last 7 days - #17
CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.
★ 910+1Star change over the last 7 days - #18★ 860+1Star change over the last 7 days
- #19
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
★ 672+13Star change over the last 7 days - #20
An Optimized Speech-to-Text Pipeline for the Whisper Model Supporting Multiple Inference Engine
★ 578+1Star change over the last 7 days - #21
A SOTA Industrial-Grade Voice Activity Detection & Audio Event Detection, supporting 100+ languages, outperforming Silero-VAD, TEN-VAD, FunASR-VAD and WebRTC-VAD
★ 521+6Star change over the last 7 days - #22
Automatically synchronize and translate subtitles, or create new ones by transcribing, using pre-trained DNNs, Forced Alignments and Transformers. https://subaligner.readthedocs.io/
★ 509+0Star change over the last 7 days - #23
Android Voice Activity Detection (VAD) library. Supports WebRTC VAD GMM, Silero VAD DNN, Yamnet VAD DNN models.
★ 506+1Star change over the last 7 days