speech-processing
Tracked open-source repos tagged speech-processing, sorted by stars.
Related topics
Topics that frequently appear alongside speech-processing on the same repo.
Recent risers
Repos created in the last 90 days, tagged speech-processing.
No new repos tagged with this topic in the last 90 days.
- #1
A PyTorch-based Speech Toolkit
★ 11,792+22Star change over the last 7 days - #2
Neural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
★ 10,484+38Star change over the last 7 days - #3
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
★ 10,083+52Star change over the last 7 days - #4
Become a cracked AI/ML researcher/engineer with this unconventional textbook covering maths, computing, and ML with intuition.
★ 7,379+33Star change over the last 7 days - #5
Reading list for research topics in multimodal machine learning
★ 6,925+0Star change over the last 7 days - #6
Foundation Architecture for (M)LLMs
★ 3,138+1Star change over the last 7 days - #7
Multilingual Automatic Speech Recognition with word-level timestamps and confidence
★ 2,841+2Star change over the last 7 days - #8
AI powered speech denoising and enhancement
★ 2,401+6Star change over the last 7 days - #9★ 2,250+4Star change over the last 7 days
- #10
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2,208+0Star change over the last 7 days - #11
A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources.
★ 1,893+2Star change over the last 7 days - #12
💎 A list of accessible speech corpora for ASR, TTS, and other Speech Technologies
★ 1,399+0Star change over the last 7 days - #13
General Speech Restoration
★ 1,376+3Star change over the last 7 days - #14
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
★ 1,362+21Star change over the last 7 days - #15
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+2Star change over the last 7 days - #16★ 1,146+0Star change over the last 7 days
- #17
A Framework for Speech, Language, Audio, Music Processing with Large Language Model
★ 1,056+0Star change over the last 7 days - #18
You can find the speech algorithms you want here
★ 878+1Star change over the last 7 days - #19
[NeurIPS 2021] Multiscale Benchmarks for Multimodal Representation Learning
★ 636+0Star change over the last 7 days - #20
语音方向实验室/公司/资源/实习等,欢迎推荐或自荐
★ 609+0Star change over the last 7 days