speech
Tracked open-source repos tagged speech, sorted by stars.
- #31★ 2,818+1Star change over the last 7 days
- #32★ 2,750-1Star change over the last 7 days
- #33★ 2,628+0Star change over the last 7 days
- #34★ 2,256+7Star change over the last 7 days
- #35
Controllable and fast Text-to-Speech for over 7000 languages!
★ 2,207-1Star change over the last 7 days - #36★ 2,170+3Star change over the last 7 days
- #37
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2,075+5Star change over the last 7 days - #38
Praat: Doing Phonetics By Computer
★ 1,970+1Star change over the last 7 days - #39★ 1,935+1Star change over the last 7 days
- #40
Community list of startups working with AI in audio and music technology
★ 1,769+1Star change over the last 7 days - #41★ 1,521+3Star change over the last 7 days
- #42
General Speech Restoration
★ 1,375+0Star change over the last 7 days - #43
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
★ 1,371+13Star change over the last 7 days - #44★ 1,333+1Star change over the last 7 days
- #45★ 1,305+5Star change over the last 7 days
- #46
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+0Star change over the last 7 days - #47
Videos, notes and experiments to understand deep learning
★ 1,198+0Star change over the last 7 days - #48★ 1,149-1Star change over the last 7 days
- #49
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
★ 1,132+1Star change over the last 7 days - #50★ 1,040+1Star change over the last 7 days
- #51
An Open-Sourced LLM-empowered Foundation TTS System
★ 920+1Star change over the last 7 days - #52
CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.
★ 910+1Star change over the last 7 days - #53★ 885+0Star change over the last 7 days
- #54★ 881-1Star change over the last 7 days
- #55
基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型
★ 870-1Star change over the last 7 days - #56
Voxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn ML framework
★ 819+2Star change over the last 7 days - #57
A real-time interactive Omni Avatar built on LiveKit, which allows you to seamlessly integrate with any open source Avatar components (real-time model, visual, voice, memory, search, etc.).
★ 813+4Star change over the last 7 days - #58★ 771+0Star change over the last 7 days
- #59
Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser
★ 755+9Star change over the last 7 days - #60
Allosaurus is a pretrained universal phone recognizer for more than 2000 languages
★ 746+4Star change over the last 7 days