Skip to main content
buildradar
Sign in
Topic · speech

speech

Tracked open-source repos tagged speech, sorted by stars.

74 repos
  • MARS5-TTS@Camb-ai

    MARS5 speech model (TTS) from CAMB.AI

    2,818+1Star change over the last 7 days
  • speechgpt@hahahumble

    💬 SpeechGPT is a web application that enables you to converse with ChatGPT.

    2,750-1Star change over the last 7 days
  • gTTS@pndurette

    Python library and CLI tool to interface with Google Translate's text-to-speech API

    2,628+0Star change over the last 7 days
  • ten-vad@TEN-framework

    Voice Activity Detector (VAD) : low-latency, high-performance and lightweight

    2,256+7Star change over the last 7 days
  • IMS-Toucan@DigitalPhonetics

    Controllable and fast Text-to-Speech for over 7000 languages!

    2,207-1Star change over the last 7 days
  • soloud@jarikomppa

    Free, easy, portable audio engine for games

    2,170+3Star change over the last 7 days
  • openai-edge-tts@travisvn

    Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs

    2,075+5Star change over the last 7 days
  • Praat: Doing Phonetics By Computer

    1,970+1Star change over the last 7 days
  • julius@julius-speech

    Open-Source Large Vocabulary Continuous Speech Recognition Engine

    1,935+1Star change over the last 7 days
  • Community list of startups working with AI in audio and music technology

    1,769+1Star change over the last 7 days
  • SALMONN@bytedance

    SALMONN family: A suite of advanced multi-modal LLMs

    1,521+3Star change over the last 7 days
  • voicefixer@haoheliu

    General Speech Restoration

    1,375+0Star change over the last 7 days
  • CrisperWhisper@nyrahealth

    Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.

    1,371+13Star change over the last 7 days
  • nlp-paper@DengBoCong

    自然语言处理领域下的相关论文(附阅读笔记),复现模型以及数据处理等(代码含TensorFlow和PyTorch两版本)

    1,333+1Star change over the last 7 days
  • MTools@HG-ha

    MTools 是一个功能强大的多功能桌面应用程序,集成了音视频处理、图片编辑、文本操作和编码工具,内置AI增强功能。旨在简化您的工作流程,提升生产效率

    1,305+5Star change over the last 7 days
  • StreamSpeech@ictnlp

    StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

    1,288+0Star change over the last 7 days
  • Deep-Learning-Experiments@roatienza

    Videos, notes and experiments to understand deep learning

    1,198+0Star change over the last 7 days
  • lhotse@lhotse-speech

    Tools for handling multimodal data in machine learning projects.

    1,149-1Star change over the last 7 days
  • conformer@sooftware

    [Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)

    1,132+1Star change over the last 7 days
  • pykaldi@pykaldi

    A Python wrapper for Kaldi

    1,040+1Star change over the last 7 days
  • FireRedTTS@FireRedTeam

    An Open-Sourced LLM-empowered Foundation TTS System

    920+1Star change over the last 7 days
  • inaSpeechSegmenter@ina-foss

    CNN-based audio segmentation toolkit. Allows to detect speech, music, noise and speaker gender. Has been designed for large scale gender equality studies based on speech time per gender.

    910+1Star change over the last 7 days
  • diffwave@lmnt-com

    DiffWave is a fast, high-quality neural vocoder and waveform synthesizer.

    885+0Star change over the last 7 days
  • AnyGPT@OpenMOSS

    A unified multimodal language model based on discrete sequence modeling

    881-1Star change over the last 7 days
  • PPASR@yeyupiaoling

    基于PaddlePaddle实现端到端中文语音识别,从入门到实战,超简单的入门案例,超实用的企业项目。支持当前最流行的DeepSpeech2、Conformer、Squeezeformer模型

    870-1Star change over the last 7 days
  • Voxtral ASR & TTS running natively and in the browser. A Rust implementation of Mistral's Voxtral mini realtime ASR / TTS using the Burn ML framework

    819+2Star change over the last 7 days
  • AlphaAvatar@AlphaAvatar

    A real-time interactive Omni Avatar built on LiveKit, which allows you to seamlessly integrate with any open source Avatar components (real-time model, visual, voice, memory, search, etc.).

    813+4Star change over the last 7 days
  • XR3Player@goxr3plus

    🎧 🎼 The MOST ADVANCED JavaFX Media Player

    771+0Star change over the last 7 days
  • cboard@cboard-org

    Augmentative and Alternative Communication (AAC) system with text-to-speech for the browser

    755+9Star change over the last 7 days
  • allosaurus@xinjli

    Allosaurus is a pretrained universal phone recognizer for more than 2000 languages

    746+4Star change over the last 7 days
← Back to topics