text-to-audio
Tracked open-source repos tagged text-to-audio, sorted by stars.
Related topics
Topics that frequently appear alongside text-to-audio on the same repo.
Recent risers
Repos created in the last 90 days, tagged text-to-audio.
No new repos tagged with this topic in the last 90 days.
- #1
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
★ 10,276+1Star change over the last 7 days - #2★ 5,831+45Star change over the last 7 days
- #3
[CVPR 2025] MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis
★ 2,264+0Star change over the last 7 days - #4
[NeurIPS 2025] PyTorch implementation of [ThinkSound], a unified framework for generating audio from any modality, guided by Chain-of-Thought (CoT) reasoning.
★ 1,378+0Star change over the last 7 days - #5
StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.
★ 1,288+0Star change over the last 7 days - #6
A webui for different audio related Neural Networks
★ 1,246+0Star change over the last 7 days - #7★ 1,238-1Star change over the last 7 days
- #8
HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation.
★ 1,056+4Star change over the last 7 days - #9
[ICLR 2026] TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching
★ 882+0Star change over the last 7 days - #10
PyTorch Implementation of Make-An-Audio (ICML'23) with a Text-to-Audio Generative Model
★ 668+0Star change over the last 7 days - #11★ 631+1Star change over the last 7 days
- #12
🔥🔥🔥 A curated list of papers on LLMs-based multimodal generation (image, video, 3D and audio).
★ 552+0Star change over the last 7 days - #13
An Open-Source Project to Unify Audio Processing and Generation
★ 512+1Star change over the last 7 days