audio-language-model
Tracked open-source repos tagged audio-language-model, sorted by stars.
Related topics
Topics that frequently appear alongside audio-language-model on the same repo.
Recent risers
Repos created in the last 90 days, tagged audio-language-model.
No new repos tagged with this topic in the last 90 days.
- #1
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
★ 1,512+10Star change over the last 7 days - #2
We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities.
★ 1,113+0Star change over the last 7 days - #3
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
★ 677+0Star change over the last 7 days - #4
Worlds first open-source real-time end-to-end spoken dialogue model with personalized voice cloning.
★ 549-1Star change over the last 7 days