qwen
Tracked open-source repos tagged qwen, sorted by stars.
- #91
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
★ 631+48Star change over the last 7 days - #92
Welcome to the official repository of SINQ! A novel, fast and high-quality quantization method designed to make any Large Language Model smaller while preserving accuracy [ICML 2026]
★ 630+1Star change over the last 7 days - #93
ComfyUI DyPE+SEGA, enabling artifact-free 4K+ image generation: Z-Image, Qwen, Flux, Krea2
★ 623+2Star change over the last 7 days - #94
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
★ 621+0Star change over the last 7 days - #95★ 605+1Star change over the last 7 days
- #96
🔴 VERY LARGE AI TOOL LIST! 🔴 Curated list of AI Tools - Updated 2026
★ 580+2Star change over the last 7 days - #97★ 577—Star change over the last 7 days
- #98
💭 一个可二次开发 Chat Bot 单轮对话 Web 端 MVP 原型模板, 基于 Vue 3, Vite8, TypeScript, Naive UI, Pinia(v3), UnoCSS 等主流技术构建, 🧤简单集成大模型 API, 采用单轮 AI 问答对话模式, 每次提问独立响应, 无需上下文, 支持 SSE 打字机效果流式输出, 集成 markdown-it Mermaid/KaTex/LaTex 公式高亮预览, 星火, 智谱, 硅基流动, Deepseek V4/V3/R1 深度思考推理模型预览, 兼容 <think> 标签, 包含蒸馏 skill 💼 易于定制和快速搭建 Chat 类大语言模型产品 (附示例截图)
★ 575+1Star change over the last 7 days - #99
校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。
★ 573+0Star change over the last 7 days - #100
Run Qwen3-TTS text-to-speech locally on Mac (M1/M2/M3/M4). Voice cloning, voice design, custom voices. 100% offline using MLX.
★ 561+2Star change over the last 7 days - #101★ 560+0Star change over the last 7 days
- #102
FlashRT is a high-performance realtime inference engine for small-batch, latency-sensitive AI workloads. The flagship integration is production VLA control for Pi0, Pi0.5, GROOT N1.6, and Pi0-FAST. Also support llm e.g, qwen3.6-27B
★ 555+12Star change over the last 7 days - #103
[ICLR 2026] A Framework for LLM-based Multi-Agent Reinforced Training and Inference
★ 555+1Star change over the last 7 days - #104
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 550+18Star change over the last 7 days - #105
CLI-first orchestration for AI coding agents: run selected agents and models in parallel, compare raw answers, and coordinate review or implementation workflows.
★ 515+5Star change over the last 7 days - #106
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 508+3Star change over the last 7 days - #107
A curated list of awesome tools, extensions, and resources for Gemini CLI.
★ 504—Star change over the last 7 days