sglang
Tracked open-source repos tagged sglang, sorted by stars.
Related topics
Topics that frequently appear alongside sglang on the same repo.
Recent risers
Repos created in the last 90 days, tagged sglang.
- #1
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
★ 880
- #1
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.
★ 9,537+9Star change over the last 7 days - #2★ 7,952+44Star change over the last 7 days
- #3
Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.
★ 6,470+40Star change over the last 7 days - #4
OpenClaw-RL: Train any agent simply by talking
★ 5,665+7Star change over the last 7 days - #5
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
★ 5,593+21Star change over the last 7 days - #6
Control panel for VLLM, Sglang, llama.cpp, exllamav3
★ 1,749+7Star change over the last 7 days - #7
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+33Star change over the last 7 days - #8
A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包
★ 1,599+8Star change over the last 7 days - #9
A multilingual model for long-form, multi-speaker dialogue synthesis with flexible speaker control and zero-shot voice cloning
★ 1,394+4Star change over the last 7 days - #10★ 1,275+130Star change over the last 7 days
- #11
LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.
★ 1,248+0Star change over the last 7 days - #12
Train speculative decoding models effortlessly and port them smoothly to SGLang serving.
★ 1,144+14Star change over the last 7 days - #13★ 1,108+3Star change over the last 7 days
- #14
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,105+130Star change over the last 7 days - #15★ 933+15Star change over the last 7 days
- #16
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
★ 880+76Star change over the last 7 days - #17★ 613+1Star change over the last 7 days
- #18
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
★ 510—Star change over the last 7 days - #19
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 506—Star change over the last 7 days