Skip to main content
buildradar
Sign in
Topic · vllm

vllm

Tracked open-source repos tagged vllm, sorted by stars.

Repos
52
Total stars
195,358
Avg. stars
3,757
Share
0.01%

Topics that frequently appear alongside vllm on the same repo.

Recent risers

Repos created in the last 90 days, tagged vllm.

  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,209
  • A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

    882
  • kaggle-tpu-lab@ARahim3

    Qwen3.8-27B (bf16) on a free Kaggle TPU: OpenAI-compatible endpoint, 262k context, ~130 tok/s, works with Claude Code, Codex, Opencode and Pi.

    146
  • FunASR@modelscope

    Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

    20,136+64Star change over the last 7 days
  • llama-cookbook@meta-llama

    Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services

    18,556-3Star change over the last 7 days
  • ✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记

    13,226+6Star change over the last 7 days
  • AI-Research-SKILLs@Orchestra-Research

    Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

    12,256+104Star change over the last 7 days
  • LMCache@LMCache

    LMCache: Supercharge Your LLM with the Fastest KV Cache Layer

    11,622+60Star change over the last 7 days
  • OpenRLHF@OpenRLHF

    An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)

    9,969+9Star change over the last 7 days
  • inference@xorbitsai

    Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all through one unified, production-ready inference API.

    9,537+9Star change over the last 7 days
  • dynamo@ai-dynamo

    A Datacenter Scale Distributed Inference Serving Framework

    7,952+44Star change over the last 7 days
  • Mooncake@kvcache-ai

    Mooncake is the serving platform for Kimi, a leading LLM service provided by Moonshot AI.

    6,470+40Star change over the last 7 days
  • kserve@kserve

    Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

    5,853+13Star change over the last 7 days
  • UltraRAG@OpenBMB

    A Low-Code MCP Framework for Building Complex and Innovative RAG Pipelines

    5,678+4Star change over the last 7 days
  • gpustack@gpustack

    A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

    5,593+21Star change over the last 7 days
  • llama-swap@mostlygeek

    Reliable model swapping for any local OpenAI/Anthropic compatible server - llama.cpp, vllm, etc

    5,562+49Star change over the last 7 days
  • semantic-router@vllm-project

    A programmable Mixture-of-Models router for heterogeneous LLM inference

    5,506+101Star change over the last 7 days
  • Awesome-LLM-Inference@xlite-dev

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    5,482+3Star change over the last 7 days
  • sparrow@katanaml

    Structured data extraction, instruction calling and agentic workflows with ML, LLM and Vision LLM

    5,214+7Star change over the last 7 days
  • tiny-llm@skyzh

    learn LLM inference system on Apple Silicon for systems engineers: build a tiny vLLM + Qwen

    4,537+8Star change over the last 7 days
  • cascadeflow@lemony-ai

    Cascading runtime for AI agents. Optimize cost, latency, quality, and policy decisions inside the agent loop.

    3,970-9Star change over the last 7 days
  • FastDeploy@PaddlePaddle

    High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

    3,714+3Star change over the last 7 days
  • ramalama@containers

    RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.

    3,031+6Star change over the last 7 days
  • vllm-ascend@vllm-project

    Community maintained hardware plugin for vLLM on Ascend

    2,749+19Star change over the last 7 days
  • local-studio@sybil-solutions

    Control panel for VLLM, Sglang, llama.cpp, exllamav3

    1,749+7Star change over the last 7 days
  • InferenceX@SemiAnalysisAI

    Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

    1,613+33Star change over the last 7 days
  • auto-round@intel

    A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包

    1,599+8Star change over the last 7 days
  • vllm-mlx@waybarrios

    High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.

    1,557+6Star change over the last 7 days
  • kvcached@ovg-project

    Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

    1,275+130Star change over the last 7 days
  • kubeai@kubeai-project

    AI Inference Operator for Kubernetes. The easiest way to serve ML models in production. Supports VLMs, LLMs, embeddings, and speech-to-text.

    1,256+2Star change over the last 7 days
  • GPTQModel@ModelCloud

    LLM model quantization (compression) toolkit with HW acceleration support for Nvidia, AMD, Intel GPU and Intel/AMD/Apple CPU via HF, vLLM, and SGLang.

    1,248+0Star change over the last 7 days
  • BricksLLM@bricks-cloud

    🔒 Enterprise-grade API gateway that helps you monitor and impose cost or rate limits per API key. Get fine-grained access control and monitoring per user, application, or environment. Supports OpenAI, Azure OpenAI, Anthropic, vLLM, and open-source LLMs.

    1,228-1Star change over the last 7 days
  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,209+316Star change over the last 7 days
← Back to topics