vllm
Tracked open-source repos tagged vllm, sorted by stars.
- #31
Evaluate your LLM's response with Prometheus and GPT4 💯
★ 1,111+4Star change over the last 7 days - #32
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1,094+15Star change over the last 7 days - #33★ 959+64Star change over the last 7 days
- #34★ 935+17Star change over the last 7 days
- #35
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
★ 889+1Star change over the last 7 days - #36
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
★ 881+77Star change over the last 7 days - #37
An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).
★ 873+3Star change over the last 7 days - #38
End-to-end documentation to set up your own local & fully private LLM server on Debian. Equipped with chat, web search, RAG, model management, MCP servers, image generation, and TTS.
★ 834-2Star change over the last 7 days - #39
Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more)
★ 829+4Star change over the last 7 days - #40
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
★ 787+1Star change over the last 7 days - #41
[EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.
★ 747+4Star change over the last 7 days - #42
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 676+18Star change over the last 7 days - #43★ 673+5Star change over the last 7 days
- #44
EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.
★ 667+9Star change over the last 7 days - #45★ 613+1Star change over the last 7 days
- #46
[CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository.
★ 564+0Star change over the last 7 days - #47★ 554+1Star change over the last 7 days
- #48
A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.
★ 530+8Star change over the last 7 days - #49
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
★ 519+14Star change over the last 7 days - #50
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
★ 512—Star change over the last 7 days - #51
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 507—Star change over the last 7 days - #52★ 505-1Star change over the last 7 days