Skip to main content
buildradar
Sign in
Topic · vllm

vllm

Tracked open-source repos tagged vllm, sorted by stars.

52 repos
  • prometheus-eval@prometheus-eval

    Evaluate your LLM's response with Prometheus and GPT4 💯

    1,111+4Star change over the last 7 days
  • tiny-vllm@jmaczan

    Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

    1,094+15Star change over the last 7 days
  • verl-omni@verl-project

    Multimodal RL training framework for diffusion & omni models

    959+64Star change over the last 7 days
  • UniRL@Tencent-Hunyuan

    UniRL is a Framework for Unified Multimodal Model Reinforcement Learning

    935+17Star change over the last 7 days
  • llm_note@harleyszhang

    LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

    889+1Star change over the last 7 days
  • A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

    881+77Star change over the last 7 days
  • local_ai_ocr@th1nhhdk

    An local, offline (after initial setup), portable OCR software that can process images and PDF files, using DeepSeek-OCR-2 AI (running directly on your machine).

    873+3Star change over the last 7 days
  • llm-server-docs@varunvasudeva1

    End-to-end documentation to set up your own local & fully private LLM server on Debian. Equipped with chat, web search, RAG, model management, MCP servers, image generation, and TTS.

    834-2Star change over the last 7 days
  • llmcord@jakobdylanc

    Make Discord your LLM frontend - Supports any OpenAI compatible API (OpenRouter, Ollama and more)

    829+4Star change over the last 7 days
  • BambooAI@pgalko

    A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.

    787+1Star change over the last 7 days
  • LightCompress@ModelTC

    [EMNLP 2024 & AAAI 2026] A powerful toolkit for compressing large models including LLMs, VLMs, and video generative models.

    747+4Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    676+18Star change over the last 7 days
  • vidur@microsoft

    Accurate, large-scale, and extensible simulator for LLM inference Systems

    673+5Star change over the last 7 days
  • efficientsam3@SimonZeng7108

    EfficientSAM3 compresses SAM3 into lightweight, edge-friendly models via progressive knowledge distillation for fast promptable concept segmentation and tracking.

    667+9Star change over the last 7 days
  • FlashTTS@HuiResearch

    基于SparkTTS、OrpheusTTS等模型,提供高质量中文语音合成与声音克隆服务。

    613+1Star change over the last 7 days
  • RoboBrain@FlagOpen

    [CVPR 2025] RoboBrain: A Unified Brain Model for Robotic Manipulation from Abstract to Concrete. Official Repository.

    564+0Star change over the last 7 days
  • crater@raids-lab

    Crater is a cloud-native AI training & inference platform.

    554+1Star change over the last 7 days
  • vllm-playground@micytao

    A modern web interface for managing and interacting with vLLM servers (www.github.com/vllm-project/vllm). Supports both GPU and CPU modes, with special optimizations for macOS Apple Silicon and enterprise deployment on OpenShift/Kubernetes.

    530+8Star change over the last 7 days
  • taOS@jaylfc

    Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).

    519+14Star change over the last 7 days
  • smg@smg-project

    Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.

    512Star change over the last 7 days
  • ome@ome-projects

    Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

    507Star change over the last 7 days
  • vllm-cli@Chen-zexi

    A command-line interface tool for serving LLM using vLLM.

    505-1Star change over the last 7 days
← Back to topics