vllm
標記 vllm 主題、收錄中的開源專案,依星數排序。
- #31
使用 Prometheus 與 GPT4 評估你的 LLM 回應 💯
★ 1,114+7近 7 天星數變化 - #32★ 1,100+10近 7 天星數變化
- #33★ 983+49近 7 天星數變化
- #34★ 939+15近 7 天星數變化
- #35★ 890+1近 7 天星數變化
- #36
為期 10 週、每天 30 分鐘的 LLM 推論服務與最佳化學習藍圖。包含 vLLM、SGLang、量化、推測解碼、效能評測。
★ 887+21近 7 天星數變化 - #37
一款本機、離線(初始設定後)、可攜式的 OCR 軟體,可使用 DeepSeek-OCR-2 AI(直接在您的機器上執行)處理圖片和 PDF 檔案。
★ 874+2近 7 天星數變化 - #38
在 Debian 上建置您自己的本地端且完全私有 LLM 伺服器的端到端文件。配備聊天、網路搜尋、RAG、模型管理、MCP 伺服器、圖像生成以及 TTS。
★ 833-3近 7 天星數變化 - #39★ 830+1近 7 天星數變化
- #40★ 788+1近 7 天星數變化
- #41
[EMNLP 2024 & AAAI 2026] 一個強大的大型模型壓縮工具包,包含 LLM、VLM 及影片生成模型。
★ 747+3近 7 天星數變化 - #42★ 682+11近 7 天星數變化
- #43★ 674+3近 7 天星數變化
- #44
EfficientSAM3 透過漸進式知識蒸餾將 SAM3 壓縮為輕量且適合邊緣裝置的模型,以實現快速的提示式概念分割與追蹤。
★ 671+10近 7 天星數變化 - #45★ 614+2近 7 天星數變化
- #46★ 564+0近 7 天星數變化
- #47★ 555+1近 7 天星數變化
- #48
用於管理與互動 vLLM 伺服器的現代化網頁介面。支援 GPU 與 CPU 模式,並針對 macOS Apple Silicon 與 OpenShift/Kubernetes 企業部署進行特別優化。
★ 531+3近 7 天星數變化 - #49
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
★ 522+14近 7 天星數變化 - #50
Engine-agnostic LLM gateway in Rust. Full OpenAI & Anthropic API compatibility across vLLM, TRT-LLM, TokenSpeed, SGLang, OpenAI, Gemini & more. Industry-first gRPC pipeline, KV cache-aware routing, chat history, tokenization caching, Responses API, embeddings, WASM plugins, MCP, and multi-tenant auth.
★ 517+14近 7 天星數變化 - #51
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 508+4近 7 天星數變化 - #52★ 505-1近 7 天星數變化
- #53
sparkrun - launch, manage, and stop LLM inference workloads on NVIDIA DGX Spark systems. Live support: https://discord.com/invite/GH5kRgv6ZD
★ 504—近 7 天星數變化