llm-serving
Tracked open-source repos tagged llm-serving, sorted by stars.
Related topics
Topics that frequently appear alongside llm-serving on the same repo.
Recent risers
Repos created in the last 90 days, tagged llm-serving.
No new repos tagged with this topic in the last 90 days.
- #1★ 90,817+402Star change over the last 7 days
- #2
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43,688+41Star change over the last 7 days - #3
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 24,995+21Star change over the last 7 days - #4
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★ 14,536+38Star change over the last 7 days - #5
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12,524+1Star change over the last 7 days - #6
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
★ 10,550+16Star change over the last 7 days - #7
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
★ 8,817+4Star change over the last 7 days - #8
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
★ 5,593+21Star change over the last 7 days - #9
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5,318+1Star change over the last 7 days - #10★ 3,827+1Star change over the last 7 days
- #11
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3,714+3Star change over the last 7 days - #12
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
★ 2,997-1Star change over the last 7 days - #13
Community maintained hardware plugin for vLLM on Ascend
★ 2,749+19Star change over the last 7 days - #14★ 2,174+1Star change over the last 7 days
- #15★ 2,076+0Star change over the last 7 days
- #16
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+33Star change over the last 7 days - #17
Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere
★ 1,372+4Star change over the last 7 days - #18★ 1,325+6Star change over the last 7 days
- #19★ 1,275+130Star change over the last 7 days
- #20★ 975+1Star change over the last 7 days
- #21★ 908+0Star change over the last 7 days
- #22
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
★ 902+0Star change over the last 7 days - #23
♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️
★ 805+1Star change over the last 7 days - #24
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 677+14Star change over the last 7 days - #25
LLM (Large Language Model) FineTuning
★ 579+1Star change over the last 7 days