Skip to main content
buildradar
Sign in
Topic · llm-serving

llm-serving

Tracked open-source repos tagged llm-serving, sorted by stars.

Repos
25
Total stars
244,794
Avg. stars
9,792
Share
0.01%

Topics that frequently appear alongside llm-serving on the same repo.

Recent risers

Repos created in the last 90 days, tagged llm-serving.

No new repos tagged with this topic in the last 90 days.

  • vllm@vllm-project

    A high-throughput and memory-efficient inference and serving engine for LLMs

    90,817+402Star change over the last 7 days
  • ray@ray-project

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    43,688+41Star change over the last 7 days
  • llm-action@liguodongiot

    本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)

    24,995+21Star change over the last 7 days
  • TensorRT-LLM@NVIDIA

    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

    14,536+38Star change over the last 7 days
  • OpenLLM@bentoml

    Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.

    12,524+1Star change over the last 7 days
  • skypilot@skypilot-org

    The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

    10,550+16Star change over the last 7 days
  • BentoML@bentoml

    The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

    8,817+4Star change over the last 7 days
  • gpustack@gpustack

    A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.

    5,593+21Star change over the last 7 days
  • superduper@superduper-io

    Superduper: End-to-end framework for building custom AI applications and agents.

    5,318+1Star change over the last 7 days
  • lorax@predibase

    Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

    3,827+1Star change over the last 7 days
  • FastDeploy@PaddlePaddle

    High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle

    3,714+3Star change over the last 7 days
  • chitu@thu-pacman

    High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

    2,997-1Star change over the last 7 days
  • vllm-ascend@vllm-project

    Community maintained hardware plugin for vLLM on Ascend

    2,749+19Star change over the last 7 days
  • MoBA@MoonshotAI

    MoBA: Mixture of Block Attention for Long-Context LLMs

    2,174+1Star change over the last 7 days
  • aici@microsoft

    AICI: Prompts as (Wasm) Programs

    2,076+0Star change over the last 7 days
  • InferenceX@SemiAnalysisAI

    Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

    1,613+33Star change over the last 7 days
  • parallax@GradientHQ

    Parallax is a distributed model serving framework that lets you build your own AI cluster anywhere

    1,372+4Star change over the last 7 days
  • rtp-llm@alibaba

    RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

    1,325+6Star change over the last 7 days
  • kvcached@ovg-project

    Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

    1,275+130Star change over the last 7 days
  • Nanoflow@efeslab

    A throughput-oriented high-performance serving framework for LLMs

    975+1Star change over the last 7 days
  • ZhiLight@zhihu

    A highly optimized LLM inference acceleration engine for Llama and its variants.

    908+0Star change over the last 7 days
  • mosec@mosecorg

    A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

    902+0Star change over the last 7 days
  • helix@helixml

    ♾️ Private Agent Fleet with Spec Coding. Each agent gets their own GPU-accelerated desktop. Run Claude, Codex, Gemini and open models on a full private AI Stack ♾️

    805+1Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    677+14Star change over the last 7 days
  • LLM-FineTuning-Large-Language-Models@rohan-paul

    LLM (Large Language Model) FineTuning

    579+1Star change over the last 7 days
← Back to topics