Skip to main content
buildradar
Sign in
Topic · model-serving

model-serving

Tracked open-source repos tagged model-serving, sorted by stars.

Repos
29
Total stars
164,226
Avg. stars
5,663
Share
0.01%

Topics that frequently appear alongside model-serving on the same repo.

Recent risers

Repos created in the last 90 days, tagged model-serving.

No new repos tagged with this topic in the last 90 days.

  • vllm@vllm-project

    A high-throughput and memory-efficient inference and serving engine for LLMs

    90,415+745Star change over the last 7 days
  • BentoML@bentoml

    The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!

    8,813+18Star change over the last 7 days
  • vllm-omni@vllm-project

    A framework for efficient model inference with omni-modality models

    6,457+229Star change over the last 7 days
  • kserve@kserve

    Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes

    5,840+25Star change over the last 7 days
  • Olares@beclab

    Open-Source Personal Cloud OS for Always-On Agents

    5,243+10Star change over the last 7 days
  • In this repository, I will share some useful notes and references about deploying deep learning-based models in production.

    4,376+1Star change over the last 7 days
  • 🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for Machine Learning, LLM (Large Language Model), GenAI (Generative AI). 🍻 OSDI, NSDI, SIGCOMM, SoCC, MLSys, etc. 🗃️ Llama3, Mistral, etc. 🧑‍💻 Video Tutorials.

    4,318+27Star change over the last 7 days
  • LightLLM@ModelTC

    LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.

    4,251+21Star change over the last 7 days
  • FedML@FedML-AI

    FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.

    4,061-1Star change over the last 7 days
  • lorax@predibase

    Multi-LoRA inference server that scales to 1000s of fine-tuned LLMs

    3,826-2Star change over the last 7 days
  • chitu@thu-pacman

    High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.

    2,998+5Star change over the last 7 days
  • vllm-ascend@vllm-project

    Community maintained hardware plugin for vLLM on Ascend

    2,730+53Star change over the last 7 days
  • openlake@openlake-project

    OpenLake is a high performance storage engine for efficient LLM inference and GPU Training

    2,346+9Star change over the last 7 days
  • envd@tensorchord

    🏕️ Reproducible development environment for humans and agents

    2,227+3Star change over the last 7 days
  • aici@microsoft

    AICI: Prompts as (Wasm) Programs

    2,076+0Star change over the last 7 days
  • mlrun@mlrun

    MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.

    1,694+0Star change over the last 7 days
  • kitops@kitops-ml

    An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.

    1,409+5Star change over the last 7 days
  • rtp-llm@alibaba

    RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

    1,320+6Star change over the last 7 days
  • hopsworks@logicalclocks

    Hopsworks - Data-Intensive AI platform with a Feature Store

    1,303+1Star change over the last 7 days
  • truss@basetenlabs

    The simplest way to serve AI/ML models in production

    1,196+1Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    978+70Star change over the last 7 days
  • Nanoflow@efeslab

    A throughput-oriented high-performance serving framework for LLMs

    974+0Star change over the last 7 days
  • model_server@openvinotoolkit

    A scalable inference server for models optimized with OpenVINO™

    921+3Star change over the last 7 days
  • ZhiLight@zhihu

    A highly optimized LLM inference acceleration engine for Llama and its variants.

    908+0Star change over the last 7 days
  • mosec@mosecorg

    A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine

    902+0Star change over the last 7 days
  • ServerlessLLM@ServerlessLLM

    Serverless LLM Serving for Everyone.

    711+6Star change over the last 7 days
  • timber@kossisoroyce

    Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.

    687-1Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    660+4Star change over the last 7 days
  • fastapi-ml-skeleton@eightBEC

    FastAPI Skeleton App to serve machine learning models production-ready.

    603-1Star change over the last 7 days
← Back to topics