model-serving
Tracked open-source repos tagged model-serving, sorted by stars.
Related topics
Topics that frequently appear alongside model-serving on the same repo.
Recent risers
Repos created in the last 90 days, tagged model-serving.
No new repos tagged with this topic in the last 90 days.
- #1★ 90,415+745Star change over the last 7 days
- #2
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
★ 8,813+18Star change over the last 7 days - #3★ 6,457+229Star change over the last 7 days
- #4
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
★ 5,840+25Star change over the last 7 days - #5★ 5,243+10Star change over the last 7 days
- #6
In this repository, I will share some useful notes and references about deploying deep learning-based models in production.
★ 4,376+1Star change over the last 7 days - #7
🚀 Awesome System for Machine Learning ⚡️ AI System Papers and Industry Practice. ⚡️ System for Machine Learning, LLM (Large Language Model), GenAI (Generative AI). 🍻 OSDI, NSDI, SIGCOMM, SoCC, MLSys, etc. 🗃️ Llama3, Mistral, etc. 🧑💻 Video Tutorials.
★ 4,318+27Star change over the last 7 days - #8
LightLLM is a Python-based LLM (Large Language Model) inference and serving framework, notable for its lightweight design, easy scalability, and high-speed performance.
★ 4,251+21Star change over the last 7 days - #9
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
★ 4,061-1Star change over the last 7 days - #10★ 3,826-2Star change over the last 7 days
- #11
High-performance inference framework for large language models, focusing on efficiency, flexibility, and availability.
★ 2,998+5Star change over the last 7 days - #12
Community maintained hardware plugin for vLLM on Ascend
★ 2,730+53Star change over the last 7 days - #13
OpenLake is a high performance storage engine for efficient LLM inference and GPU Training
★ 2,346+9Star change over the last 7 days - #14★ 2,227+3Star change over the last 7 days
- #15★ 2,076+0Star change over the last 7 days
- #16
MLRun is an open source MLOps platform for quickly building and managing continuous ML applications across their lifecycle. MLRun integrates into your development and CI/CD environment and automates the delivery of production data, ML pipelines, and online applications.
★ 1,694+0Star change over the last 7 days - #17
An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
★ 1,409+5Star change over the last 7 days - #18★ 1,320+6Star change over the last 7 days
- #19★ 1,303+1Star change over the last 7 days
- #20★ 1,196+1Star change over the last 7 days
- #21
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 978+70Star change over the last 7 days - #22★ 974+0Star change over the last 7 days
- #23
A scalable inference server for models optimized with OpenVINO™
★ 921+3Star change over the last 7 days - #24★ 908+0Star change over the last 7 days
- #25
A high-performance ML model serving framework, offers dynamic batching and CPU/GPU pipelines to fully exploit your compute machine
★ 902+0Star change over the last 7 days - #26
Serverless LLM Serving for Everyone.
★ 711+6Star change over the last 7 days - #27
Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.
★ 687-1Star change over the last 7 days - #28
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 660+4Star change over the last 7 days - #29
FastAPI Skeleton App to serve machine learning models production-ready.
★ 603-1Star change over the last 7 days