inference
Tracked open-source repos tagged inference, sorted by stars.
- #61
Manages Unified Access to Generative AI Services built on Envoy Gateway
★ 1,990+16Star change over the last 7 days - #62★ 1,927+4Star change over the last 7 days
- #63
The inference module for AgiBot X1.
★ 1,835+1Star change over the last 7 days - #64★ 1,711+1Star change over the last 7 days
- #65
Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀
★ 1,690+0Star change over the last 7 days - #66★ 1,687+4Star change over the last 7 days
- #67
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+33Star change over the last 7 days - #68
Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.
★ 1,594+8Star change over the last 7 days - #69
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
★ 1,552+11Star change over the last 7 days - #70
a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.
★ 1,550+1Star change over the last 7 days - #71
Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML
★ 1,550+4Star change over the last 7 days - #72★ 1,343+0Star change over the last 7 days
- #73★ 1,325+6Star change over the last 7 days
- #74
Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀
★ 1,305-1Star change over the last 7 days - #75
Package for causal inference in graphs and in the pairwise settings. Tools for graph structure recovery and dependencies are included.
★ 1,237+0Star change over the last 7 days - #76
Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.
★ 1,224+0Star change over the last 7 days - #77★ 1,201+5Star change over the last 7 days
- #78
Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
★ 1,195+1Star change over the last 7 days - #79
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
★ 1,118+143Star change over the last 7 days - #80
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
★ 1,115+245Star change over the last 7 days - #81
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1,094+15Star change over the last 7 days - #82★ 975+1Star change over the last 7 days
- #83
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
★ 964+5Star change over the last 7 days - #84
Bolt is a deep learning library with high performance and heterogeneous flexibility.
★ 957+0Star change over the last 7 days - #85
📚 Introduction to Modern Statistics - A college-level open-source textbook with a modern approach highlighting multivariable relationships and simulation-based inference.
★ 944+1Star change over the last 7 days - #86★ 943+0Star change over the last 7 days
- #87
A scalable inference server for models optimized with OpenVINO™
★ 927+7Star change over the last 7 days - #88★ 867+1Star change over the last 7 days
- #89★ 851-1Star change over the last 7 days
- #90
PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.
★ 848+0Star change over the last 7 days