inference
Tracked open-source repos tagged inference, sorted by stars.
- #91
TensorRT C++ API Tutorial
★ 811+0Star change over the last 7 days - #92★ 802-1Star change over the last 7 days
- #93
Example notebooks for working with SageMaker Studio Lab. Sign up for an account at the link below!
★ 772+1Star change over the last 7 days - #94
Small, dependency-free, fast Python package to infer binary file types checking the magic numbers signature
★ 772+2Star change over the last 7 days - #95
AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data security. Including but not limited to: built-in local exclusive ChatGPT, DeepSeek, Phi, Qwen and other models, one-click batch intelligent processing of pictures, videos, audio, etc.
★ 769+0Star change over the last 7 days - #96
⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.
★ 760+8Star change over the last 7 days - #97★ 753+5Star change over the last 7 days
- #98★ 753+0Star change over the last 7 days
- #99★ 748+3Star change over the last 7 days
- #100
yolort is a runtime stack for yolov5 on specialized accelerators such as tensorrt, libtorch, onnxruntime, tvm and ncnn.
★ 730+0Star change over the last 7 days - #101
Privacy Meter: An open-source library to audit data privacy in statistical and machine learning algorithms.
★ 725+1Star change over the last 7 days - #102★ 694—Star change over the last 7 days
- #103
Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.
★ 691+2Star change over the last 7 days - #104
Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.
★ 687+0Star change over the last 7 days - #105
Train a state-of-the-art yolov3 object detector from scratch!
★ 677+0Star change over the last 7 days - #106
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 677+18Star change over the last 7 days - #107★ 673+5Star change over the last 7 days
- #108
A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support.
★ 652+0Star change over the last 7 days - #109
High-performance serverless GPU task orchestration — the scheduling layer behind wavespeed.ai
★ 632+0Star change over the last 7 days - #110
Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.
★ 630+0Star change over the last 7 days - #111
Deep Learning Streamer (DL Streamer) Pipeline Framework is an open-source streaming media analytics framework, based on GStreamer* multimedia framework, for creating complex media analytics pipelines for the Cloud or at the Edge.
★ 622-1Star change over the last 7 days - #112
convert mmdetection model to tensorrt, support fp16, int8, batch input, dynamic shape etc.
★ 596+0Star change over the last 7 days - #113★ 561—Star change over the last 7 days
- #114
Python library for YOLO small object detection and instance segmentation
★ 557+0Star change over the last 7 days - #115
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 543—Star change over the last 7 days - #116★ 539+2Star change over the last 7 days
- #117★ 532+1Star change over the last 7 days
- #118
Triton Model Analyzer is a CLI tool to help with better understanding of the compute and memory requirements of the Triton Inference Server models.
★ 526+0Star change over the last 7 days - #119
A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.
★ 514+8Star change over the last 7 days - #120★ 512+0Star change over the last 7 days