Skip to main content
buildradar
Sign in
Topic · inference

inference

Tracked open-source repos tagged inference, sorted by stars.

120 repos
  • tensorrt-cpp-api@cyrusbehr

    TensorRT C++ API Tutorial

    811+0Star change over the last 7 days
  • cppflow@serizba

    Run TensorFlow models in C++ without installation and without Bazel

    802-1Star change over the last 7 days
  • studio-lab-examples@aws

    Example notebooks for working with SageMaker Studio Lab. Sign up for an account at the link below!

    772+1Star change over the last 7 days
  • filetype.py@h2non

    Small, dependency-free, fast Python package to infer binary file types checking the magic numbers signature

    772+2Star change over the last 7 days
  • APT@rnchg

    AI Productivity Tool - Free and open source, improve user productivity, and protect privacy and data security. Including but not limited to: built-in local exclusive ChatGPT, DeepSeek, Phi, Qwen and other models, one-click batch intelligent processing of pictures, videos, audio, etc.

    769+0Star change over the last 7 days
  • SwiftLM@SharpAI

    ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

    760+8Star change over the last 7 days
  • emlearn@emlearn

    Machine Learning inference engine for Microcontrollers and Embedded devices

    753+5Star change over the last 7 days
  • InferLLM@MegEngine

    a lightweight LLM model inference framework

    753+0Star change over the last 7 days
  • identYwaf@stamparm

    Blind WAF identification tool

    748+3Star change over the last 7 days
  • yolort@zhiqwang

    yolort is a runtime stack for yolov5 on specialized accelerators such as tensorrt, libtorch, onnxruntime, tvm and ncnn.

    730+0Star change over the last 7 days
  • ml_privacy_meter@privacytrustlab

    Privacy Meter: An open-source library to audit data privacy in statistical and machine learning algorithms.

    725+1Star change over the last 7 days
  • reef@Human-Agent-Society

    Continual learning infra for self-improving agents

    694Star change over the last 7 days
  • kglab@DerwenAI

    Graph Data Science: an abstraction layer in Python for building knowledge graphs, integrated with popular graph libraries – atop Pandas, NetworkX, RAPIDS, RDFlib, pySHACL, PyVis, morph-kgc, pslpython, pyarrow, etc.

    691+2Star change over the last 7 days
  • timber@kossisoroyce

    Ollama for classical ML models. AOT compiler that turns XGBoost, LightGBM, scikit-learn, CatBoost & ONNX models into native C99 inference code. One command to load, one command to serve. 336x faster than Python inference.

    687+0Star change over the last 7 days
  • TrainYourOwnYOLO@AntonMu

    Train a state-of-the-art yolov3 object detector from scratch!

    677+0Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    677+18Star change over the last 7 days
  • vidur@microsoft

    Accurate, large-scale, and extensible simulator for LLM inference Systems

    673+5Star change over the last 7 days
  • libonnx@xboot

    A lightweight, portable pure C99 onnx inference engine for embedded devices with hardware acceleration support.

    652+0Star change over the last 7 days
  • waverless@WaveSpeedAI

    High-performance serverless GPU task orchestration — the scheduling layer behind wavespeed.ai

    632+0Star change over the last 7 days
  • Deepdive-llama3-from-scratch@therealoliver

    Achieve the llama3 inference step-by-step, grasp the core concepts, master the process derivation, implement the code.

    630+0Star change over the last 7 days
  • dlstreamer@open-edge-platform

    Deep Learning Streamer (DL Streamer) Pipeline Framework is an open-source streaming media analytics framework, based on GStreamer* multimedia framework, for creating complex media analytics pipelines for the Cloud or at the Edge.

    622-1Star change over the last 7 days
  • convert mmdetection model to tensorrt, support fp16, int8, batch input, dynamic shape etc.

    596+0Star change over the last 7 days
  • inferoa@agentic-in

    Inference-native Tokenmaxxing Agent Harness for Loop Engineering

    561Star change over the last 7 days
  • Python library for YOLO small object detection and instance segmentation

    557+0Star change over the last 7 days
  • BigMoeOnEdge@Helldez

    Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp

    543Star change over the last 7 days
  • aikit@kaito-project

    🏗️ Fine-tune, build, and deploy open-source LLMs easily!

    539+2Star change over the last 7 days
  • kubedl@kubedl-io

    Run your deep learning workloads on Kubernetes more easily and efficiently.

    532+1Star change over the last 7 days
  • model_analyzer@triton-inference-server

    Triton Model Analyzer is a CLI tool to help with better understanding of the compute and memory requirements of the Triton Inference Server models.

    526+0Star change over the last 7 days
  • nanoowl@NVIDIA-AI-IOT

    A project that optimizes OWL-ViT for real-time inference with NVIDIA TensorRT.

    514+8Star change over the last 7 days
  • blindai@mithril-security

    Confidential AI deployment with secure enclaves :lock:

    512+0Star change over the last 7 days
← Back to topics