Skip to main content
buildradar
Sign in
Topic · inference

inference

Tracked open-source repos tagged inference, sorted by stars.

120 repos
  • ai-gateway@envoyproxy

    Manages Unified Access to Generative AI Services built on Envoy Gateway

    1,990+16Star change over the last 7 days
  • picolm@RightNow-AI

    Run a 1-billion parameter LLM on a $10 board with 256MB RAM

    1,927+4Star change over the last 7 days
  • agibot_x1_infer@AgibotTech

    The inference module for AgiBot X1.

    1,835+1Star change over the last 7 days
  • NNPACK@Maratyszcza

    Acceleration package for neural networks on multi-core CPUs

    1,711+1Star change over the last 7 days
  • Efficient, scalable and enterprise-grade CPU/GPU inference server for 🤗 Hugging Face transformer models 🚀

    1,690+0Star change over the last 7 days
  • uzu@trymirai

    A high-performance inference engine for AI models

    1,687+4Star change over the last 7 days
  • InferenceX@SemiAnalysisAI

    Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

    1,613+33Star change over the last 7 days
  • llmgateway@theopenco

    Route, manage, and analyze your LLM requests across multiple providers with a unified API interface.

    1,594+8Star change over the last 7 days
  • xllm@xLLM-AI

    A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

    1,552+11Star change over the last 7 days
  • a fast and user-friendly runtime for transformer inference (Bert, Albert, GPT2, Decoders, etc) on CPU and GPU.

    1,550+1Star change over the last 7 days
  • Nvidia GPU exporter for prometheus using nvidia-smi binary OR using NVML

    1,550+4Star change over the last 7 days
  • budgetml@ebhy

    Deploy a ML inference service on a budget in less than 10 lines of code.

    1,343+0Star change over the last 7 days
  • rtp-llm@alibaba

    RTP-LLM: Alibaba's high-performance LLM inference engine for diverse applications.

    1,325+6Star change over the last 7 days
  • EmbedAnything@StarlightSearch

    Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust 🦀

    1,305-1Star change over the last 7 days
  • CausalDiscoveryToolbox@FenTechSolutions

    Package for causal inference in graphs and in the pairwise settings. Tools for graph structure recovery and dependencies are included.

    1,237+0Star change over the last 7 days
  • kubetorch@run-house

    Distribute and run AI workloads on Kubernetes magically in Python, like PyTorch for ML infra.

    1,224+0Star change over the last 7 days
  • kvpress@NVIDIA

    LLM KV cache compression made easy

    1,201+5Star change over the last 7 days
  • ai-hub-models@qualcomm

    Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.

    1,195+1Star change over the last 7 days
  • sglang-omni@sgl-project

    SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.

    1,118+143Star change over the last 7 days
  • mlx-serve@ddalcu

    Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

    1,115+245Star change over the last 7 days
  • tiny-vllm@jmaczan

    Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

    1,094+15Star change over the last 7 days
  • Nanoflow@efeslab

    A throughput-oriented high-performance serving framework for LLMs

    975+1Star change over the last 7 days
  • mlxstudio@jjang-ai

    MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)

    964+5Star change over the last 7 days
  • bolt@huawei-noah

    Bolt is a deep learning library with high performance and heterogeneous flexibility.

    957+0Star change over the last 7 days
  • ims@OpenIntroStat

    📚 Introduction to Modern Statistics - A college-level open-source textbook with a modern approach highlighting multivariable relationships and simulation-based inference.

    944+1Star change over the last 7 days
  • neuropod@uber

    A uniform interface to run deep learning models from multiple frameworks

    943+0Star change over the last 7 days
  • model_server@openvinotoolkit

    A scalable inference server for models optimized with OpenVINO™

    927+7Star change over the last 7 days
  • bark.cpp@PABannier

    Suno AI's Bark model in C/C++ for fast text-to-speech generation

    867+1Star change over the last 7 days
  • pipeless@pipeless-ai

    An open-source computer vision framework to build and deploy apps in minutes

    851-1Star change over the last 7 days
  • pytriton@triton-inference-server

    PyTriton is a Flask/FastAPI-like interface that simplifies Triton's deployment in Python environments.

    848+0Star change over the last 7 days
← Back to topics