Skip to main content
buildradar
Sign in
Topic · inference-engine

inference-engine

Tracked open-source repos tagged inference-engine, sorted by stars.

Repos
24
Total stars
48,964
Avg. stars
2,040
Share
0.00%

Topics that frequently appear alongside inference-engine on the same repo.

Recent risers

Repos created in the last 90 days, tagged inference-engine.

  • kimi-k3-in-c@FareedKhan-dev

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

    6,888
  • mlx-dspark@ARahim3

    Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

    624
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    552
  • ✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记

    13,220+3Star change over the last 7 days
  • kimi-k3-in-c@FareedKhan-dev

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

    6,888+485Star change over the last 7 days
  • FedML@FedML-AI

    FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.

    4,061-1Star change over the last 7 days
  • KuiperInfer@zjhellofss

    校招、秋招、春招、实习好项目!带你从零实现一个高性能的深度学习推理库,支持大模型 llama2 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library step by step

    3,497+1Star change over the last 7 days
  • grule-rule-engine@hyperjumptech

    Rule engine implementation in Golang

    2,522+0Star change over the last 7 days
  • onediff@siliconflow

    OneDiff: An out-of-the-box acceleration library for diffusion models.

    1,964+1Star change over the last 7 days
  • sonar@dphnAI

    Large-scale LLM inference engine

    1,844+2Star change over the last 7 days
  • MTPLX@youssofal

    3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.

    1,805+235Star change over the last 7 days
  • xllm@xLLM-AI

    A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

    1,541+12Star change over the last 7 days
  • ai-hub-models@qualcomm

    Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.

    1,194+3Star change over the last 7 days
  • kvcached@ovg-project

    Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

    1,145+5Star change over the last 7 days
  • Paddle.js@PaddlePaddle

    Paddle.js is a web project for Baidu PaddlePaddle, which is an open source deep learning framework running in the browser. Paddle.js can either load a pre-trained model, or transforming a model from paddle-hub with model transforming tools provided by Paddle.js. It could run in every browser with WebGL/WebGPU/WebAssembly supported. It could also run in Baidu Smartprogram and WX miniprogram.

    1,103+0Star change over the last 7 days
  • nobodywho@nobodywho-ooo

    NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.

    1,081+11Star change over the last 7 days
  • mlxstudio@jjang-ai

    MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)

    960+11Star change over the last 7 days
  • ZhiLight@zhihu

    A highly optimized LLM inference acceleration engine for Llama and its variants.

    908+0Star change over the last 7 days
  • Savant@insight-platform

    Python Computer Vision & Video Analytics Framework With Batteries Included

    849+1Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    660+4Star change over the last 7 days
  • mlx-dspark@ARahim3

    Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

    624+98Star change over the last 7 days
  • yalm@andrewkchan

    Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O

    597+1Star change over the last 7 days
  • astroid@pylint-dev

    A common base representation of python source code for pylint and other projects

    582+1Star change over the last 7 days
  • KuiperLLama@zjhellofss

    校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。

    572+9Star change over the last 7 days
  • tessera@zengxiao-he

    From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.

    552+35Star change over the last 7 days
  • krasis@brontoguana

    Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

    516+3Star change over the last 7 days
  • OpenArc@SearchSavior

    Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.

    513+4Star change over the last 7 days
← Back to topics