inference-engine
Tracked open-source repos tagged inference-engine, sorted by stars.
Related topics
Topics that frequently appear alongside inference-engine on the same repo.
Recent risers
Repos created in the last 90 days, tagged inference-engine.
- #1
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 6,888 - #2
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 624 - #3
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 552
- #1
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
★ 13,220+3Star change over the last 7 days - #2
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 6,888+485Star change over the last 7 days - #3
FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.
★ 4,061-1Star change over the last 7 days - #4
校招、秋招、春招、实习好项目!带你从零实现一个高性能的深度学习推理库,支持大模型 llama2 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library step by step
★ 3,497+1Star change over the last 7 days - #5
Rule engine implementation in Golang
★ 2,522+0Star change over the last 7 days - #6★ 1,964+1Star change over the last 7 days
- #7★ 1,844+2Star change over the last 7 days
- #8
3x faster speeds on MLX | Qwen 3.8 27B | Native MTP Speculative Decoding On Apple Silicon With No External Drafter.
★ 1,805+235Star change over the last 7 days - #9
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
★ 1,541+12Star change over the last 7 days - #10
Qualcomm® AI Hub Models is our collection of state-of-the-art machine learning models optimized for performance (latency, memory etc.) and ready to deploy on Qualcomm® devices.
★ 1,194+3Star change over the last 7 days - #11★ 1,145+5Star change over the last 7 days
- #12
Paddle.js is a web project for Baidu PaddlePaddle, which is an open source deep learning framework running in the browser. Paddle.js can either load a pre-trained model, or transforming a model from paddle-hub with model transforming tools provided by Paddle.js. It could run in every browser with WebGL/WebGPU/WebAssembly supported. It could also run in Baidu Smartprogram and WX miniprogram.
★ 1,103+0Star change over the last 7 days - #13
NobodyWho is an inference engine that lets you run LLMs locally and efficiently on any device.
★ 1,081+11Star change over the last 7 days - #14
MLX Studio - Home of JANG_Q - Image Gen/Edit + Chat/Code All in one - + OpenClaw (Anthropic API)
★ 960+11Star change over the last 7 days - #15★ 908+0Star change over the last 7 days
- #16★ 849+1Star change over the last 7 days
- #17
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 660+4Star change over the last 7 days - #18
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 624+98Star change over the last 7 days - #19★ 597+1Star change over the last 7 days
- #20★ 582+1Star change over the last 7 days
- #21
校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。
★ 572+9Star change over the last 7 days - #22
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
★ 552+35Star change over the last 7 days - #23
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
★ 516+3Star change over the last 7 days - #24
Inference engine for Intel devices. Serve LLMs, VLMs, Whisper, Kokoro-TTS, Embedding and Rerank models over OpenAI endpoints.
★ 513+4Star change over the last 7 days