llm-inference
Tracked open-source repos tagged llm-inference, sorted by stars.
Related topics
Topics that frequently appear alongside llm-inference on the same repo.
Recent risers
Repos created in the last 90 days, tagged llm-inference.
- #1
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 7,205 - #2
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
★ 6,668 - #3
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,209 - #4
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
★ 882 - #5
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 645
- #1
GPT4All: Run Local LLMs on Any Device. Open-source and available for commercial use.
★ 77,387-9Star change over the last 7 days - #2
Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.
★ 43,688+41Star change over the last 7 days - #3★ 29,078+70Star change over the last 7 days
- #4
本项目旨在分享大模型相关技术原理以及实战经验(大模型工程化、大模型应用落地)
★ 24,995+21Star change over the last 7 days - #5
20+ high-performance LLMs with recipes to pretrain, finetune and deploy at scale.
★ 13,645+9Star change over the last 7 days - #6
Run any open-source LLMs, such as DeepSeek and Llama, as OpenAI compatible API endpoint in the cloud.
★ 12,524+1Star change over the last 7 days - #7
OpenVINO™ is an open source toolkit for optimizing and deploying AI inference
★ 10,793+30Star change over the last 7 days - #8
High-speed Large Language Model Serving for Local Deployment
★ 9,765+7Star change over the last 7 days - #9
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
★ 8,817+4Star change over the last 7 days - #10★ 8,041+8Star change over the last 7 days
- #11★ 7,952+44Star change over the last 7 days
- #12
Open-source implementation of AlphaEvolve
★ 7,307+22Star change over the last 7 days - #13
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 7,205+468Star change over the last 7 days - #14
Plano is an AI-native proxy server and data plane for agentic apps. Smart LLM routing, observability, agent orchestration, and guardrails so you stay focused on your agents core logic.
★ 7,033+11Star change over the last 7 days - #15
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
★ 6,668+178Star change over the last 7 days - #16
FlashInfer: Kernel Library for LLM Serving
★ 6,321+38Star change over the last 7 days - #17
Quantization, kernels, runtime and inference engine for mobiles, wearables, smart home and robots.
★ 5,975+20Star change over the last 7 days - #18
Standardized Distributed Generative and Predictive AI Inference Platform for Scalable, Multi-Framework Deployment on Kubernetes
★ 5,853+13Star change over the last 7 days - #19
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
★ 5,816+6Star change over the last 7 days - #20
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
★ 5,593+21Star change over the last 7 days - #21
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
★ 5,585+55Star change over the last 7 days - #22
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
★ 5,482+3Star change over the last 7 days - #23
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5,318+1Star change over the last 7 days - #24
Eko (Eko Keeps Operating) - Build Production-ready Agentic Workflow with Natural Language - eko.fellou.ai
★ 4,954+2Star change over the last 7 days - #25
RuVector is a High Performance, Real-Time, Self-Learning Ai, Vector GNN, Memory DB built in Rust.
★ 4,473+6Star change over the last 7 days - #26★ 4,260+3Star change over the last 7 days
- #27
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
★ 4,169+4Star change over the last 7 days - #28★ 3,827+1Star change over the last 7 days
- #29
A collection of scientific methods, processes, algorithms, and systems to build stories & models.
★ 3,676+2Star change over the last 7 days - #30
Add a real-time analytics node to your operational database. Spice is a portable, accelerated SQL query, search, and LLM-inference engine in Rust for data-grounded AI apps and agents.
★ 3,075+1Star change over the last 7 days