llm-inference
Tracked open-source repos tagged llm-inference, sorted by stars.
- #31
Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.
★ 3,050+4Star change over the last 7 days - #32
Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2,771+0Star change over the last 7 days - #33
Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).
★ 2,521+7Star change over the last 7 days - #34
An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.
★ 2,439+41Star change over the last 7 days - #35
The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that interact with your data and UI.
★ 2,086+21Star change over the last 7 days - #36★ 2,076+0Star change over the last 7 days
- #37
Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.
★ 1,936-1Star change over the last 7 days - #38★ 1,764+4Star change over the last 7 days
- #39
Declarative way to run AI models in React Native on device, powered by ExecuTorch.
★ 1,707+5Star change over the last 7 days - #40
Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3
★ 1,613+33Star change over the last 7 days - #41
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
★ 1,552+11Star change over the last 7 days - #42
The non profit fostering AI in India through credits, grants, resources
★ 1,461+56Star change over the last 7 days - #43
Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser
★ 1,449+1Star change over the last 7 days - #44
List of software that allows searching the web with the assistance of AI: https://hf.co/spaces/felladrin/awesome-ai-web-search
★ 1,426+3Star change over the last 7 days - #45
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V
★ 1,413+20Star change over the last 7 days - #46
Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs
★ 1,396+12Star change over the last 7 days - #47
A curated collection of top-tier penetration testing tools and productivity utilities across multiple domains. Join us to explore, contribute, and enhance your hacking toolkit!
★ 1,383+1Star change over the last 7 days - #48
LLMs as Copilots for Theorem Proving in Lean
★ 1,318+1Star change over the last 7 days - #49★ 1,275+130Star change over the last 7 days
- #50
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,218+334Star change over the last 7 days - #51
Streamlines and simplifies prompt design for both developers and non-technical users with a low code approach.
★ 1,154+0Star change over the last 7 days - #52
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
★ 1,094+15Star change over the last 7 days - #53
OpenAlpha_Evolve is an open-source Python framework inspired by the groundbreaking research on autonomous coding agents like DeepMind's AlphaEvolve.
★ 1,050+0Star change over the last 7 days - #54
“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。
★ 935+12Star change over the last 7 days - #55
List of awesome hosting sorted by minimal plan price
★ 926+4Star change over the last 7 days - #56
Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)
★ 925+4Star change over the last 7 days - #57★ 908+0Star change over the last 7 days
- #58
LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.
★ 890+2Star change over the last 7 days - #59
A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
★ 882+78Star change over the last 7 days - #60
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
★ 874+3Star change over the last 7 days