llm-inference
Tracked open-source repos tagged llm-inference, sorted by stars.
- #61
Llama 3+ inference in pure Java
★ 816+0Star change over the last 7 days - #62★ 809+8Star change over the last 7 days
- #63LLM-PowerHouse-A-Curated-Guide-for-Large-Language-Models-with-Custom-Training-and-Inferencing↗@ghimiresunil
LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for custom training and inferencing.
★ 731+0Star change over the last 7 days - #64★ 707+0Star change over the last 7 days
- #65
USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference
★ 691+3Star change over the last 7 days - #66
Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2
★ 676+17Star change over the last 7 days - #67★ 676+4Star change over the last 7 days
- #68
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 645+30Star change over the last 7 days - #69★ 609+46Star change over the last 7 days
- #70
Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )
★ 602+5Star change over the last 7 days - #71★ 596-1Star change over the last 7 days
- #72
Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.
★ 594+35Star change over the last 7 days - #73
Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space
★ 585+0Star change over the last 7 days - #74★ 582+0Star change over the last 7 days
- #75
Local LLM, image&video&music generator, vibecode like cursor with local models on your phone
★ 580+9Star change over the last 7 days - #76
LLM (Large Language Model) FineTuning
★ 579+1Star change over the last 7 days - #77
校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。
★ 573+1Star change over the last 7 days - #78
面向大模型开发者的 SGLang 系统化开源教程:从推理基础与环境搭建开始,逐步学习模型部署、结构化生成、服务开发和性能优化, 结合实战案例带你从 0 到 1 掌握 SGLang,构建高性能 LLM 推理应用
★ 566—Star change over the last 7 days - #79★ 559+0Star change over the last 7 days
- #80
A low-latency & high-throughput serving engine for LLMs
★ 521+0Star change over the last 7 days - #81
Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).
★ 521+17Star change over the last 7 days - #82
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
★ 519+3Star change over the last 7 days - #83
Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton
★ 507—Star change over the last 7 days - #84★ 505-1Star change over the last 7 days