Skip to main content
buildradar
Sign in
Topic · llm-inference

llm-inference

Tracked open-source repos tagged llm-inference, sorted by stars.

84 repos
  • llama3.java@mukel

    Llama 3+ inference in pure Java

    816+0Star change over the last 7 days
  • lws@kubernetes-sigs

    LeaderWorkerSet: An API for deploying a group of pods as a unit of replication

    809+8Star change over the last 7 days
  • LLM-PowerHouse: Unleash LLMs' potential through curated tutorials, best practices, and ready-to-use code for custom training and inferencing.

    731+0Star change over the last 7 days
  • llmflows@stoyan-stoyanov

    LLMFlows - Simple, Explicit and Transparent LLM Apps

    707+0Star change over the last 7 days
  • long-context-attention@feifeibear

    USP: Unified (a.k.a. Hybrid, 2D) Sequence Parallel Attention for Long Context Transformers Model Training and Inference

    691+3Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    676+17Star change over the last 7 days
  • atlas@Avarok-Cybersecurity

    Pure Rust Inference Engine

    676+4Star change over the last 7 days
  • mlx-dspark@ARahim3

    Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.

    645+30Star change over the last 7 days
  • hipfire@warpfront

    RDNA-native LLM inference engine in Rust.

    609+46Star change over the last 7 days
  • rkllama@NotPunchnox

    Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )

    602+5Star change over the last 7 days
  • yalm@andrewkchan

    Yet Another Language Model: LLM inference in C++/CUDA, no libraries except for I/O

    596-1Star change over the last 7 days
  • qvac@tetherto

    Open-source local AI SDK - run AI on-device with no cloud, no API keys. Supports GGUF, RAG, image, music, and video generation, speech-to-text, P2P inference, and more. Cross-platform: Linux, macOS, Windows, Android, iOS.

    594+35Star change over the last 7 days
  • MiniSearch@felladrin

    Minimalist web-searching platform with an AI assistant that runs directly from your browser. Demo: https://felladrin-minisearch.hf.space

    585+0Star change over the last 7 days
  • uzi@devflowinc

    CLI for running large numbers of coding agents in parallel with git worktrees

    582+0Star change over the last 7 days
  • LLM-Hub@timmyy123

    Local LLM, image&video&music generator, vibecode like cursor with local models on your phone

    580+9Star change over the last 7 days
  • LLM-FineTuning-Large-Language-Models@rohan-paul

    LLM (Large Language Model) FineTuning

    579+1Star change over the last 7 days
  • KuiperLLama@zjhellofss

    校招、秋招、春招、实习好项目,带你从零动手实现支持LLama2/3和Qwen2.5的大模型推理框架。

    573+1Star change over the last 7 days
  • zero-to-sglang@datawhalechina

    面向大模型开发者的 SGLang 系统化开源教程:从推理基础与环境搭建开始,逐步学习模型部署、结构化生成、服务开发和性能优化, 结合实战案例带你从 0 到 1 掌握 SGLang,构建高性能 LLM 推理应用

    566Star change over the last 7 days
  • qwen600@yassa9

    Static suckless single batch CUDA-only qwen3-0.6B mini inference engine

    559+0Star change over the last 7 days
  • sarathi-serve@microsoft

    A low-latency & high-throughput serving engine for LLMs

    521+0Star change over the last 7 days
  • taOS@jaylfc

    Self-hosted AI agent OS. Your memory, chat, agents, and files stay on hardware you own, offline by default, cloud by choice. Offline AI memory (taOSmd), self-hosted multi-framework group chat, a full web desktop + app store, and auto-clustering across the consumer hardware you already have (Orange/Raspberry Pi, Mac mini, gaming PC).

    521+17Star change over the last 7 days
  • krasis@brontoguana

    Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware

    519+3Star change over the last 7 days
  • ome@ome-projects

    Open Model Engine (OME) — Kubernetes operator for LLM serving, GPU scheduling, and model lifecycle management. Works with SGLang, vLLM, TensorRT-LLM, and Triton

    507Star change over the last 7 days
  • vllm-cli@Chen-zexi

    A command-line interface tool for serving LLM using vLLM.

    505-1Star change over the last 7 days
← Back to topics