Skip to main content
buildradar
Sign in
Topic · llm-inference

llm-inference

Tracked open-source repos tagged llm-inference, sorted by stars.

84 repos
  • Distributed LLM inference. Connect home devices into a powerful cluster to accelerate LLM inference. More devices means faster inference.

    3,050+4Star change over the last 7 days
  • Medusa@FasterDecoding

    Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads

    2,771+0Star change over the last 7 days
  • EAGLE@SafeAILab

    Official Implementation of EAGLE-1 (ICML'24), EAGLE-2 (EMNLP'24), and EAGLE-3 (NeurIPS'25).

    2,521+7Star change over the last 7 days
  • nanocoder@Nano-Collective

    An open coding agent for your terminal, built by a community collective rather than a company. Bring your own model, keep your code on your machine, and owe nothing to anyone.

    2,439+41Star change over the last 7 days
  • neuron-ai@neuron-core

    The Agentic Framework of the PHP ecosystem to build production-ready AI driven applications. Connect components (LLMs, Tools, vector DBs, memory) to agents that interact with your data and UI.

    2,086+21Star change over the last 7 days
  • aici@microsoft

    AICI: Prompts as (Wasm) Programs

    2,076+0Star change over the last 7 days
  • llama2-webui@liltom-eth

    Run any Llama 2 locally with gradio UI on GPU or CPU from anywhere (Linux/Windows/Mac). Use `llama2-wrapper` as your local llama2 backend for Generative Agents/Apps.

    1,936-1Star change over the last 7 days
  • beta9@beam-cloud

    Ultrafast serverless GPU inference, sandboxes, and background jobs

    1,764+4Star change over the last 7 days
  • react-native-executorch@software-mansion

    Declarative way to run AI models in React Native on device, powered by ExecuTorch.

    1,707+5Star change over the last 7 days
  • InferenceX@SemiAnalysisAI

    Open Source Continuous Inference Benchmark Research Platform — Kimi K3 2.8T, MiniMax M3, DeepSeekv4, GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72 & soon™ TPUv6e/v7/Trainium2/3 | 开源持续推理基准研究平台 — Kimi K2.7-Code、MiniMax M3、DeepSeekv4、GLM5 - GB200 NVL72 vs MI355X vs B200 vs GB300 NVL72,即将推出™ TPUv6e/v7/Trainium2/3

    1,613+33Star change over the last 7 days
  • xllm@xLLM-AI

    A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.

    1,552+11Star change over the last 7 days
  • aigrantsindia@aigrantsindia

    The non profit fostering AI in India through credits, grants, resources

    1,461+56Star change over the last 7 days
  • BrowserAI@sauravpanda

    Run local LLMs like llama, deepseek-distill, kokoro and more inside your browser

    1,449+1Star change over the last 7 days
  • List of software that allows searching the web with the assistance of AI: https://hf.co/spaces/felladrin/awesome-ai-web-search

    1,426+3Star change over the last 7 days
  • Atomic-Chat@AtomicBot-ai

    Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V

    1,413+20Star change over the last 7 days
  • scaling-book@jax-ml

    Home for "How To Scale Your Model", a short blog-style textbook about scaling LLMs on TPUs

    1,396+12Star change over the last 7 days
  • A curated collection of top-tier penetration testing tools and productivity utilities across multiple domains. Join us to explore, contribute, and enhance your hacking toolkit!

    1,383+1Star change over the last 7 days
  • LeanCopilot@lean-dojo

    LLMs as Copilots for Theorem Proving in Lean

    1,318+1Star change over the last 7 days
  • kvcached@ovg-project

    Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond

    1,275+130Star change over the last 7 days
  • Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks

    1,218+334Star change over the last 7 days
  • prompt-poet@character-ai

    Streamlines and simplifies prompt design for both developers and non-technical users with a low code approach.

    1,154+0Star change over the last 7 days
  • tiny-vllm@jmaczan

    Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM

    1,094+15Star change over the last 7 days
  • OpenAlpha_Evolve@shyamsaktawat

    OpenAlpha_Evolve is an open-source Python framework inspired by the groundbreaking research on autonomous coding agents like DeepMind's AlphaEvolve.

    1,050+0Star change over the last 7 days
  • AI-Compass@tingaicompass

    “AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。

    935+12Star change over the last 7 days
  • List of awesome hosting sorted by minimal plan price

    926+4Star change over the last 7 days
  • Pure C++ implementation of several models for real-time chatting on your computer (CPU & GPU)

    925+4Star change over the last 7 days
  • ZhiLight@zhihu

    A highly optimized LLM inference acceleration engine for Llama and its variants.

    908+0Star change over the last 7 days
  • llm_note@harleyszhang

    LLM notes, including model inference, transformer model structure, and llm framework code analysis notes.

    890+2Star change over the last 7 days
  • A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.

    882+78Star change over the last 7 days
  • LLM.swift@eastriverlee

    LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.

    874+3Star change over the last 7 days
← Back to topics