local-llm
Tracked open-source repos tagged local-llm, sorted by stars.
Related topics
Topics that frequently appear alongside local-llm on the same repo.
Recent risers
Repos created in the last 90 days, tagged local-llm.
- #1
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,205 - #2
Run full Kimi K3 on a single device. And an OpenAI-compatible API server for local chat and coding agents.
★ 798 - #3
Up to 4× faster LLM decoding on Apple Silicon, lossless. Native MLX port of DeepSeek's DSpark & z-lab's DFlash speculative decoding — Gemma-4, Qwen3.8, Muse-Glimmer, Nemotron, LFM2.5, Ornith-1.0, ternary Bonsai-27B.
★ 645 - #4
Open-source AI companions, desktop pets, long-term memory & proactive chat. 人机恋开源项目大全。让你的家机能够脱离官端自主存在,拥有记忆和主动性。
★ 641 - #5
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
★ 628
- #1
Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
★ 47,662+140Star change over the last 7 days - #2★ 25,859+113Star change over the last 7 days
- #3
~95% on SimpleQA (e.g. Qwen3.6-27B on a 3090). Supports all local and cloud LLMs (llama.cpp, Ollama, Google, ...). 10+ search engines - arXiv, PubMed, your private documents. Everything Local & Encrypted.
★ 9,023+18Star change over the last 7 days - #4
TypeScript AI agent orchestration framework with dynamic workflows. Describe the goal, not the graph: a coordinator plans the task DAG at runtime and runs it on any LLM (Claude, ChatGPT, Gemini, DeepSeek, or local models).
★ 6,860+20Star change over the last 7 days - #5
Open source voice AI platform. Self-hosted alternative to Vapi and Retell. On Prem, BYOK across Speech to Speech or LLM/STT/TTS, with a visual workflow builder, MCP native and telephony support.
★ 5,568+45Star change over the last 7 days - #6
Local-first healthcare AI: clinical NER & HIPAA PII de-identification that runs 100% on-device. 2,200+ medical models, 21 languages, Apple MLX + Python, no cloud, no patient data leaving your network. Apache-2.0
★ 5,199+20Star change over the last 7 days - #7
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
★ 4,929+1,365Star change over the last 7 days - #8★ 4,102+1Star change over the last 7 days
- #9
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
★ 3,647+81Star change over the last 7 days - #10
Personal AI Notebooks. Organize files & webpages and generate notes from them. Open source, local & open data, open model choice (incl. local).
★ 3,554+10Star change over the last 7 days - #11
Run Claude Code 100% on-device with local AI on Apple Silicon. MLX-native Anthropic-API server. 6 fighters incl. Muse-Glimmer 30B (now multimodal — reads images, abliterated), Gemma 4 31B, Qwen 3.5 122B (65 tok/s), DeepSeek V4 Flash (1M ctx). Private, offline, airgap-ready. Built for NDA / legal / healthcare workflows.
★ 3,255+12Star change over the last 7 days - #12
Small self-contained pure-Go web server with Lua, Teal, Markdown, HTTP/2, QUIC, Redis, TypeScript, npm-less React 19, SQLite, and PostgreSQL support ++
★ 3,028+1Star change over the last 7 days - #13
A curated list of awesome platforms, tools, practices and resources that helps run LLMs locally
★ 2,823+23Star change over the last 7 days - #14
A harness optimized to smaller LLMs
★ 2,528+16Star change over the last 7 days - #15
Atomic Agent is a local-first AI agent. Runs open-weight models on your own machine via llama.cpp.
★ 2,452+0Star change over the last 7 days - #16
Translate full-length books and documents with Ollama, OpenAI-compatible, Gemini, Mistral, DeepSeek, Poe or OpenRouter. Preserves formatting. Resumes where you left off. No file size limits.
★ 2,368+31Star change over the last 7 days - #17
Free, high-quality text-to-speech API endpoint to replace OpenAI, Azure, or ElevenLabs
★ 2,075+5Star change over the last 7 days - #18★ 2,048+7Star change over the last 7 days
- #19
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
★ 1,557+6Star change over the last 7 days - #20
Row-Bot - Personal AI Sovereignty. A local-first AI assistant with integrated tools, a personal knowledge graph, voice, vision, shell, browser automation, scheduled tasks, health tracking, and messaging channels. Run locally via Ollama or add opt-in cloud models. Your data stays on your machine.
★ 1,480+11Star change over the last 7 days - #21
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V
★ 1,413+20Star change over the last 7 days - #22
Fine-tune LLMs on your Mac with Apple Silicon. SFT, DPO, GRPO, Vision, TTS, STT, Embedding, and OCR fine-tuning — natively on MLX. Unsloth-compatible API.
★ 1,398+8Star change over the last 7 days - #23
Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.
★ 1,351+185Star change over the last 7 days - #24
🌌 Give a soul to your favorite characters. Soul of Waifu is a desktop roleplay & AI companion app featuring Live2D/VRM avatars, voice chat & local LLM. Evolve together across immersive chat, RPG adventures, and your desktop.
★ 1,281+15Star change over the last 7 days - #25★ 1,254+18Star change over the last 7 days
- #26
open source grokbot. gawkbots automate your menial work via AI models and build you microapps to manage the outcome, so that you have a false sense of control.
★ 1,253+6Star change over the last 7 days - #27
Qwen3.8-27B on a single RTX 3090 with vLLM: ~1,000 tok/s at 64 concurrent (int8 tensor-core GEMMs, fp16 DeltaNet state), ~114 tok/s single-user at default sampling / ~124 greedy (MTP drafts, own-output draft vocab, calibrated int4 lm_head, split-KV verify attention), 150k-262k context; patches, requant scripts, benchmarks
★ 1,205+316Star change over the last 7 days - #28
A simple, locally hosted Web Search MCP server for use with Local LLMs
★ 1,122+8Star change over the last 7 days - #29
Synthadoc: An open-source LLM knowledge compilation engine that turns raw documents into structured, local-first wikis. A transparent, human-readable alternative to traditional RAG, which can be self-managed and self-improved without the use of any tools.
★ 1,116+2Star change over the last 7 days - #30
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
★ 1,115+245Star change over the last 7 days