gguf
Tracked open-source repos tagged gguf, sorted by stars.
Related topics
Topics that frequently appear alongside gguf on the same repo.
Recent risers
Repos created in the last 90 days, tagged gguf.
- #1
无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, on-screen translations, and text-to-speech (TTS).
★ 768 - #2
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 548 - #3
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
★ 501
- #1★ 34,759+265Star change over the last 7 days
- #2★ 25,859+113Star change over the last 7 days
- #3
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
★ 5,816+6Star change over the last 7 days - #4
Fine-tune LLMs from one YAML. Layer streaming trains an 8B model on a 4 GB laptop GPU.
★ 4,929+1,365Star change over the last 7 days - #5
The Swiss Army Knife of Offline AI. Chat, see, speak, and generate images on your phone or Mac — GGUF LLMs, vision, Whisper speech-to-text, Stable Diffusion, tool calling, and local-network servers. Runs on your CPU, GPU, or NPU. No account, no API key, zero data leaves your device.
★ 3,038+17Star change over the last 7 days - #6
Maid is a free and open source application for interfacing with llama.cpp models locally, and with Anthropic, DeepSeek, Ollama, Mistral and OpenAI models remotely.
★ 2,656+1Star change over the last 7 days - #7
动手学Ollama,CPU玩转大模型部署,在线阅读地址:https://datawhalechina.github.io/handy-ollama/
★ 2,516+3Star change over the last 7 days - #8
Atomic Agent is a local-first AI agent. Runs open-weight models on your own machine via llama.cpp.
★ 2,452+0Star change over the last 7 days - #9
LLM Agent Framework in ComfyUI includes MCP sever, Omost,GPT-sovits, ChatTTS,GOT-OCR2.0, and FLUX prompt nodes,access to Feishu,discord,and adapts to all llms with similar openai / aisuite interfaces, such as o1,ollama, gemini, grok, qwen, GLM, deepseek, kimi,doubao. Adapted to local llms, vlm, gguf such as llama-3.3 Janus-Pro, Linkage graphRAG
★ 2,348+5Star change over the last 7 days - #10
Run AI models locally on your machine with node.js bindings for llama.cpp. Enforce a JSON schema on the model output on the generation level
★ 2,170+4Star change over the last 7 days - #11
ggml speech-to-text inference for 16+ model families
★ 1,872+19Star change over the last 7 days - #12★ 1,835-1Star change over the last 7 days
- #13
A SOTA quantization toolkit for high-accuracy low-bit LLM inference|简洁且高效的量化工具包
★ 1,599+8Star change over the last 7 days - #14
Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.
★ 1,512+10Star change over the last 7 days - #15★ 1,437+0Star change over the last 7 days
- #16
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V
★ 1,413+20Star change over the last 7 days - #17
An open source DevOps tool from the CNCF for packaging and versioning AI/ML models, datasets, code, and configuration into an OCI Artifact.
★ 1,409+0Star change over the last 7 days - #18
Plug-and-play local AI studio: uncensored chat, image & video generation, coding agent. Runs abliterated LLMs + ComfyUI 100% offline. One installer, no Docker, no cloud.
★ 1,351+185Star change over the last 7 days - #19
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
★ 1,115+245Star change over the last 7 days - #20
A CLI to estimate inference memory requirements for Hugging Face models, written in Python.
★ 939+0Star change over the last 7 days - #21
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
★ 874+3Star change over the last 7 days - #22
Llama 3+ inference in pure Java
★ 816+0Star change over the last 7 days - #23★ 803+15Star change over the last 7 days
- #24
无需 ROOT 的开源 Android 屏幕实时翻译工具,适合游戏、视觉小说和漫画。支持端侧与云端 OCR、离线 LLM、多种翻译服务和文字朗读(TTS),译文可直接显示在画面上。Open-source no-root Android real-time screen translator for games, visual novels, and manga. Supports on-device and cloud OCR, offline LLMs, multiple translation services, on-screen translations, and text-to-speech (TTS).
★ 768+82Star change over the last 7 days - #25
Go with your own intelligence - Write Go applications that directly integrate llama.cpp for local inference using hardware acceleration on Linux, macOS, Windows, & WebAssembly.
★ 589+16Star change over the last 7 days - #26
ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
★ 588-1Star change over the last 7 days - #27★ 559+0Star change over the last 7 days
- #28
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 548—Star change over the last 7 days - #29
Arabic-first generative speech recognition — Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
★ 501+7Star change over the last 7 days