inference-server
Tracked open-source repos tagged inference-server, sorted by stars.
Related topics
Topics that frequently appear alongside inference-server on the same repo.
Recent risers
Repos created in the last 90 days, tagged inference-server.
No new repos tagged with this topic in the last 90 days.
- #1
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
★ 21,333+380Star change over the last 7 days - #2
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
★ 5,816+6Star change over the last 7 days - #3
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
★ 3,031+6Star change over the last 7 days - #4
Open-source inference server and production cluster for all the models your agent needs.
★ 3,030+173Star change over the last 7 days - #5
Turn any computer or edge device into a command center for your computer vision projects.
★ 2,431+2Star change over the last 7 days - #6
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native MLX, continuous batching, multimodal models, MCP tool calling, and Claude Code support.
★ 1,557+6Star change over the last 7 days - #7★ 1,199+3Star change over the last 7 days
- #8★ 1,055+16Star change over the last 7 days
- #9★ 852+0Star change over the last 7 days