inference
Tracked open-source repos tagged inference, sorted by stars.
- #31
Superduper: End-to-end framework for building custom AI applications and agents.
★ 5,318+1Star change over the last 7 days - #32
cube studio开源云原生一站式机器学习/深度学习/大模型AI平台,mlops算法链路全流程,算力租赁平台,notebook在线开发,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务VGPU虚拟化,边缘计算,标注平台自动化标注,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库,AI模型市场,支持国产cpu/gpu/npu 昇腾生态,支持RDMA,支持pytorch/tf/mxnet/deepspeed/paddle/colossalai/horovod/ray/volcano等分布式
★ 5,078+3Star change over the last 7 days - #33★ 4,882+3Star change over the last 7 days
- #34
Fast inference engine for Transformer models
★ 4,658+7Star change over the last 7 days - #35
TNN: developed by Tencent Youtu Lab and Guangying Lab, a uniform deep learning inference framework for mobile、desktop and server. TNN is distinguished by several outstanding features, including its cross-platform capability, high performance, model compression and code pruning. Based on ncnn and Rapidnet, TNN further strengthens the support and performance optimization for mobile devices, and also draws on the advantages of good extensibility and high performance from existed open source efforts. TNN has been deployed in multiple Apps from Tencent, such as Mobile QQ, Weishi, Pitu, etc. Contributions are welcome to work in collaborative with us and make TNN a better framework.
★ 4,649+1Star change over the last 7 days - #36★ 4,438+8Star change over the last 7 days
- #37
Pre-trained Deep Learning models and demos (high quality and extremely fast)
★ 4,424+0Star change over the last 7 days - #38★ 4,390+69Star change over the last 7 days
- #39
A unified inference and post-training framework for accelerated video generation.
★ 4,300+150Star change over the last 7 days - #40
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
★ 4,104+3Star change over the last 7 days - #41
Any model. Any hardware. Zero compromise. Built with @ziglang / @openxla / MLIR / @bazelbuild
★ 4,015+5Star change over the last 7 days - #42
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
★ 3,714+3Star change over the last 7 days - #43
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100% tool calling. 17 tool parsers, prompt cache, reasoning separation, cloud routing. Drop-in OpenAI replacement. Works with Claude Code, Cursor, Aider.
★ 3,647+81Star change over the last 7 days - #44
校招、秋招、春招、实习好项目!带你从零实现一个高性能的深度学习推理库,支持大模型 llama2 、Unet、Yolov5、Resnet等模型的推理。Implement a high-performance deep learning inference library step by step
★ 3,499+2Star change over the last 7 days - #45
🚀 Accelerate inference and training of 🤗 Transformers, Diffusers, TIMM and Sentence Transformers with easy to use hardware optimization tools
★ 3,477+4Star change over the last 7 days - #46
📚 Jupyter notebook tutorials for OpenVINO™
★ 3,212+3Star change over the last 7 days - #47
Open-source inference server and production cluster for all the models your agent needs.
★ 3,030+173Star change over the last 7 days - #48★ 2,961+2Star change over the last 7 days
- #49
Community maintained hardware plugin for vLLM on Ascend
★ 2,749+19Star change over the last 7 days - #50
Use Hugging Face with JavaScript
★ 2,510+3Star change over the last 7 days - #51★ 2,484+6Star change over the last 7 days
- #52
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
★ 2,469+12Star change over the last 7 days - #53
High-efficiency floating-point neural network inference operators for mobile, server, and Web
★ 2,439+2Star change over the last 7 days - #54
Turn any computer or edge device into a command center for your computer vision projects.
★ 2,431+2Star change over the last 7 days - #55
Mastering Applied AI, One Concept at a Time
★ 2,378+1Star change over the last 7 days - #56
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.
★ 2,233+4Star change over the last 7 days - #57★ 2,177+5Star change over the last 7 days
- #58
Inference Llama 2 in one file of pure 🔥
★ 2,126+0Star change over the last 7 days - #59
MII makes low-latency and high-throughput inference possible, powered by DeepSpeed.
★ 2,112+1Star change over the last 7 days - #60★ 2,076+0Star change over the last 7 days