Skip to main content
buildradar
Sign in

vllm-project/vllm

@vllm-project

A high-throughput and memory-efficient inference and serving engine for LLMs

Stars
90,817
Forks
21,619
Language
Python
License
Apache-2.0
Last push
5 days ago
Pythonopenaillmcudagpttransformerpytorchdeepseekqwenllamainferencekimiqwen3tpullm-servingmodel-servingmoedeepseek-v3amdgpt-ossblackwell