vllm-project/vllm
@vllm-projectA high-throughput and memory-efficient inference and serving engine for LLMs
Stars
90,817
Forks
21,619
Language
Python
License
Apache-2.0
Last push
5 days ago
A high-throughput and memory-efficient inference and serving engine for LLMs