vllm-project/vllm

A high-throughput and memory-efficient inference and serving engine for LLMs

View on GitHub
Python
Stars 87.1k
Forks 19.8k
License Apache-2.0
Open Issues 6056
Updated 9h ago