vLLM 是高性能LLM推理引擎,采用PagedAttention技术,是生产环境部署的首选方案。Stars 40k+。
High-throughput LLM inference engine with PagedAttention. Essential for production deployments needing low latency and efficient GPU memory management. 🔗 Stars 40k+ github.com/vllm-project/vllm