Skip
93k
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
—
—
Explore 2 GitHub repositories focused on tpu. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
A high-throughput and memory-efficient inference and serving engine for LLMs
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.