Skip
93k
vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
—
—
Explore 2 GitHub repositories focused on model-serving. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
A high-throughput and memory-efficient inference and serving engine for LLMs
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.