Skip
6.5k
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
—
—
Explore 3 GitHub repositories focused on distributed-inference. Discover top-starred projects and those trending this week.
FlashInfer: Kernel Library for LLM Serving
Achieve state of the art inference performance with modern accelerators on Kubernetes
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
FlashInfer: Kernel Library for LLM Serving
Achieve state of the art inference performance with modern accelerators on Kubernetes
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.