Skip
6.5k
flashinfer-ai/flashinfer
FlashInfer: Kernel Library for LLM Serving
—
—
Explore 3 GitHub repositories written in Cuda. Discover top-starred projects and those trending this week.
FlashInfer: Kernel Library for LLM Serving
DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers
No description.
FlashInfer: Kernel Library for LLM Serving
DeepSelect: TopK kernels for DeepSeek Sparse Attention (DSA) and Samplers
No description.