QwenLM/Qwen
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
Explore 3 GitHub repositories focused on flash-attention. Discover top-starred projects and those trending this week.
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
MoBA: Mixture of Block Attention for Long-Context LLMs
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.
The official repo of Qwen (通义千问) chat & pretrained large language model proposed by Alibaba Cloud.
MoBA: Mixture of Block Attention for Long-Context LLMs
From teacher to tiles — a from-scratch LLM distillation & serving engine: custom Triton/CUDA kernels, FSDP distillation, paged-KV continuous batching, speculative decoding, a Rust gateway, a JAX oracle, and interpretability tooling.