vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Explore 5 GitHub repositories focused on llm-serving. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
MoBA: Mixture of Block Attention for Long-Context LLMs
Open Source Inference Research Platform Standard / 开源推理研究平台
fak — the Fused Agent Kernel: one Go binary that turns a tool-using agent (Claude Code, Codex, Cursor, any OpenAI/Anthropic/MCP client) into a managed agent: cache-stable model traffic, context compaction + crash resume, nanosecond tool-call policy, local GGUF serving with SSD expert offload.
A high-throughput and memory-efficient inference and serving engine for LLMs
The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.
MoBA: Mixture of Block Attention for Long-Context LLMs
Open Source Inference Research Platform Standard / 开源推理研究平台
fak — the Fused Agent Kernel: one Go binary that turns a tool-using agent (Claude Code, Codex, Cursor, any OpenAI/Anthropic/MCP client) into a managed agent: cache-stable model traffic, context compaction + crash resume, nanosecond tool-call policy, local GGUF serving with SSD expert offload.