LMCache/LMCache
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Explore 7 GitHub repositories focused on vllm. Discover top-starred projects and those trending this week.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
Open Source Inference Research Platform Standard / 开源推理研究平台
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Open Source Inference Research Platform Standard / 开源推理研究平台
Serve large Qwen models fast on the GPUs you actually own. Qwen3.8-27B on a single 24 GB card with vLLM: 127 tok/s single-user (381 when the answer quotes the prompt), ~1,035 tok/s at 64 concurrent, 150k-262k context. vLLM patches, requant pipeline, benchmarks.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
UniRL is a Framework for Unified Multimodal Model Reinforcement Learning
A Python library powered by Language Models (LLMs) for conversational data discovery and analysis.