vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Explore 5 GitHub repositories focused on amd. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Open Source Inference Research Platform Standard / 开源推理研究平台
A high-throughput and memory-efficient inference and serving engine for LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
ODS V3 Pre-Release: Public testing and refinement ahead of the official V3 launch. Turn your PC, Mac, or Linux box into a private AI server.
Lemonade helps users discover and run local AI apps by serving optimized LLMs right from their own GPUs and NPUs. Join our discord: https://discord.gg/5xXzkMu8Zk
Open Source Inference Research Platform Standard / 开源推理研究平台