vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Explore 4 GitHub repositories focused on gpt-oss. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
TokenSpeed is a speed-of-light LLM inference engine.
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V
Get up and running with Kimi-K2.6, GLM-5.2, MiniMax, DeepSeek, gpt-oss, Qwen, Gemma and other models.
A high-throughput and memory-efficient inference and serving engine for LLMs
TokenSpeed is a speed-of-light LLM inference engine.
Local AI app and inference engine for agents. Run open-weight LLMs locally — private, 100% offline on your computer. Join our Discord: https://discord.com/invite/8wGSsvmg4V