vllm-project/vllm
A high-throughput and memory-efficient inference and serving engine for LLMs
Explore 14 GitHub repositories focused on inference. Discover top-starred projects and those trending this week.
A high-throughput and memory-efficient inference and serving engine for LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Infrastructure for continually self‑improving agents
Achieve state of the art inference performance with modern accelerators on Kubernetes
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
Open-source inference server and production cluster for all the models your agent needs.
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
Open Source Inference Research Platform Standard / 开源推理研究平台
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
A high-throughput and memory-efficient inference and serving engine for LLMs
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Find the local LLM that actually runs and performs best on your hardware. Ranked by real, recency-aware benchmarks, not parameter count. One command, run it instantly.
Native MLX runtime for Laya typed decision models — 7–14 ms short decisions on M3 Max. No text generation, PyTorch, or cloud API.
Infrastructure for continually self‑improving agents
Achieve state of the art inference performance with modern accelerators on Kubernetes
CSGHub is a brand-new open-source platform for managing LLMs, developed by the OpenCSG team. It offers both open-source and on-premise/SaaS solutions, with features comparable to Hugging Face. Gain full control over the lifecycle of LLMs, datasets, and agents, with Python SDK compatibility with Hugging Face. Join us! ⭐️
Open-source inference server and production cluster for all the models your agent needs.
cubestudio开源云原生一站式机器学习/深度学习/大模型AI平台/MaaS/mlops/人工智能平台/训推平台,算法全链路流程,多租户,算力租赁平台,token中转,拖拉拽任务流pipeline编排,多机多卡分布式训练,超参搜索,推理服务,VGPU虚拟化,云边端协同,边缘计算,自动化标注平台,deepseek等大模型sft微调/奖励模型/强化学习训练,vllm/ollama/mindie大模型多机推理,私有知识库llmops智能体,AI模型市场,支持国产异构算力调度,昇腾/寒武纪/海光/摩尔/沐曦等,支持ib/roce/RDMA,信创支持
Open Source Inference Research Platform Standard / 开源推理研究平台
SGLang-Omni empowers high-performance serving for TTS, ASR, speech and omni models.
Build your own high performance LLM inference engine in C++ and CUDA - a smaller version of vLLM
Inference-native Tokenmaxxing Agent Harness for Loop Engineering
Evidence-aware on-prem LLM inference sizing and TCO calculator