jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
Explore 3 GitHub repositories focused on inference-server. Discover top-starred projects and those trending this week.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Open-source inference server and production cluster for all the models your agent needs.
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
Open-source inference server and production cluster for all the models your agent needs.