Community recipes for serving LLMs on RTX 3090/4090/5090 CUDA gpus. Multi-engine (vLLM, llama.cpp, ik_llama) and model-agnostic. Currently shipping Qwen3.6-27B Qwen3.6 35B Gemma 4 26B Gemma 4 31B configs for 1× and 2× cards.
Why this repo matters
•Provides community-tested LLM serving recipes for popular NVIDIA GPUs.
•Supports multiple inference engines for flexibility in deployment.
•Offers configurations for recent large language models on consumer hardware.