Run Qwen3.8-Flash-Next (125B-A6B MoE) on ONE Colab A100-80GB High-RAM: pinned runtime, weights from HF, OpenAI-compatible endpoint, measured 2.8k-3.9k t/s prefill / 90-97 t/s decode, plus a verified pinned-RAM KV tier.
Why this repo matters
•Enables running large Qwen3.8-Flash-Next (125B MoE) on a single Colab A100-80GB.
•Provides an OpenAI-compatible API endpoint for easy integration.
•Achieves high performance with measured 2.8k-3.9k t/s prefill and 90-97 t/s decode.