Run a 105 GB AI model on a Mac that can't hold it. Slotstream streams Qwen3.8-Flash-Next (125B mixture of experts) from your SSD and caches the busiest experts in memory, so it runs on Macs with 16 to 64 GB. One native Swift binary on MLX and Metal, no Python, offline. Works with Claude Code, Codex and Ollama or OpenAI clients.
Why this repo matters
•Enables running large LLMs like Qwen3.8-Flash-Next on Macs with limited RAM.
•Leverages MLX and Swift for efficient LLM inference on Apple Silicon.
•Offers an Ollama-compatible API for easy integration and local LLM deployment.