Chuyển tới nội dung chính
buildradar
Đăng nhập

carloslfu/slotstream

@carloslfu

Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.

Sao
287
Fork
16
Ngôn ngữ
Swift
Giấy phép
MIT
Push gần nhất
6 ngày trước
Swiftmacosllmswiftollamallm-inferencelocal-llmqwenapple-siliconmlxmixture-of-experts