본문으로 건너뛰기
buildradar
Sign in

carloslfu/slotstream

@carloslfu

Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.

스타
287
포크
16
언어
Swift
라이선스
MIT
마지막 푸시
3일 전
Swiftmacosllmswiftollamallm-inferencelocal-llmqwenapple-siliconmlxmixture-of-experts