mixture-of-experts
Tracked open-source repos tagged mixture-of-experts, sorted by stars.
Related topics
Topics that frequently appear alongside mixture-of-experts on the same repo.
Recent risers
Repos created in the last 90 days, tagged mixture-of-experts.
- #1
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 7,533 - #2
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
★ 631 - #3
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 551 - #4
Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API.
★ 287
- #1
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43,058+0Star change over the last 7 days - #2
A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.
★ 7,533+490Star change over the last 7 days - #3★ 4,260+0Star change over the last 7 days
- #4
Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.
★ 2,515+0Star change over the last 7 days - #5
Run Mixtral-8x7B models in Colab or consumer desktops
★ 2,334+0Star change over the last 7 days - #6★ 2,322+0Star change over the last 7 days
- #7
PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538
★ 1,253+0Star change over the last 7 days - #8★ 1,088+1Star change over the last 7 days
- #9
Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4
★ 1,018+2Star change over the last 7 days - #10
⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)
★ 1,002+1Star change over the last 7 days - #11
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
★ 932+9Star change over the last 7 days - #12★ 912+0Star change over the last 7 days
- #13
From scratch implementation of a sparse mixture of experts language model inspired by Andrej Karpathy's makemore :)
★ 814+0Star change over the last 7 days - #14
Swiftlet is a Swift and Metal runtime that runs large Qwen Mixture-of-Experts models locally on Apple devices by streaming expert weights from storage, enabling 35B and 80B models to run with low RAM, including on iPhone.
★ 631+51Star change over the last 7 days - #15
中文Mixtral混合专家大模型(Chinese Mixtral MoE LLMs)
★ 612+0Star change over the last 7 days - #16
Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp
★ 551+19Star change over the last 7 days - #17
Krasis is a Hybrid LLM runtime which focuses on efficient running of larger models on consumer grade VRAM limited hardware
★ 519+3Star change over the last 7 days - #18
A library for easily merging multiple LLM experts, and efficiently train the merged LLM.
★ 516+0Star change over the last 7 days