Skip to main content
buildradar
Sign in
Topic · moe

moe

Tracked open-source repos tagged moe, sorted by stars.

Repos
24
Total stars
278,965
Avg. stars
11,624
Share
0.01%

Topics that frequently appear alongside moe on the same repo.

Recent risers

Repos created in the last 90 days, tagged moe.

  • FreeToken@FlashML-org

    FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

    12,036
  • kimi-k3-in-c@FareedKhan-dev

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

    7,187
  • BigMoeOnEdge@Helldez

    Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp

    548
  • vllm@vllm-project

    A high-throughput and memory-efficient inference and serving engine for LLMs

    90,817+402Star change over the last 7 days
  • LlamaFactory@hiyouga

    Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

    74,532+89Star change over the last 7 days
  • sglang@sgl-project

    SGLang is a high-performance serving framework for large language models and multimodal models.

    33,538+791Star change over the last 7 days
  • ms-swift@modelscope

    Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).

    15,488+88Star change over the last 7 days
  • TensorRT-LLM@NVIDIA

    TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

    14,536+38Star change over the last 7 days
  • FreeToken@FlashML-org

    FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

    12,036+2,438Star change over the last 7 days
  • kimi-k3-in-c@FareedKhan-dev

    A 2.78-trillion-parameter Kimi K3 running inference on a single CPU in 8.24 GB of RAM. Portable C99: no BLAS, no framework, no GPU.

    7,187+468Star change over the last 7 days
  • flashinfer@flashinfer-ai

    FlashInfer: Kernel Library for LLM Serving

    6,321+38Star change over the last 7 days
  • Bangumi@czy0729

    :electron: An unofficial https://bgm.tv ui first app client for Android and iOS, built with React Native. 一个无广告、以爱好为驱动、不以盈利为目的、专门做 ACG 的类似豆瓣的追番记录,bgm.tv 第三方客户端。为移动端重新设计,内置大量加强的网页端难以实现的功能,且提供了相当的自定义选项。 目前已适配 iOS / Android。

    5,900+6Star change over the last 7 days
  • GLM-4.5@zai-org

    GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

    4,422+1Star change over the last 7 days
  • MoE-LLaVA@PKU-YuanGroup

    【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models

    2,322+0Star change over the last 7 days
  • MoBA@MoonshotAI

    MoBA: Mixture of Block Attention for Long-Context LLMs

    2,174+1Star change over the last 7 days
  • uccl@uccl-project

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    1,502+3Star change over the last 7 days
  • mixture-of-experts@davidmrau

    PyTorch Re-Implementation of "The Sparsely-Gated Mixture-of-Experts Layer" by Noam Shazeer et al. https://arxiv.org/abs/1701.06538

    1,253+0Star change over the last 7 days
  • Tutel@microsoft

    Tutel MoE: Optimized Mixture-of-Experts Library, Support GptOss/DeepSeek/Kimi-K2/Qwen3 using FP8/NVFP4/MXFP4

    1,017+3Star change over the last 7 days
  • llama-moe@pjlab-sys4nlp

    ⛷️ LLaMA-MoE: Building Mixture-of-Experts from LLaMA with Continual Pre-training (EMNLP 2024)

    1,003+2Star change over the last 7 days
  • cudnn-frontend@NVIDIA

    cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.

    931+13Star change over the last 7 days
  • Adan@sail-sg

    Adan: Adaptive Nesterov Momentum Algorithm for Faster Optimizing Deep Models

    822+0Star change over the last 7 days
  • DeepSeek-671B-SFT-Guide@ScienceOne-AI

    An open-source solution for full parameter fine-tuning of DeepSeek-V3/R1 671B, including complete code and scripts from training to inference, as well as some practical experiences and conclusions. (DeepSeek-V3/R1 满血版 671B 全参数微调的开源解决方案,包含从训练到推理的完整代码和脚本,以及实践中积累一些经验和结论。)

    812+0Star change over the last 7 days
  • SwiftLM@SharpAI

    ⚡ Native MLX Swift LLM inference server for Apple Silicon. OpenAI-compatible API, SSD streaming for 100B+ MoE models, TurboQuant KV cache compression, MACOS + iOS iPhone app.

    760+8Star change over the last 7 days
  • YOLO-Master@Tencent

    [CVPR2026]🚀🚀🚀Official code for the paper "YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection." *(YOLO = You Only Look Once)* 🔥🔥🔥

    709+13Star change over the last 7 days
  • pegainfer@pegainfer-project

    Pure Rust + CUDA LLM inference engine — no PyTorch, OpenAI-compatible, serves Qwen3 to Kimi-K2

    676+18Star change over the last 7 days
  • Chinese-Mixtral@ymcui

    中文Mixtral混合专家大模型(Chinese Mixtral MoE LLMs)

    612+0Star change over the last 7 days
  • BigMoeOnEdge@Helldez

    Run MoE models bigger than your RAM. Frontier-size MoE on a 12 GB phone, CPU only, lossless, on stock llama.cpp

    548Star change over the last 7 days
← Back to topics