Skip to main content
buildradar
Sign in
Topic · distributed-training

distributed-training

Tracked open-source repos tagged distributed-training, sorted by stars.

Repos
18
Total stars
171,169
Avg. stars
9,509
Share
0.01%

Topics that frequently appear alongside distributed-training on the same repo.

Recent risers

Repos created in the last 90 days, tagged distributed-training.

No new repos tagged with this topic in the last 90 days.

  • Made-With-ML@GokuMohandas

    Learn how to develop, deploy and iterate on production-grade ML applications.

    49,268+92Star change over the last 7 days
  • pytorch-image-models@huggingface

    The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more

    37,108+31Star change over the last 7 days
  • Paddle@PaddlePaddle

    PArallel Distributed Deep LEarning: Machine Learning Framework from Industrial Practice (『飞桨』核心框架,深度学习&机器学习高性能单机、分布式训练和跨平台部署)

    24,062+4Star change over the last 7 days
  • PaddleNLP@PaddlePaddle

    Easy-to-use and powerful LLM and SLM library with awesome model zoo.

    12,970+5Star change over the last 7 days
  • skypilot@skypilot-org

    The AI Compute Platform for frontier teams. SkyPilot turns fragmented AI compute into one AI supercomputer, so frontier AI teams build custom intelligence faster.

    10,534+16Star change over the last 7 days
  • metaflow@Netflix

    Build, Manage and Deploy AI/ML Systems

    10,250+21Star change over the last 7 days
  • rllm@rllm-org

    Democratizing Reinforcement Learning for LLMs

    5,808+14Star change over the last 7 days
  • Fengshenbang-LM@IDEA-CCNL

    Fengshenbang-LM(封神榜大模型)是IDEA研究院认知计算与自然语言研究中心主导的大模型开源体系,成为中文AIGC和认知智能的基础设施。

    4,123-3Star change over the last 7 days
  • FedML@FedML-AI

    FEDML - The unified and scalable ML library for large-scale distributed training, model serving, and federated learning. FEDML Launch, a cross-cloud scheduler, further enables running any AI jobs on any GPU cloud or on-premise cluster. Built on this library, TensorOpera AI (https://TensorOpera.ai) is your generative AI platform at scale.

    4,061-1Star change over the last 7 days
  • determined@determined-ai

    Determined is an open-source machine learning platform that simplifies distributed training, hyperparameter tuning, experiment tracking, and resource management. Works with PyTorch and TensorFlow.

    3,237+2Star change over the last 7 days
  • hivemind@learning-at-home

    Decentralized deep learning in PyTorch. Built to train models on thousands of volunteers across the world.

    2,516+2Star change over the last 7 days
  • dlrover@intelligent-machine-learning

    DLRover: An Automatic Distributed Deep Learning System

    1,680+3Star change over the last 7 days
  • gloo@pytorch

    Collective communications library with various primitives for multi-machine training.

    1,447+0Star change over the last 7 days
  • DeepRec@DeepRec-AI

    DeepRec is a high-performance recommendation deep learning framework based on TensorFlow. It is hosted in incubation in LF AI & Data Foundation.

    1,197+0Star change over the last 7 days
  • efficient-dl-systems@mryab

    Efficient Deep Learning Systems course materials

    1,028+4Star change over the last 7 days
  • oat@sail-sg

    🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.

    670-1Star change over the last 7 days
  • Best practices & guides on how to write distributed pytorch training code

    629+0Star change over the last 7 days
  • Relax@redai-infra

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    581+2Star change over the last 7 days
← Back to topics