Skip to main content
buildradar
Sign in
Topic · rlhf

rlhf

Tracked open-source repos tagged rlhf, sorted by stars.

Repos
44
Total stars
223,657
Avg. stars
5,083
Share
0.01%

Topics that frequently appear alongside rlhf on the same repo.

Recent risers

Repos created in the last 90 days, tagged rlhf.

No new repos tagged with this topic in the last 90 days.

  • LlamaFactory@hiyouga

    Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)

    74,532+89Star change over the last 7 days
  • Open-Assistant@LAION-AI

    OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.

    37,401-7Star change over the last 7 days
  • LLMSurvey@RUCAIBox

    The official GitHub page for the survey paper "A Survey of Large Language Models".

    12,208+2Star change over the last 7 days
  • InternLM@InternLM

    Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).

    7,277+4Star change over the last 7 days
  • 中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)

    7,120+0Star change over the last 7 days
  • alignment-handbook@huggingface

    Robust recipes to align language models with human and AI preferences

    5,670-3Star change over the last 7 days
  • OpenClaw-RL@Gen-Verse

    OpenClaw-RL: Train any agent simply by talking

    5,665+7Star change over the last 7 days
  • transformerlab-app@transformerlab

    The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.

    5,185+3Star change over the last 7 days
  • reasoning-from-scratch@rasbt

    Implement a reasoning LLM in PyTorch from scratch, step by step

    5,138+49Star change over the last 7 days
  • argilla@argilla-io

    Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets

    5,093+5Star change over the last 7 days
  • Kiln@Kiln-AI

    Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

    5,042+5Star change over the last 7 days
  • align-anything@PKU-Alignment

    Align Anything: Training All-modality Model with Feedback

    4,671+3Star change over the last 7 days
  • awesome-RLHF@opendilab

    A curated list of reinforcement learning with human feedback resources (continually updated)

    4,424+2Star change over the last 7 days
  • hands-on-modern-rl@walkinglabs

    🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

    4,199+54Star change over the last 7 days
  • docta@Docta-ai

    A Doctor for your data

    3,485-1Star change over the last 7 days
  • distilabel@argilla-io

    Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.

    3,384+3Star change over the last 7 days
  • ROLL@alibaba

    An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models

    3,380+6Star change over the last 7 days
  • TorchLeet@Exorust

    LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.

    2,461+12Star change over the last 7 days
  • rlhf-book@natolambert

    Textbook on reinforcement learning from human feedback

    2,357+11Star change over the last 7 days
  • alpaca_eval@tatsu-lab

    An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.

    2,012-1Star change over the last 7 days
  • AgentsMeetRL@thinkwee

    Awesome List for Agentic RL

    1,832+29Star change over the last 7 days
  • ImageReward@zai-org

    [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation

    1,703+1Star change over the last 7 days
  • safe-rlhf@PKU-Alignment

    Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback

    1,613+1Star change over the last 7 days
  • WebGLM@THUDM

    WebGLM: An Efficient Web-enhanced Question Answering System (KDD 2023)

    1,601-1Star change over the last 7 days
  • Recipes to train reward model for RLHF.

    1,541+0Star change over the last 7 days
  • MOSS-RLHF@OpenLMLab

    Secrets of RLHF in Large Language Models Part I: PPO

    1,430+0Star change over the last 7 days
  • xtreme1@xtreme1-io

    Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.

    1,239+2Star change over the last 7 days
  • verl-omni@verl-project

    Multimodal RL training framework for diffusion & omni models

    959+64Star change over the last 7 days
  • SimPO@princeton-nlp

    [NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward

    959+1Star change over the last 7 days
  • AI-Compass@tingaicompass

    “AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。

    933+10Star change over the last 7 days
← Back to topics