Skip to main content
buildradar
Sign in
Topic · rlhf

rlhf

Tracked open-source repos tagged rlhf, sorted by stars.

44 repos
  • HALOs@ContextualAI

    A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).

    909-1Star change over the last 7 days
  • OpenJudge@agentscope-ai

    OpenJudge: A Unified Framework for Holistic Evaluation and Quality Rewards

    817+12Star change over the last 7 days
  • A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models

    786+26Star change over the last 7 days
  • reward-bench@allenai

    RewardBench: the first evaluation tool for reward models.

    735+2Star change over the last 7 days
  • Trinity-RFT@agentscope-ai

    Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).

    698+2Star change over the last 7 days
  • oat@sail-sg

    🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.

    669-1Star change over the last 7 days
  • Open-AgentRL@Gen-Verse

    RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios

    632+9Star change over the last 7 days
  • LLamaTuner@jianzhnie

    Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.

    621+0Star change over the last 7 days
  • Relax@redai-studio

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    591+10Star change over the last 7 days
  • SPPO@uclaml

    The official implementation of Self-Play Preference Optimization (SPPO)

    589+0Star change over the last 7 days
  • TextRL@voidful

    Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)

    564+0Star change over the last 7 days
  • Online-RLHF@RLHFlow

    A recipe for online RLHF and online iterative DPO.

    544+0Star change over the last 7 days
  • A curated collection of papers and resources on On-Policy Distillation for Large Language Models.

    534+9Star change over the last 7 days
  • dLLM-RL@Gen-Verse

    [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.

    520+1Star change over the last 7 days
← Back to topics