rlhf
Tracked open-source repos tagged rlhf, sorted by stars.
- #31
A library with extensible implementations of DPO, KTO, PPO, ORPO, and other human-aware loss functions (HALOs).
★ 909-1Star change over the last 7 days - #32★ 817+12Star change over the last 7 days
- #33
A curated collection of papers, technical reports, frameworks, and tools for on-policy distillation (OPD) of large language models
★ 786+26Star change over the last 7 days - #34
RewardBench: the first evaluation tool for reward models.
★ 735+2Star change over the last 7 days - #35
Trinity-RFT is a general-purpose, flexible and scalable framework designed for reinforcement fine-tuning (RFT) of large language models (LLM).
★ 698+2Star change over the last 7 days - #36
🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
★ 669-1Star change over the last 7 days - #37
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 632+9Star change over the last 7 days - #38
Easy and Efficient Finetuning LLMs. (Supported LLama, LLama2, LLama3, Qwen, Baichuan, GLM , Falcon) 大模型高效量化训练+部署.
★ 621+0Star change over the last 7 days - #39
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 591+10Star change over the last 7 days - #40★ 589+0Star change over the last 7 days
- #41
Implementation of ChatGPT RLHF (Reinforcement Learning with Human Feedback) on any generation model in huggingface's transformer (blommz-176B/bloom/gpt/bart/T5/MetaICL)
★ 564+0Star change over the last 7 days - #42
A recipe for online RLHF and online iterative DPO.
★ 544+0Star change over the last 7 days - #43
A curated collection of papers and resources on On-Policy Distillation for Large Language Models.
★ 534+9Star change over the last 7 days - #44
[ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.
★ 520+1Star change over the last 7 days