rlhf
Tracked open-source repos tagged rlhf, sorted by stars.
Related topics
Topics that frequently appear alongside rlhf on the same repo.
Recent risers
Repos created in the last 90 days, tagged rlhf.
No new repos tagged with this topic in the last 90 days.
- #1
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74,532+89Star change over the last 7 days - #2
OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37,401-7Star change over the last 7 days - #3
The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12,208+2Star change over the last 7 days - #4
Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7,277+4Star change over the last 7 days - #5
中文LLaMA-2 & Alpaca-2大模型二期项目 + 64K超长上下文模型 (Chinese LLaMA-2 & Alpaca-2 LLMs with 64K long context models)
★ 7,120+0Star change over the last 7 days - #6
Robust recipes to align language models with human and AI preferences
★ 5,670-3Star change over the last 7 days - #7
OpenClaw-RL: Train any agent simply by talking
★ 5,665+7Star change over the last 7 days - #8
The open source research environment for AI researchers to seamlessly train, evaluate, and scale models from local hardware to GPU clusters.
★ 5,185+3Star change over the last 7 days - #9
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 5,138+49Star change over the last 7 days - #10
Argilla is a collaboration tool for AI engineers and domain experts to build high-quality datasets
★ 5,093+5Star change over the last 7 days - #11
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
★ 5,042+5Star change over the last 7 days - #12
Align Anything: Training All-modality Model with Feedback
★ 4,671+3Star change over the last 7 days - #13
A curated list of reinforcement learning with human feedback resources (continually updated)
★ 4,424+2Star change over the last 7 days - #14
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
★ 4,199+54Star change over the last 7 days - #15★ 3,485-1Star change over the last 7 days
- #16
Distilabel is a framework for synthetic data and AI feedback for engineers who need fast, reliable and scalable pipelines based on verified research papers.
★ 3,384+3Star change over the last 7 days - #17
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3,380+6Star change over the last 7 days - #18
LeetCode for PyTorch — 65 ML/AI interview problems from real interviews at Google, Meta, Anthropic. Jupyter notebooks, an auto-grader, and an MCP AI tutor.
★ 2,461+12Star change over the last 7 days - #19★ 2,357+11Star change over the last 7 days
- #20
An automatic evaluator for instruction-following language models. Human-validated, high-quality, cheap, and fast.
★ 2,012-1Star change over the last 7 days - #21
Awesome List for Agentic RL
★ 1,832+29Star change over the last 7 days - #22
[NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1,703+1Star change over the last 7 days - #23
Safe RLHF: Constrained Value Alignment via Safe Reinforcement Learning from Human Feedback
★ 1,613+1Star change over the last 7 days - #24★ 1,601-1Star change over the last 7 days
- #25
Recipes to train reward model for RLHF.
★ 1,541+0Star change over the last 7 days - #26★ 1,430+0Star change over the last 7 days
- #27
Xtreme1 is an all-in-one data labeling and annotation platform for multimodal data training and supports 3D LiDAR point cloud, image, and LLM.
★ 1,239+2Star change over the last 7 days - #28★ 959+64Star change over the last 7 days
- #29
[NeurIPS 2024] SimPO: Simple Preference Optimization with a Reference-Free Reward
★ 959+1Star change over the last 7 days - #30
“AI-Compass”将为社区指引在 AI 技术海洋中航行的方向,无论你是初学者还是进阶开发者,都能在这里找到通往 AI 各大方向的路径。旨在帮助开发者系统性地了解 AI 的核心概念、主流技术、前沿趋势,并通过实践掌握从理论到落地的全过程。
★ 933+10Star change over the last 7 days