grpo
Tracked open-source repos tagged grpo, sorted by stars.
Related topics
Topics that frequently appear alongside grpo on the same repo.
Recent risers
Repos created in the last 90 days, tagged grpo.
No new repos tagged with this topic in the last 90 days.
- #1
Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15,488+88Star change over the last 7 days - #2
Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
★ 12,256+104Star change over the last 7 days - #3
Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!
★ 10,689+10Star change over the last 7 days - #4
https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG | 转行大模型 | 大模型面试 | 算法工程师 | 面试题库 | 强化学习|数据合成
★ 9,116+169Star change over the last 7 days - #5★ 6,018+4Star change over the last 7 days
- #6
OpenClaw-RL: Train any agent simply by talking
★ 5,665+7Star change over the last 7 days - #7
Implement a reasoning LLM in PyTorch from scratch, step by step
★ 5,138+49Star change over the last 7 days - #8
🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.
★ 4,199+54Star change over the last 7 days - #9
Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.
★ 3,168+0Star change over the last 7 days - #10
verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"
★ 2,274+17Star change over the last 7 days - #11
[ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1,177+0Star change over the last 7 days - #12
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
★ 1,060+2Star change over the last 7 days - #13★ 955+60Star change over the last 7 days
- #14
🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.
★ 669-1Star change over the last 7 days - #15
[NeurIPS 2025] AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning
★ 640+6Star change over the last 7 days - #16
RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios
★ 632+9Star change over the last 7 days - #17
Explore the Multimodal “Aha Moment” on 2B Model
★ 623+0Star change over the last 7 days - #18
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
★ 591+10Star change over the last 7 days - #19
A curated list of papers on reinforcement learning for video generation
★ 591+1Star change over the last 7 days - #20
Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
★ 501—Star change over the last 7 days