Skip to main content
buildradar
Sign in
Topic · grpo

grpo

Tracked open-source repos tagged grpo, sorted by stars.

Repos
20
Total stars
81,418
Avg. stars
4,071
Share
0.00%

Topics that frequently appear alongside grpo on the same repo.

Recent risers

Repos created in the last 90 days, tagged grpo.

No new repos tagged with this topic in the last 90 days.

  • ms-swift@modelscope

    Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).

    15,488+88Star change over the last 7 days
  • AI-Research-SKILLs@Orchestra-Research

    Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.

    12,256+104Star change over the last 7 days
  • ART@OpenPipe

    Agent Reinforcement Trainer: train multi-step agents for real-world tasks using GRPO. Give your agents on-the-job training. Reinforcement learning for Qwen3.6, GPT-OSS, Llama, and more!

    10,689+10Star change over the last 7 days
  • AgentGuide@adongwanai

    https://adongwanai.github.io/AgentGuide | AI Agent开发指南 | LangGraph实战 | 高级RAG | 转行大模型 | 大模型面试 | 算法工程师 | 面试题库 | 强化学习|数据合成

    9,116+169Star change over the last 7 days
  • VLM-R1@om-ai-lab

    Solve Visual Understanding with Reinforced VLMs

    6,018+4Star change over the last 7 days
  • OpenClaw-RL@Gen-Verse

    OpenClaw-RL: Train any agent simply by talking

    5,665+7Star change over the last 7 days
  • reasoning-from-scratch@rasbt

    Implement a reasoning LLM in PyTorch from scratch, step by step

    5,138+49Star change over the last 7 days
  • hands-on-modern-rl@walkinglabs

    🚀 An open-source, hands-on curriculum bridging the gap from basic RL concepts to LLM alignment, RLVR, and advanced Agentic systems.

    4,199+54Star change over the last 7 days
  • Skywork-R1V@SkyworkAI

    Skywork-R1V is an advanced multimodal AI model series developed by Skywork AI, specializing in vision-language reasoning.

    3,168+0Star change over the last 7 days
  • verl-agent@langfengQ

    verl-agent is an extension of veRL, designed for training LLM/VLM agents via RL. verl-agent is also the official code for paper "Group-in-Group Policy Optimization for LLM Agent Training"

    2,274+17Star change over the last 7 days
  • MixGRPO@Tencent-Hunyuan

    [ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

    1,177+0Star change over the last 7 days
  • judgeval@JudgmentLabs

    The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.

    1,060+2Star change over the last 7 days
  • verl-omni@verl-project

    Multimodal RL training framework for diffusion & omni models

    955+60Star change over the last 7 days
  • oat@sail-sg

    🌾 OAT: A research-friendly framework for LLM online alignment, including reinforcement learning, preference learning, etc.

    669-1Star change over the last 7 days
  • AutoVLA@ucla-mobility

    [NeurIPS 2025] AutoVLA: A Vision-Language-Action Model for End-to-End Autonomous Driving with Adaptive Reasoning and Reinforcement Fine-Tuning

    640+6Star change over the last 7 days
  • Open-AgentRL@Gen-Verse

    RLAnything (ICML 2026) & AutoTool (ICML 2026), DemyAgent: Open-Source RL for LLMs and Agentic Scenarios

    632+9Star change over the last 7 days
  • VisualThinker-R1-Zero@turningpoint-ai

    Explore the Multimodal “Aha Moment” on 2B Model

    623+0Star change over the last 7 days
  • Relax@redai-studio

    An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale

    591+10Star change over the last 7 days
  • A curated list of papers on reinforcement learning for video generation

    591+1Star change over the last 7 days
  • GDPO@NVlabs

    Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

    501Star change over the last 7 days
← Back to topics