rl
Tracked open-source repos tagged rl, sorted by stars.
- #31
[EMNLP 2026] MobileGym: A Verifiable and Highly Parallel Simulation Platform for Mobile GUI Agent Research · 浏览器里运行的安卓模拟器 · Browser-hosted Android Simulator · Verifiable Evaluation · Scalable Online RL Training
★ 786+9Star change over the last 7 days - #32
Contrib package for Stable-Baselines3 - Experimental reinforcement learning (RL) code
★ 736+5Star change over the last 7 days - #33
A curated list of Monte Carlo tree search papers with implementations.
★ 715+1Star change over the last 7 days - #34
AI Research Platform for Reinforcement Learning from Real Panoramic Images.
★ 713+2Star change over the last 7 days - #35
Official repository for DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
★ 704+3Star change over the last 7 days - #36
Implementation of Inverse Reinforcement Learning (IRL) algorithms in Python/Tensorflow. Deep MaxEnt, MaxEnt, LPIRL
★ 681+2Star change over the last 7 days - #37
Hearthstone simulator using C++ with some reinforcement learning
★ 680+1Star change over the last 7 days - #38
BenchMARL is a library for benchmarking Multi-Agent Reinforcement Learning (MARL). BenchMARL allows to quickly compare different MARL algorithms, tasks, and models while being systematically grounded in its two core tenets: reproducibility and standardization.
★ 656+2Star change over the last 7 days - #39
Latest Advances on Long Chain-of-Thought Reasoning
★ 647-1Star change over the last 7 days - #40
[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
★ 589+3Star change over the last 7 days - #41
ROS2-Control implementations for Quadruped robots, include sim2real
★ 566+3Star change over the last 7 days - #42
Meta Agents Research Environments is a comprehensive platform designed to evaluate AI agents in dynamic, realistic scenarios. Unlike static benchmarks, this platform introduces evolving environments where agents must adapt their strategies as new information becomes available, mirroring real-world challenges.
★ 550+3Star change over the last 7 days - #43
Multi-Objective Reinforcement Learning algorithms implementations.
★ 533+1Star change over the last 7 days - #44
DeepThinkVLA: Enhancing Reasoning Capability of Vision-Language-Action Models
★ 529+0Star change over the last 7 days - #45★ 506+1Star change over the last 7 days
- #46
Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
★ 501—Star change over the last 7 days