NVlabs/GDPO
@NVlabsOfficial implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Sao
501
Fork
36
Ngôn ngữ
Python
Giấy phép
Apache-2.0
Push gần nhất
4 tháng trước
Intel liên quan (0)
Chưa có intel liên quan
Kho mã này chưa xuất hiện trong bất kỳ nguồn nào radar theo dõi. Bộ thu thập chạy theo lịch — hãy quay lại khi nó bao quát kho mã này.