跳到主要内容
buildradar
Sign in

voidful/TextRL

@voidful

在 huggingface transformer 的任何生成模型(blommz-176B/bloom/gpt/bart/T5/MetaICL)上实现 ChatGPT RLHF(人类反馈强化学习)。

星数
564
Fork 数
61
语言
Python
许可
MIT
最后推送
5个月前
Pythonlanguage-modelpytorchchatgptreinforcement-learningnlpgpt-3rlhfgpt-2controlled-nlgnlg

还没有相关情报

radar 追踪的来源里还没有出现过这个 repo。收集器按计划运行——等它覆盖到这个 repo 再回来看看。