Ian Cole
codingiancole
AI & ML interests
Reinforcement learning, reward modeling, RLHF, policy optimization, offline RL
Recent Activity
liked a dataset about 20 hours ago
yitingxie/rlhf-reward-datasets liked a dataset about 20 hours ago
beyond/rlhf-reward-single-round-trans_chinese liked a dataset about 20 hours ago
SeeWye/NFA_OCR_reinforcement_learning_format_TEST5Organizations
None yet