← Research ReadingThe RLxx Zoo: what the suffix actually names, and what has held up
The RLxx Zoo: what the suffix actually names, and what has held up
Do a survey of the state of the art in various RL methods being experimented to improve language models. RLHF and RLVR are known, but Jev uses RLCD and in the process I started seeing a variety of RLxx algorithms. They seem to describe the shape of the data that the reward is being derived off. What are people up to and what has been working?