The Fine-Tuning Index / Training Frameworks / #161
RLHFlow/Reinforce-Ada
by RLHFlow · Training Frameworks · updated 9mo ago
[COLM 2026] An adaptive sampling framework for Reinforce-style LLM post training.
22
momentum
96
stars
17
forks
#161
rank