The Fine-Tuning Index / Training Frameworks / #161

RLHFlow/Reinforce-Ada

by RLHFlow · Training Frameworks · updated 9mo ago

[COLM 2026] An adaptive sampling framework for Reinforce-style LLM post training.

22
momentum
96
stars
17
forks
#161
rank
View on GitHub →