The Fine-Tuning Index / RLHF & Preference / #110
RLHFlow/Online-RLHF
by RLHFlow · RLHF & Preference · updated 1y ago
A recipe for online RLHF and online iterative DPO.
31
momentum
545
stars
48
forks
#110
rank
llama3llmrlhf
View on GitHub →