The Fine-Tuning Index / RLHF & Preference / #175
OpenMOSE/RWKV-LM-RLHF
by OpenMOSE · RLHF & Preference · updated 11mo ago
Reinforcement Learning Toolkit for RWKV.(v6,v7,ARWKV) Distillation,SFT,RLHF(DPO,ORPO), infinite context training, Aligning. Exploring the possibilities for deeper fine-tuning of RWKV.
20
momentum
63
stars
6
forks
#175
rank