The Fine-Tuning Index / RLHF & Preference / #45
raghavc/LLM-RLHF-Tuning-with-PPO-and-DPO
by raghavc · RLHF & Preference · updated 9d ago
Comprehensive toolkit for Reinforcement Learning from Human Feedback (RLHF) training, featuring instruction fine-tuning, reward model training, and support for PPO and DPO algorithms with various configurations for the Transformer 5 models including Qwen 3.0
56
momentum
193
stars
19
forks
#45
rank