The Fine-Tuning Index / RLHF & Preference / #91

zht8506/Easy-LLM-Post-Training

by zht8506 · RLHF & Preference · updated 3mo ago

Implement popular LLM post-training algorithms (SFT, DFT, DPO, GRPO, etc.) in PyTorch with easy code!

39
momentum
116
stars
10
forks
#91
rank
View on GitHub →