The Fine-Tuning Index / RLHF & Preference / #29
DaoyuanLi2816/mini-verl
by DaoyuanLi2816 · RLHF & Preference · updated today
Run common resolved verl PPO/GRPO configs directly on one NVIDIA GPU, with typed semantic lowering, exact recovery and portable artifacts.
68
momentum
303
stars
75
forks
#29
rank
agentic-rlalignmentconsumer-gpugrpoknowledge-distillationllmllm-agentsllm-alignmenton-policy-distillationpeftpost-trainingppo
View on GitHub →