The Fine-Tuning Index / RLHF & Preference / #29

DaoyuanLi2816/mini-verl

by DaoyuanLi2816 · RLHF & Preference · updated today

Run common resolved verl PPO/GRPO configs directly on one NVIDIA GPU, with typed semantic lowering, exact recovery and portable artifacts.

68
momentum
303
stars
75
forks
#29
rank
agentic-rlalignmentconsumer-gpugrpoknowledge-distillationllmllm-agentsllm-alignmenton-policy-distillationpeftpost-trainingppo
View on GitHub →