The Fine-Tuning Index / RLHF & Preference / #17

radixark/miles

by radixark · RLHF & Preference · updated today

Miles is an enterprise-facing reinforcement learning framework for LLM and VLM post-training, forked from and co-evolving with slime.

71
momentum
2,810
stars
482
forks
#17
rank
View on GitHub →