The Fine-Tuning Index / RLHF & Preference / #55

wxhcore/bumblecore

by wxhcore · RLHF & Preference · updated 19d ago

An LLM training framework built from the ground up, featuring a custom BumbleBee architecture and end-to-end support for multiple open-source models across Pretraining → SFT → RLHF/DPO.

51
momentum
102
stars
13
forks
#55
rank
aideepseekfine-tuninggenerative-aigptinstruction-tuningllmloranlppeftpretrainqwen
View on GitHub →