The Fine-Tuning Index / RLHF & Preference / #141
jackaduma/Vicuna-LoRA-RLHF-PyTorch
by jackaduma · RLHF & Preference · updated 2y ago
A full pipeline to finetune Vicuna LLM with LoRA and RLHF on consumer hardware. Implementation of RLHF (Reinforcement Learning with Human Feedback) on top of the Vicuna architecture. Basically ChatGPT but with Vicuna
26
momentum
220
stars
18
forks
#141
rank
chatgptfinetunegptllamallmlorapeftppopytorchreward-modelsrlhfvicuna
View on GitHub →