The Fine-Tuning Index / RLHF & Preference / #81

zhoujx4/llm-atlas

by zhoujx4 · RLHF & Preference · updated 1mo ago

LLM 训练算法知识图谱:SFT / LoRA / DPO / RLHF / Agent — A knowledge atlas of LLM training algorithms

41
momentum
31
stars
3
forks
#81
rank
View on GitHub →