The Fine-Tuning Index / Fine-Tuning Tools / #185
jinpz/q_sharp
by jinpz · Fine-Tuning Tools · updated 1y ago
The official code release for Q#: Provably Optimal Distributional RL for LLM Post-Training
15
momentum
21
stars
1
forks
#185
rank
The official code release for Q#: Provably Optimal Distributional RL for LLM Post-Training