The Fine-Tuning Index / Fine-Tuning Tools / #185

jinpz/q_sharp

by jinpz · Fine-Tuning Tools · updated 1y ago

The official code release for Q#: Provably Optimal Distributional RL for LLM Post-Training

15
momentum
21
stars
1
forks
#185
rank
View on GitHub →