The Fine-Tuning Index / Distillation / #62

yangqi0/petitgpt

by yangqi0 · Distillation · updated today

A 124.6M LLM trained from scratch on 13B tokens on a single RTX 4090 — tokenizer, pretraining, post-training, evaluation, and reproducible inference.

50
momentum
41
stars
7
forks
#62
rank
distillationfrom-scratchgenerative-aiinstruction-tuninglanguage-modelllmmodel-evaluationpost-trainingpretrainingpytorchreproducibilitysingle-gpu
View on GitHub →