Blog
Blog
Training a 19M Parameter Story Generator from Scratch on 2xT4 GPUs
April 01, 2026
We trained a GPT language model from scratch on the TinyStories dataset using flaxchat, our JAX/Flax NNX port of nanochat. The entire pipeline ran end-to-end on 2x NVIDIA T4 GPUs via Kaggle with data-parallel training, generating coherent children’s stories in under an hour.