← All projectsML Engineering

Makemore

PythonPyTorchNeural Network ImplementationPaper ReadingNatural Language ProcessingDeep Learning

The problem

Understanding a language model requires more than calling an API. I wanted to follow the mathematics from a bigram baseline to a working GPT.

How I solved it

Built a character-level GPT from scratch using PyTorch, following Karpathy's Makemore progression to understand the fundamentals behind ChatGPT. Developed models from bigram to neural networks, implementing core operations from scratch to understand the math. Used techniques like He initialization and batch normalization to solve vanishing gradients. Trained a character-level GPT on Indonesian Twitter poems.

  • Built GPT implementation based on Karpathy's Makemore
  • Developed models from bigram to neural networks
  • Implemented core operations from scratch
  • Used He initialization and batch normalization
  • Trained on Indonesian Twitter poems dataset

The results

parameters built and trained from scratch
10M+
character context window
256

Delivered a configurable GPT module with six Transformer blocks, six attention heads per block and a 256-character context window, plus training/inference scripts and generated poetry examples.

Evaluation context. Tutorial-inspired by Andrej Karpathy. The reviewed 384-dimensional, six-block architecture has over 10 million trainable parameters (10,738,944 plus 769 per vocabulary character). Counts describe this configuration, not a generation-quality benchmark or every saved checkpoint.

πŸ€—