Watch one token flow through a 1,952-parameter mini GPT — every lookup, matrix multiply and softmax computed live in your browser
Transformer Visualisation walks one token through an entire working GPT. Not a diagram of a transformer — a complete, trained model, shrunk until every number fits on screen: 1,952 parameters, an 18-word vocabulary, 36 training sentences, and the exact architecture of GPT-2 scaled down to 2 blocks, 2 attention heads and width 8. Type a prompt, and every lookup, matrix multiplication, normalisation and softmax recomputes live in your browser from the trained weights.
The step-by-step explorer has 37 steps from tokenisation to sampling, plus a pretraining chapter showing one training step and a post-training chapter with one PPO/RLHF iteration. Hover over any number to see exactly how it was computed — all arithmetic in 64-bit floats, cells shown to 2 decimals. The model itself was trained with Adam for 6,000 full-batch steps in plain NumPy with a hand-written, gradient-checked backward pass, reaching a loss of 0.4994 nats per token — essentially perfect for the corpus.
The author, xtcntr, announced it on Hacker News on 2026-10-07 with the note "So here it is, made with Opus 5.5." As an education tool it's the rare artefact that treats the learner as an adult: no metaphor-only explanations, just the actual numbers, all of them, computed for real. Tech stack per the repo: TypeScript, Vite, NumPy, hosted on Vercel.
curated by the VibeFix editors