In 1989, Yann LeCun and colleagues published what may be the earliest real-world application of a neural network trained end-to-end with backpropagation. The network had roughly 1,000 neurons, the dataset contained 7,291 grayscale 16x16 images of handwritten digits, and training took three days on a workstation. Thirty-three years later, the paper reads remarkably modern. It describes a dataset, an architecture, a loss function, optimization, and test error rates. Everything a contemporary deep learning paper contains, just smaller.
Andrej Karpathy reproduced the work in PyTorch to use it as a case study on the nature of progress in the field. The results are instructive.
Reproduction and Immediate Speedup
Karpathy's PyTorch implementation trained in about 90 seconds on an M1 MacBook Air, roughly a 3,000x speedup over the original. The reproduction achieved 4.09% test error versus the paper's reported 5.00%. An exact match proved impossible because the original dataset has been lost, forcing Karpathy to simulate it by downsampling MNIST. Other ambiguities lurked in the paper too: weight initialization was described as "2 4 / F" where F is fan-in, likely meaning 2.4 over square root of F, and the sparse connectivity between layers was "chosen according to a scheme that will not be discussed here."
Modern Techniques, Old Hardware
The most interesting experiment asked: if you could time travel to 1989 with modern knowledge, how much could you improve the result without changing the hardware or dataset? Karpathy swapped MSE regression for softmax cross-entropy, replaced SGD with AdamW, added data augmentation (1-pixel shifts), introduced dropout, and switched from tanh to ReLU. The error rate fell from 4.09% to roughly 1.5%, a 60% reduction. The cost was nearly quadrupling training time from 3 days to 12. Further gains would require a larger model and more compute.
Scaling the dataset from 7,291 to 50,000 examples, combined with the same modern techniques, pushed the error down to 1.25%. Data alone, without algorithmic changes, also helped significantly.
What Has Not Changed
The macro picture looks eerily stable. We still build differentiable neural architectures from layers of neurons and optimize them end-to-end with backpropagation and stochastic gradient descent. The 1989 net had 9,760 parameters and 64K MACs. Today's vision models run into the billions of parameters and trillions of operations. Datasets have grown by roughly 100 million times in total pixel information. A state-of-the-art classifier that took three days now trains in 90 seconds on a fanless laptop.
What a 2055 Time Traveler Would See
If the pattern holds, neural nets in 2055 will look basically the same as today, just bigger. Our current datasets and models will seem like jokes, probably 10 million times larger. A 2026 frontier model will train in about a minute on personal hardware. Today's models are not optimally formulated, and changing details of the architecture, loss, or optimizer could roughly halve the error. Further gains will require more compute and dedicated research on training at scale.
The Bigger Shift
The most important trend may be the decline of training from scratch. Foundation models trained by a few well-resourced institutions are increasingly the starting point for everyone else. Fine-tuning, prompt engineering, and distillation into smaller inference networks are becoming the norm. In its most extreme form, you may not train neural networks at all. You will ask a billion-parameter megabrain to perform a task in English, and if you ask nicely enough, it will oblige. You could train a network too. But why would you?