Skip to content
— CH. 1 · INTRODUCTION —

AlexNet

13 min listen · Ch. 1 of 8
8 sections
  • AlexNet entered the ImageNet Large Scale Visual Recognition Challenge on the 30th of September 2012, submitted by a three-person team calling itself SuperVision. The network posted a top-5 error rate of 15.3 percent, finishing more than 10.8 percent ahead of the next best entry. It came from a convolutional neural network built by graduate student Alex Krizhevsky. He worked with fellow graduate student Ilya Sutskever and his PhD advisor, Geoffrey Hinton, at the University of Toronto. AlexNet held 60 million parameters and 650,000 neurons, and it sorted images into 1,000 distinct object categories. It is often described as the first deep convolutional network to gain broad recognition for large-scale visual recognition. The original paper argued that depth was what made the model work. That depth came at a computational cost, one that only became affordable once the team turned to graphics processing units instead of relying on CPUs alone.

    What did it take to train a network of that size before anyone had a standard tool for doing so? And how did a single contest result end up reordering an entire research field? The path runs through a bedroom at Krizhevsky's parents' house. It also runs through a dataset that one researcher spent years trying to get others to take seriously.

  • Nvidia's GTX 580 GPU had only 3 gigabytes of video memory, not enough to hold AlexNet's full structure on its own. Krizhevsky split the network, except for the last layer, into two identical copies, one running on each of two GTX 580 cards. The result was eight layers in all. Five of them handled convolution, with pooling folded into some, and the final three were fully connected. Written out, the structure ran as two repeats of a convolution, normalization and pooling block. Next came a further convolution and pooling stage. Then two blocks of fully connected layers with dropout followed, ending in a final linear layer and a softmax output.

    Layers three, four and five broke from that pattern. They fed directly into each other with nothing in between: no pooling, no normalization. Instead of the tanh or sigmoid functions common at the time, AlexNet used the non-saturating ReLU activation function. The team found it trained the network faster and better than either alternative.

    Getting that structure to learn would demand training tricks nobody had needed at this scale before.

  • AlexNet trained against 1.2 million images, the full ImageNet training set. It worked through that set 90 times over. The whole process took five to six days on two Nvidia GTX 580 GPUs. Each GTX 580 offered a theoretical 1.581 teraflops of float32 performance and sold for about US$500 when it was released. A single forward pass through AlexNet took roughly 1.43 gigaflops. That meant the pair, working together, could in theory crank out more than 2,200 forward passes each second, given ideal conditions. Stored as JPEG files, the dataset filled 27 gigabytes on disk. Running the network required 2 gigabytes of memory on each GPU, plus roughly 5 gigabytes more in system RAM. The graphics cards handled all the training computation, while ordinary processors handled loading images off disk and applying augmentation on the fly.

    The team optimized the network with momentum-based gradient descent: a batch size of 128, momentum set to 0.9, and weight decay of 0.0005. The learning rate began at 10 to the negative 2. It was manually cut by a factor of ten three times whenever validation error stalled, ending at 10 to the negative 5. Two forms of data augmentation ran on the CPU at no extra computational cost. Every image was first rescaled so its shorter side measured 256 pixels. It was then cropped to a central 256 by 256 patch, then normalized using ImageNet's mean values of 0.485, 0.456 and 0.406, and standard deviations of 0.229, 0.224 and 0.225. From that crop, the team pulled random 224 by 224 patches, along with their horizontal reflections, a trick that expanded the effective training set 2,048-fold. They also nudged each image's RGB values along its three principal color directions. The 224 pixel size was not arbitrary: subtracting 16 pixels from each of the four sides of a 256 pixel image leaves exactly 224.

    Dropout regularization ran at a 0.5 probability, alongside local response normalization. Every weight started as a random value drawn from a gaussian distribution with a mean of zero and a standard deviation of 0.01. Biases in convolutional layers two, four and five, and in every fully connected layer, were fixed at 1 to head off the dying ReLU problem. At test time, a new image was scaled and cropped the same way. It was then split into five 224 by 224 patches, the four corners plus the center, along with their mirror images. That gave ten patches per image, and the network's ten resulting predictions were averaged into a single final answer.

    But the version submitted to the 2012 contest that September was not just this one network.

  • The entry that actually won the 2012 contest was not a single AlexNet but an ensemble of seven. Five of those seven were standard AlexNets, each trained the same way on the ILSVRC-2012 training set. The other two were variants, built by adding one extra convolutional layer above the last pooling layer. Those two variants were trained first on the entire ImageNet Fall 2011 release, a set of 15 million images spread across 22,000 categories. Only afterward were they finetuned on the smaller ILSVRC-2012 training data. The final system averaged the predicted probabilities from all seven networks to produce its winning entry.

    That averaging trick was a detail few in computer vision had used at this scale before. It came from a field where neural networks had spent two decades on the losing side of the argument.

  • In 1980, Kunihiko Fukushima proposed an early convolutional network called the neocognitron, trained through an unsupervised learning algorithm. Roughly a decade later, Yann LeCun and colleagues built LeNet-5 in 1989. They trained it through supervised learning with backpropagation, using an architecture essentially the same as AlexNet's, just at a much smaller scale. Max pooling itself dates to 1990, first applied to speech processing as a one-dimensional version of the technique. Its use in image processing followed in 1992, inside a system called Cresceptron.

    Through the 2000s, as graphics hardware improved, researchers began repurposing GPUs for general computing tasks, including training neural networks. In 2006, K. Chellapilla and colleagues trained a CNN on a GPU four times faster than an equivalent CPU setup. By 2009, Raina and colleagues had trained a deep belief network with 100 million parameters on an Nvidia GeForce GTX 280. It reached speedups of up to 70 times over CPU training. In 2011, a deep CNN built by Dan Cireșan and colleagues at IDSIA ran 60 times faster than its CPU equivalent. Between the 15th of May 2011 and the 10th of September 2012, that network won four separate image competitions. It also set the state of the art on multiple image databases. The AlexNet paper itself called Cireșan's earlier network "somewhat similar." Both were written in CUDA to run on GPU hardware.

    None of that history explains why computer vision as a field had spent two decades treating neural networks as a losing bet.

  • Between 1990 and 2010, neural networks routinely lost out to other machine learning methods: kernel regression, support vector machines, AdaBoost, and structured estimation among them. Progress in computer vision during that period came largely from manually engineered features, tools like SIFT, SURF, and HoG. Some of that work built bags of visual words from those features instead. The idea that a network could learn its own features directly from data was a minority position, one that only became dominant after AlexNet's win.

    In 2011, Geoffrey Hinton began asking colleagues a pointed question: "What do I have to do to convince you that neural networks are the future?" One of those colleagues, Jitendra Malik, was himself a sceptic of neural networks, and he suggested Hinton try the PASCAL Visual Object Classes challenge. Hinton judged that dataset too small to make the case. Malik then pointed him toward ImageNet instead.

    That dataset had been built by Fei-Fei Li and her collaborators starting in 2007, aiming to advance visual recognition through sheer scale of data. Li's project eventually gathered over 14 million labeled images across 22,000 categories, far larger than anything that had come before it. Workers on Amazon Mechanical Turk supplied the labels, and the WordNet hierarchy gave the whole set its structure. ImageNet was initially met with skepticism, but it went on to become the foundation of the ILSVRC contest and a central resource in deep learning's rise.

    None of that data would have meant anything without two graduate students willing to spend months making it run on hardware never meant for the job.

  • Ilya Sutskever and Alex Krizhevsky were both graduate students when the ImageNet idea took hold. Krizhevsky had already written cuda-convnet before 2011, a tool for training small convolutional networks on the CIFAR-10 dataset using a single GPU. Sutskever was convinced that Krizhevsky's skill with general-purpose GPU programming could scale further. He persuaded Krizhevsky to try training a CNN on ImageNet, with Geoffrey Hinton serving as principal investigator. Krizhevsky then extended cuda-convnet so it could train across multiple GPUs at once.

    The actual training happened on two Nvidia GTX 580 cards set up in Krizhevsky's bedroom, at his parents' house. Through 2012, Krizhevsky ran hyperparameter optimization on the network again and again, refining it until it won the ImageNet competition later that year. Hinton later summed up the division of labor bluntly: "Ilya thought we should do it, Alex made it work, and I got the Nobel Prize." At the 2012 European Conference on Computer Vision, held soon after AlexNet's win, Yann LeCun offered his own verdict. He called the model "an unequivocal turning point in the history of computer vision."

    AlexNet's win depended on three developments that had each been maturing for roughly a decade. Those were large labeled datasets, general-purpose GPU computing, and better training methods for deep networks. ImageNet supplied the data. Nvidia's CUDA platform made training large models practical. Algorithmic improvements did the rest, and together the three let AlexNet perform at a level nothing before it had reached. Looking back in a 2024 interview, Fei-Fei Li described the moment this way: "That moment was pretty symbolic to the world of AI because three fundamental elements of modern AI converged for the first time." AlexNet and LeNet, despite the gap of more than two decades between them, share essentially the same design and algorithm. What changed was scale. AlexNet was far larger, trained on a far larger dataset, and run on far faster hardware. That was the product of 20 years in which both data and computing power became cheap to obtain.

    That cheap compute and cheap data would soon let other teams build on AlexNet's design rather than start from nothing.

  • By early 2025, the AlexNet paper had been cited more than 184,000 times, according to Google Scholar. At the time it was first published, no ready-made framework existed for training or running neural networks on GPUs at all. The team released its own codebase under a BSD license, and it saw common use in neural network research for several years afterward.

    GoogLeNet appeared in 2014, followed the same year by VGGNet. Next came the Highway network in 2015 and ResNet later that year, together forming one line of research chasing deeper, more accurate networks. A second line of research, including SqueezeNet in 2016, MobileNet in 2017, and EfficientNet in 2019, aimed instead to match AlexNet's performance at a lower cost.

    Hinton, Sutskever, and Krizhevsky went on to form a company called DNNResearch, which they sold to Google along with AlexNet's source code. The original 2012 version of that code, the one that won the contest, remains available today under a BSD-2 license through the Computer History Museum.

Common questions

What is AlexNet?

AlexNet is a convolutional neural network that sorts images into 1,000 categories, built for large-scale visual recognition. It became known for winning the 2012 ImageNet Large Scale Visual Recognition Challenge with a top-5 error rate of 15.3 percent, and it is regarded as an early landmark for deep convolutional networks at that scale.

Who created AlexNet?

AlexNet was developed in 2012 by Alex Krizhevsky, working with fellow graduate student Ilya Sutskever and Krizhevsky's PhD advisor, Geoffrey Hinton, at the University of Toronto.

When did AlexNet win the ImageNet competition?

A team called SuperVision submitted AlexNet to the ImageNet Large Scale Visual Recognition Challenge on the 30th of September 2012. It won with a top-5 error rate of 15.3 percent, more than 10.8 percent ahead of the runner-up.

Where was AlexNet trained?

AlexNet was trained on two Nvidia GTX 580 GPUs set up in Alex Krizhevsky's bedroom at his parents' house, over a period of five to six days.

Why was AlexNet significant for computer vision?

AlexNet showed that a deep convolutional network could outperform hand-engineered feature methods like SIFT and SURF at large scale, a position that had been a minority view before its win. Its success came from combining a large labeled dataset, ImageNet, with GPU computing and improved training methods.

How many layers and parameters does AlexNet have?

AlexNet contains eight layers, five convolutional and three fully connected, with 60 million parameters and 650,000 neurons. Because the network was too large to fit on a single GPU's memory, it was split across two Nvidia GTX 580 cards.

All sources

25 references cited across the entry

  1. 1JournalImageNet classification with deep convolutional neural networksAlex Krizhevsky et al. — 2017-05-24
  2. 6JournalNeocognitronK. Fukushima — 2007
  3. 8Backpropagation Applied to Handwritten Zip Code RecognitionY. LeCun — MIT Press - Journals — 1989
  4. 10A Neural Network for Speaker-Independent Isolated Word RecognitionKouichi Yamaguchi et al. — November 1990
  5. 12BookTenth International Workshop on Frontiers in Handwriting RecognitionKumar Chellapilla et al. — Suvisoft — 2006
  6. 17Book2012 IEEE Conference on Computer Vision and Pattern RecognitionDan Cireșan et al. — Institute of Electrical and Electronics Engineers (IEEE) — June 2012
  7. 18JournalMax-Margin Markov NetworksBen Taskar et al. — MIT Press — 2003
  8. 19BookDive into deep learningAston Zhang et al. — Cambridge University Press — 2024
  9. 20BookThe worlds I see: curiosity, exploration, and discovery at the dawn of AIFei Fei Li — Moment of Lift Books; Flatiron Books — 2023
  10. 21CHM Releases AlexNet Source Codehhackford — 2025-03-20
  11. 25computerhistory/AlexNet-Source-CodeComputer History Museum — 2025-03-22