Skip to content
— CH. 1 · INTRODUCTION —

ImageNet

15 min listen · Ch. 1 of 7
7 sections
  • ImageNet sorts photographs into thousands of categories, from something as ordinary as a balloon to something as specific as a strawberry. For each one, it records exactly what the picture shows. ImageNet does not own most of the pictures in its collection. It only keeps a directory of links to images hosted elsewhere, free for anyone to use. Once a year, a contest built on top of that directory asks software to sort a narrowed set of one thousand categories. The results have been credited with helping launch an industry-wide race toward artificial intelligence. Where did a library like this come from? And why would the people who built it later start cutting pieces of it away?

  • Fei-Fei Li began working on the idea for ImageNet in 2006, at a time when most AI research focused on models and algorithms rather than data. She wanted to expand and improve the raw material available for training those algorithms.

    In 2007, Li met with Princeton professor Christiane Fellbaum, one of the creators of WordNet, to discuss the project. The conversation set her building plan: start from WordNet's roughly 22,000 nouns and borrow many of its existing features. A 1987 estimate that the average person recognizes roughly 30,000 kinds of objects gave her a sense of the scale required.

    As an assistant professor at Princeton, Li assembled a team of researchers and turned to Amazon Mechanical Turk to help classify images. Labeling started in July 2008 and wrapped up in April 2010. Across that stretch, 49,000 workers spread across 167 countries filtered and labeled more than 160 million candidate images. The budget stretched far enough to have each image checked three times.

    The original plan called for 10,000 images in each of 40,000 categories, or 400 million images total, with each one verified three times. Tests showed a person can classify at most two images a second, which put the full plan at an estimated 19 human-years of nonstop labor.

    The team first showed their work as a poster at the 2009 Conference on Computer Vision and Pattern Recognition in Florida. It was titled "ImageNet: A Preview of a Large-scale Hierarchical Dataset," and the same poster reappeared months later at the Vision Sciences Society meeting.

    In 2009, researcher Alex Berg suggested adding object localization as a task, and Li approached organizers of the PASCAL Visual Object Classes contest about teaming up. That collaboration produced an annual challenge, launched in 2010 with 1,000 classes and object localization included. PASCAL VOC that same year had just 20 classes and 19,737 images.

    By 2012, ImageNet had become the largest academic user of Mechanical Turk in the world, with the average worker there identifying 50 images a minute.

  • Image-level annotations record only whether an object class shows up in a picture. The source describes it plainly: "there are tigers in this image" or "there are no tigers in this image." Object-level annotations go a step further, drawing a bounding box around the visible part of that object.

    Across its full library, ImageNet holds more than 20,000 categories, and a typical one holds several hundred images apiece. ImageNet uses a variant of the broad WordNet schema to organize its objects, adding 120 categories of dog breeds specifically to demonstrate fine-grained classification.

    A snapshot recorded on the 30th of April 2010 counted 1,034,908 images carrying bounding box annotations. The same snapshot logged 1,000 synsets with SIFT features, spanning 1.2 million images in total.

    WordNet 3.0 held more than 100,000 synsets, most of them nouns, upward of 80,000 by that count. ImageNet kept only the synsets that were countable nouns capable of being visually illustrated. A synset can bundle several synonyms together, as with "kitty" and "young cat," both folded into a single concept.

    Each synset in WordNet 3.0 carries a WordNet ID, or wnid, built by joining a part of speech with a unique identifying number called an offset. Every wnid begins with the letter n, since ImageNet includes only nouns; the synset for "dog, domestic dog, Canis familiaris" carries the wnid n02084071.

    ImageNet's categories are arranged across nine levels, moving from broad groupings down to increasingly narrow, specific ones.

    Images were scraped from online search engines, including Google, Picsearch, MSN, Yahoo and Flickr, using synonyms gathered in multiple languages. For the category "German shepherd," those synonyms included German police dog, Alsatian, the Spanish "ovejero alemán," the Italian "pastore tedesco," and the Chinese "德国牧羊犬."

    Every image is stored in RGB format, though resolutions vary widely. In the "fish" category of ImageNet 2012, for instance, sizes ranged from 4288 by 2848 pixels down to just 75 by 56. Machine learning pipelines typically resize and whiten these images before feeding them to a neural network. In PyTorch, ImageNet pixel values are first divided down to a range between 0 and 1. They are then adjusted using the means 0.485, 0.456 and 0.406, and the standard deviations 0.229, 0.224 and 0.225.

    Each image in the dataset carries exactly one wnid as its label. Bounding boxes were made available for about 3,000 popular synsets, averaging 150 images apiece. Some images also carry attributes, with 25 of them released for roughly 400 popular synsets. Those attributes span colors such as black, blue and violet, along with patterns like spotted and striped. They also cover shapes such as long and rectangular, and textures such as furry, shiny and wooden.

    That same labeling scheme would end up spread across several differently sized editions of the dataset.

  • The complete original collection is known as ImageNet-21K, holding 14,197,122 images across 21,841 classes. Some papers round that up and call it ImageNet-22k. The full set was released in the fall of 2011 as a single archive, fall11_whole.tar, with no official split between training, validation and test data. Class sizes varied widely, with some containing as few as 1-10 samples while others held thousands.

    One of the most heavily used subsets is the ILSVRC 2012-2017 image classification and localization dataset, known in research literature as ImageNet-1K or ILSVRC2017. It contains 1,281,167 training images, 50,000 validation images and 100,000 test images. Every category in ImageNet-1K is a leaf category, with no child nodes below it, while ImageNet-21K still keeps broader parent categories sitting above narrower ones. Dense SIFT features, meaning raw descriptors, quantized codewords and their coordinates, were released separately for ImageNet-1K, built for use in bag-of-visual-words models.

    ImageNet-C is an adversarially perturbed version of the dataset, built in 2019. ImageNetV2 followed the same methodology to build three separate test sets of 10,000 images each. ImageNet-21K-P is a filtered and cleaned subset containing 12,358,688 images across 11,221 categories, with every image resized to 224 by 224 pixels.

    None of these editions matched the roughly 50 million images across 50,000 synsets that the earliest plans had imagined. That original ambition was never completed.

  • On the 30th of September 2012, a convolutional neural network called AlexNet achieved a top-5 error of 15.3% in that year's challenge. That result was more than 10.8 percentage points lower than the runner-up. Training networks like it only became practical once researchers started using graphics processing units, an ingredient central to what became known as the deep learning revolution. The Economist captured the shift at the time: "Suddenly people started to pay attention, not just within the AI community but across the technology industry as a whole."

    By 2015, AlexNet had already been outperformed by a Microsoft network more than 100 layers deep, which won that year's contest with a 3.57% error rate.

    In 2014, researcher Andrej Karpathy estimated that with concentrated personal effort he could reach a 5.1% error rate on the same task. About 10 people from his lab reached roughly 12-13% with less effort, and maximal human effort was estimated to top out around 2.4%.

    Those same percentages just described one contest year among many. The competition itself ran for most of a decade, crowning a new winner almost every time.

  • The trimmed list used in competition years included just 1,000 image categories, among them 90 of the 120 dog breeds recognized across the full ImageNet schema. The decade that followed saw dramatic progress in image processing methods.

    The first competition in 2010 drew 11 participating teams, and a linear support vector machine took the win. Its features were a dense grid of HoG and LBP descriptors, sparsified through local coordinate coding and pooling. That system reached 52.9% classification accuracy and 71.8% top-5 accuracy. It trained for four days on three eight-core machines, built around dual quad-core 2 GHz Intel Xeon processors.

    The second competition in 2011 drew fewer teams, and another SVM won, this time at a 25% top-5 error rate. The winning team, XRCE, was built by Florent Perronnin and Jorge Sanchez, running a linear SVM over quantized Fisher vectors that reached 74.2% top-5 accuracy.

    In 2012, a deep convolutional neural network overtook the field, while Oxford's VGG team took second place using the previous generation's toolkit: SVMs, SIFT, color statistics and Fisher vectors. Over the next couple of years, top-5 accuracy across the competition climbed above 90%. Commentators later noted that the 2012 breakthrough "combined pieces that were all there before," yet its scale of improvement still marked the start of an industry-wide artificial intelligence boom.

    In 2013, most of the highest-ranking entries relied on convolutional neural networks. OverFeat, an architecture built for simultaneous classification and localization, won the object localization track, while Clarifai's ensemble of multiple CNNs won classification.

    By 2014, more than 50 institutions were entering the competition. GoogLeNet won the classification track that year, and VGGNet won localization.

    In 2015, ResNet won the competition and exceeded human performance on the benchmark. Olga Russakovsky, one of the challenge's organizers, pointed out that year that the contest covered only 1,000 categories, while humans can recognize far more and judge an image's context in ways the programs could not.

    In 2016, an ensemble called CUImage won, combining six separate networks: Inception v3, Inception v4, Inception ResNet v2, ResNet 200, Wide ResNet 68 and Wide ResNet 3. The runner-up, ResNeXt, combined the Inception module with ResNet's architecture.

    In 2017, the Squeeze-and-Excitation Network, or SENet, won by cutting the top-5 error down to 2.251%. That same year, 29 of the 38 competing teams cleared 95% accuracy. The organizers announced that 2017 would be the competition's last year, since they judged the benchmark solved and no longer challenging. They floated a harder follow-up for 2018, built around classifying 3D objects using natural language. Since 3D data cost more to gather, the new dataset was expected to be smaller. Its applications were expected to range from robotic navigation to augmented reality. That planned 3D competition never materialized.

  • Researchers estimate that more than 6% of labels in the ImageNet-1k validation set are wrong. A further study found that around 10% of ImageNet-1k contains ambiguous or erroneous labels. When shown both a model's prediction and the original ImageNet label side by side, human annotators tended to prefer the model's prediction. That model, trained on the original ImageNet-1k in 2020, suggested the dataset had become saturated.

    A 2019 study traced the layered history of ImageNet and WordNet's taxonomy, object classes and labeling. It described how bias runs deep in most classification approaches, across all kinds of images, and noted that ImageNet has continued working to address it.

    WordNet's "person" subtree, the branch ImageNet was built on, originally contained 2,832 synsets. Between 2018 and 2020, the project removed the ImageNet-21k download entirely while filtering that branch. Of those 2,832 synsets, 1,593 were judged "potentially offensive." Of the 1,239 that remained, 1,081 were judged not truly "visual," leaving just 158 synsets, only 139 of which held more than 100 images suitable for further use.

    In the winter of 2021, ImageNet-21k was updated again, removing 2,702 categories from the "person" subtree to prevent "problematic behaviors" in any model trained on it. Only 130 person-related synsets remained afterward. That same year, ImageNet-1k was updated by blurring faces across its 997 non-person categories. Out of 1,431,093 images in ImageNet-1k, 243,198 of them, about 17%, were found to contain at least one face, adding up to 562,626 faces in total. Researchers found that training models with those faces blurred out caused only a minimal loss in performance.

    One acknowledged downside of building on WordNet is that its categories can skew more "elevated" than practical for everyday use. As the project's own assessment put it: "Most people are more interested in Lady Gaga or the iPod Mini than in this rare kind of diplodocus."

Up Next

Common questions

Who created ImageNet?

Fei-Fei Li began developing ImageNet in 2006 and, starting in 2007, worked with Princeton professor Christiane Fellbaum, one of WordNet's creators, to build the project from WordNet's roughly 22,000 nouns.

What is ImageNet used for?

ImageNet is a large visual database used to train and test software that recognizes objects in images. It is organized into more than 20,000 categories and is used to run the annual ImageNet Large Scale Visual Recognition Challenge.

When did the ImageNet Large Scale Visual Recognition Challenge start?

The ILSVRC launched in 2010 as a collaboration between Fei-Fei Li and organizers of the PASCAL Visual Object Classes contest, using a trimmed list of 1,000 classes.

How many images are in ImageNet's original dataset?

ImageNet's original collection, known as ImageNet-21K, contains 14,197,122 images across 21,841 classes, gathered through crowdsourced labeling between July 2008 and April 2010.

Why was AlexNet significant for ImageNet?

On the 30th of September 2012, the convolutional neural network AlexNet reached a top-5 error of 15.3% in the ImageNet challenge, more than 10.8 percentage points better than the runner-up. The result is credited with helping start the deep learning revolution.

What bias issues has ImageNet faced?

Studies found more than 6% of labels in the ImageNet-1k validation set are wrong and about 10% of the set contains ambiguous or erroneous labels. ImageNet later removed thousands of "person" subtree categories and blurred faces across its non-person categories after review found many were "potentially offensive" or not truly visual.

All sources

62 references cited across the entry

  1. 2NewsFor Web Images, Creating New Technology to Seek and FindJohn Markoff — 19 November 2012
  2. 3ImageNet2020-09-07
  3. 6MagazineFei-Fei Li's Quest to Make AI Better for HumanityJesse Hempel — 13 November 2018
  4. 9Where have we been? Where are we going?Fei-Fei Li et al. — 2017
  5. 13The data that transformed AI research—and possibly the worldDave Gershgorn — Atlantic Media Co. — 26 July 2017
  6. 142009 conference on Computer Vision and Pattern RecognitionJia Deng et al. — 2009
  7. 17JournalImageNet classification with deep convolutional neural networksAlex Krizhevsky et al. — June 2017
  8. 20Book2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)Kaiming He et al. — 2016
  9. 23JournalImageNet Large Scale Visual Recognition ChallengeOlga Russakovsky et al. — 2015-12-01
  10. 28ImageNet2013-04-05
  11. 32BookTrends and Topics in Computer VisionOlga Russakovsky et al. — Springer — 2012
  12. 33ImageNet-21K Pretraining for the MassesTal Ridnik et al. — 2021-08-05
  13. 35BookProceedings of the 2020 Conference on Fairness, Accountability, and TransparencyKaiyu Yang et al. — ACM — 2020-01-27
  14. 38JournalA Study of Face Obfuscation in ImageNetKaiyu Yang et al. — PMLR — 2022-06-28
  15. 39Benchmarking Neural Network Robustness to Common Corruptions and PerturbationsDan Hendrycks et al. — 2019
  16. 40JournalDo ImageNet Classifiers Generalize to ImageNet?Benjamin Recht et al. — PMLR — 2019-05-24
  17. 42BookCVPR 2011Yuanqing Lin et al. — IEEE — June 2011
  18. 43BookCVPR 2011Jorge Sanchez et al. — IEEE — June 2011
  19. 44BookComputer Vision – ECCV 2010Florent Perronnin et al. — Springer — 2010
  20. 48OverFeat: Integrated Recognition, Localization and Detection using Convolutional NetworksPierre Sermanet et al. — 2013
  21. 49Book2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)Christian Szegedy et al. — IEEE — June 2015
  22. 55Squeeze-and-Excitation NetworksJie Hu et al. — 2017
  23. 56Pervasive Label Errors in Test Sets Destabilize Machine Learning BenchmarksCurtis G. Northcutt et al. — 2021-11-07
  24. 57Are we done with ImageNet?Lucas Beyer et al. — 2020-06-12
  25. 60Excavating AI: The Politics of Training Sets for Machine LearningKate Crawford et al. — 19 September 2019
  26. 61Excavating "Excavating AI": The Elephant in the GalleryMichael J. Lyons — 2020