Skip to content
— CH. 1 · INTRODUCTION —

Convolutional neural network

10 min listen · Ch. 1 of 6
6 sections
  • Convolutional neural networks power the face-unlock on your phone, flag tumors in mammograms, and helped a machine beat the world's best Go players. Yet the core idea behind them traces back to experiments on cats in the 1950s and 1960s, when researchers David Hubel and Wiesel studied how individual neurons in a cat's visual cortex responded to nothing more than light falling on a tiny patch of the animal's visual field. That biological observation, painstakingly replicated in laboratories over decades, became the blueprint for an entire class of artificial networks that now underpin computer vision. How did a curiosity about cat eyes become the de-facto standard architecture for image recognition? And what exactly happens inside one of these networks when it looks at a photograph?

  • Hubel and Wiesel's 1968 paper identified two distinct types of neurons in the visual brain. Simple cells gave their strongest response to straight edges oriented in a specific direction inside their receptive field. Complex cells had larger receptive fields and were indifferent to the exact position of an edge. The researchers also proposed a cascading model in which these two cell types could work together for pattern recognition.

    Kunihiko Fukushima took that cascading idea and built a computational counterpart. In 1969 he introduced a multilayer visual feature-detection network in which every element within a given layer shared the same interconnecting coefficients. That shared-weight arrangement is the essential core of what we now call a convolutional network, though in this early version the weights were not trained by any algorithm. In the same paper, Fukushima also introduced the ReLU activation function, a mathematical operation that would later become standard across the entire field.

    Fukushima went further in 1980 with the neocognitron. It formalized two types of layers that persist in modern architectures: an S-layer, which is a shared-weights receptive-field layer later known as a convolutional layer, and a C-layer, a downsampling layer whose units compute a weighted average of activations in a patch. The downsampling and competitive inhibition in the C-layer allowed the network to recognize features even when objects shifted position in the image. Separate supervised and unsupervised learning algorithms were proposed over the following decades to train the neocognitron's weights, though today CNN architectures are almost universally trained through backpropagation.

  • The term convolution first appeared in a neural-network paper at the Conference on Neural Information Processing Systems in 1987, written by Toshiteru Homma, Les Atlas, and Robert Marks II. They replaced multiplication with convolution in the time domain, providing shift invariance and linking the approach to signal-processing concepts. Their demonstration ran on a speech recognition task.

    Also in 1987, Alex Waibel and colleagues introduced the time delay neural network, or TDNN, for phoneme recognition. It was the first convolutional network to combine weight sharing with training by gradient descent via backpropagation. TDNNs processed speech signals time-invariantly, and a 1990 variant by Hampshire and Waibel added a second dimension of convolution so the system became invariant to both time and frequency shifts.

    Yann LeCun and colleagues in 1989 applied backpropagation to learn convolution kernel coefficients directly from images of hand-written numbers, making learning fully automatic. That same year, Denker and colleagues had designed a two-dimensional CNN to recognize hand-written ZIP codes, but had to hand-design all the kernel coefficients for lack of an efficient training method. LeCun's automatic approach outperformed manual coefficient design and proved suited to a broader range of image types.

    Wei Zhang and colleagues had independently used backpropagation to train convolution kernels for alphabet recognition in 1988, calling the model a shift-invariant pattern recognition neural network before the CNN name was coined in the early 1990s. They later removed the final fully connected layer and applied the same architecture to medical image segmentation in 1991 and to breast cancer detection in mammograms in 1994.

  • LeNet-5, a seven-level convolutional network published by LeCun and colleagues in 1995, classified hand-written numbers on checks digitized in 32x32-pixel images. It outperformed other commercial check-reading systems available at the time. NCR integrated it into check-reading systems deployed at several American banks starting in June 1996, where it read millions of checks per day.

    For roughly two decades after that, CNNs remained computationally expensive. The breakthrough in the 2000s required fast implementations on graphics processing units. In 2004, K. S. Oh and K. Jung showed that standard neural networks could run about 20 times faster on a GPU than on a CPU. The first GPU implementation of a CNN specifically was described in 2006 by K. Chellapilla and colleagues, who achieved a four-times speedup over an equivalent CPU implementation.

    Dan Ciresan and colleagues at IDSIA trained deep feedforward networks on GPUs in 2010, and extended the approach to CNNs in 2011, reporting a 60-times speedup over CPU training. That same year their network won an image recognition contest, achieving superhuman performance for the first time. AlexNet, a GPU-based CNN by Alex Krizhevsky and colleagues, then won the ImageNet Large Scale Visual Recognition Challenge in 2012 and has been described as an early catalytic event for the broader AI boom.

  • A CNN processes data as a tensor with four dimensions: number of inputs, height, width, and channels. Each convolutional layer slides a small filter, called a kernel, across the input to produce a feature map. A 5x5 kernel, for instance, requires only 25 shared weights, compared to the 10,000 weights that each neuron in a fully connected layer would need to process a 100x100 image. That dramatic compression is what allows CNNs to grow deeper without the vanishing-gradient problems that plagued earlier networks.

    Pooling layers follow convolutional layers and reduce the spatial dimensions of feature maps. The most common form takes the maximum value in each small region, typically a 2x2 window with a stride of 2, which discards 75% of activations. Max pooling was introduced by Yamaguchi and colleagues in 1990 as part of a speaker-independent word recognition system. J. Weng and colleagues introduced it into the vision field in 1993 through a neocognitron variant called the cresceptron.

    The ReLU activation function removes negative values from feature maps by setting them to zero, introducing nonlinearity without distorting the receptive fields of convolution layers. In 2011, Xavier Glorot, Antoine Bordes, and Yoshua Bengio found that ReLU enabled better training of deeper networks compared to activation functions that had been dominant before that year. ReLU trains networks several times faster than those older alternatives without a significant penalty to accuracy.

    Fully connected layers sit at the end of the network and handle final classification, connecting every neuron to every neuron in the adjacent layer. Regularization techniques combat overfitting throughout training: dropout, introduced in 2014, randomly drops individual nodes from the network during each training stage with probability of 0.5, effectively forcing the network to learn more robust features that transfer better to new data.

  • In 2012, an error rate of 0.23% on the MNIST handwriting database was reported using a CNN. At the ImageNet Large Scale Visual Recognition Challenge in 2014, almost every highly ranked team used a CNN as their core framework. The winning network, GoogLeNet, applied more than 30 layers and raised the mean average precision of object detection to 0.439329 while cutting classification error to 0.06656. Performance at that level was close to human accuracy on the same test set.

    In December 2014, Clark and Storkey published a paper showing that a CNN trained on a database of human professional Go games could outperform the traditional program GNU Go. A subsequent large 12-layer CNN correctly predicted the professional move in 55% of positions, matching the accuracy of a 6-dan human player, and beat GNU Go in 97% of games when used directly to play without any search. AlphaGo, which became the first system to beat the best human Go player, used two CNNs: one as a policy network to choose candidate moves and one as a value network to evaluate board positions.

    In 2015, Atomwise introduced AtomNet, the first deep learning neural network for structure-based drug design. The system learns chemical features such as aromaticity, sp3 carbons, and hydrogen bonding from three-dimensional representations of chemical interactions, in the same way image networks compose local features into larger structures. AtomNet was subsequently used to predict candidate treatments for Ebola and multiple sclerosis.

    From 1999 to 2001, Fogel and Chellapilla published papers showing how a CNN could learn to play checkers through co-evolution, without ever studying prior human games. Their program, Blondie24, was tested on 165 games and ranked in the highest 0.4% of players. It also earned a win against the program Chinook at its expert level of play. The same architecture that reads mammograms and recognizes speech also learned checkers strategy from scratch.

Common questions

What is a convolutional neural network and how does it work?

A convolutional neural network (CNN) is a type of feedforward neural network that learns features by optimizing small filters, or kernels, that slide across input data to produce feature maps. Pooling layers then reduce the spatial dimensions of those maps, and fully connected layers perform final classification. Shared weights across the filters dramatically reduce the number of parameters compared to traditional fully connected networks.

Who invented the convolutional neural network?

The foundational architecture traces to Kunihiko Fukushima, who introduced the neocognitron in 1980 with its shared-weights convolutional layers and downsampling layers. Yann LeCun and colleagues in 1989 made training fully automatic by applying backpropagation to learn kernel coefficients directly from images, and formalized this in the seven-layer LeNet-5 network published in 1995.

What was the biological inspiration for convolutional neural networks?

CNNs were inspired by work Hubel and Wiesel conducted in the 1950s and 1960s on the cat visual cortex. Their 1968 paper identified simple cells, which respond to edges of specific orientations within a small receptive field, and complex cells, which have larger receptive fields and are insensitive to exact edge position. Fukushima used this cascading model as the direct blueprint for the neocognitron.

When did convolutional neural networks become practical with GPU acceleration?

The first GPU implementation of a CNN was described in 2006 by K. Chellapilla and colleagues, running four times faster than a CPU equivalent. In 2011, Dan Ciresan and colleagues at IDSIA achieved a 60-times speedup over CPU training and won an image recognition contest with superhuman performance. AlexNet's win at the ImageNet Large Scale Visual Recognition Challenge in 2012 is widely cited as the catalytic event for the modern AI boom.

What are the main applications of convolutional neural networks?

CNNs are used in image and video recognition, medical image analysis, natural language processing, drug discovery, financial time series analysis, brain-computer interfaces, recommender systems, and game playing. In 2015, Atomwise's AtomNet used CNNs for structure-based drug design, predicting candidate treatments for Ebola and multiple sclerosis.

How did AlphaGo use convolutional neural networks to beat human Go players?

AlphaGo used two CNNs driving a Monte Carlo tree search: a policy network to select candidate moves and a value network to evaluate board positions. Before AlphaGo, a 12-layer CNN trained on professional human games had already correctly predicted the professional move in 55% of positions, matching the accuracy of a 6-dan human player, and beat the traditional program GNU Go in 97% of games.

All sources

168 references cited across the entry

  1. 1JournalDeep learningYann LeCun et al. — 2015-05-28
  2. 2JournalBackpropagation Applied to Handwritten Zip Code RecognitionY. LeCun et al. — December 1989
  3. 3BookConvolutional Neural Networks in Visual Computing: A Concise GuideRagav Venkatesan et al. — CRC Press — 2017-10-23
  4. 4BookRecent Trends and Advances in Artificial Intelligence and Internet of ThingsValentina E. Balas et al. — Springer Nature — 2019-11-19
  5. 5JournalPowder-Bed Fusion Process Monitoring by Machine Vision With Hybrid Convolutional Neural NetworksYingjie Zhang et al. — September 2020
  6. 7BookGuide to convolutional neural networks: a practical application to traffic-sign detection and classificationHamed Habibi Aghdam et al. — Springer — 2017-05-30
  7. 8JournalApplication of the residue number system to reduce hardware costs of the convolutional neural network implementationM.V. Valueva et al. — Elsevier BV — 2020
  8. 9BookDeep content-based music recommendationAaron van den Oord et al. — Curran Associates, Inc. — 2013-01-01
  9. 10BookProceedings of the 25th international conference on Machine learning - ICML '08Ronan Collobert et al. — ACM — 2008-01-01
  10. 12Book2017 IEEE 19th Conference on Business Informatics (CBI)Avraam Tsantekidis et al. — IEEE — July 2017
  11. 15BookArtificial Intelligence ResearchCoenraad Mouton et al. — Springer International Publishing — 2020
  12. 16JournalHidden bias in the DUD-E dataset leads to misleading performance of deep learning in structure-based virtual screeningThomas Kurtzman — August 20, 2019
  13. 17JournalNeocognitronK. Fukushima — 2007
  14. 21Xception: Deep Learning with Depthwise Separable ConvolutionsFrançois Chollet — 2017-04-04
  15. 23A Neural Network for Speaker-Independent Isolated Word RecognitionKouichi Yamaguchi et al. — November 1990
  16. 24Book2012 IEEE Conference on Computer Vision and Pattern RecognitionDan Ciresan et al. — Institute of Electrical and Electronics Engineers (IEEE) — June 2012
  17. 25Multi-Scale Context Aggregation by Dilated ConvolutionsFisher Yu et al. — 2016-04-30
  18. 26Rethinking Atrous Convolution for Semantic Image SegmentationLiang-Chieh Chen et al. — 2017-12-05
  19. 27Contextual Convolutional Neural NetworksIonut Cosmin Duta et al. — 2021-08-16
  20. 29Book2011 International Conference on Computer VisionMatthew D. Zeiler et al. — IEEE — November 2011
  21. 30A guide to convolution arithmetic for deep learningVincent Dumoulin et al. — 2018-01-11
  22. 31JournalDeconvolution and Checkerboard ArtifactsAugustus Odena et al. — 2016-10-17
  23. 32JournalComparing Object Recognition in Humans and Deep Convolutional Neural Networks—An Eye Tracking StudyLeonard Elia van Dyck et al. — 2021
  24. 33JournalReceptive fields and functional architecture of monkey striate cortexD. H. Hubel et al. — 1968-03-01
  25. 34BookBrain and visual perception: the story of a 25-year collaborationDavid H. Hubel and Torsten N. Wiesel — Oxford University Press US — 2005
  26. 35JournalReceptive fields of single neurones in the cat's striate cortexDH Hubel et al. — October 1959
  27. 36JournalVisual feature extraction by a multilayered network of analog threshold elementsK. Fukushima — 1969
  28. 37Annotated History of Modern AI and Deep LearningJuergen Schmidhuber — 2022
  29. 39Searching for Activation FunctionsPrajit Ramachandran et al. — October 16, 2017
  30. 43Convolutional networks for images, speech, and time seriesYann LeCun et al. — The MIT press — 1995
  31. 48Book1993 (4th) International Conference on Computer VisionJ Weng et al. — IEEE — 1993
  32. 49JournalDeep LearningJürgen Schmidhuber — 2015
  33. 50BookLearning algorithms for classification: A comparison on handwritten digit recognitionLecun, Y. et al. — World Scientific — August 1995
  34. 51JournalGradient-based learning applied to document recognitionY. Lecun et al. — November 1998
  35. 58JournalGPU implementation of neural networks.KS Oh et al. — 2004
  36. 5912th International Conference on Document Analysis and Recognition (ICDAR 2005)Dave Steinkraus et al. — 2005
  37. 60BookTenth International Workshop on Frontiers in Handwriting RecognitionKumar Chellapilla et al. — Suvisoft — 2006
  38. 61JournalA fast learning algorithm for deep belief nets.GE Hinton et al. — Jul 2006
  39. 62JournalGreedy Layer-Wise Training of Deep NetworksYoshua Bengio et al. — 2007
  40. 64BookProceedings of the 26th Annual International Conference on Machine LearningR Raina et al. — ICML '09: Proceedings of the 26th Annual International Conference on Machine Learning — 14 June 2009
  41. 65JournalDeep big simple neural nets for handwritten digit recognition.Dan Ciresan et al. — 2010
  42. 69JournalCHAOS: a parallelization scheme for training convolutional neural networks on Intel Xeon PhiAndre Viebke et al. — 2019
  43. 702015 IEEE 17th International Conference on High Performance Computing and Communications, 2015 IEEE 7th International Symposium on Cyberspace Safety and Security, and 2015 IEEE 12th International Conference on Embedded Software and SystemsAndre Viebke et al. — IEEE 2015 — 2015
  44. 72BookHands-on Machine Learning with Scikit-Learn, Keras, and TensorFlowAurélien Géron — O'Reilly Media — 2019
  45. 73JournalA Survey of Convolutional Neural Networks: Analysis, Applications, and ProspectsZewen Li et al. — December 2022
  46. 75JournalPooling in convolutional neural networks for medical image analysis: a survey and an empirical studyRajendran Nirthika et al. — 2022-04-01
  47. 78Fractional Max-PoolingBenjamin Graham — 2014-12-18
  48. 79Striving for Simplicity: The All Convolutional NetJost Tobias Springenberg et al. — 2014-12-21
  49. 80JournalFine-Grained Vehicle Classification With Channel Max Pooling Modified CNNsZhanyu Ma et al. — Institute of Electrical and Electronics Engineers (IEEE) — 2019
  50. 81JournalA Comparison of Pooling Methods for Convolutional Neural NetworksAfia Zafar et al. — 2022-08-29
  51. 82Pooling Methods in Deep Neural Networks, a ReviewHossein Gholamalinezhad et al. — 2020-09-16
  52. 84JournalImageNet classification with deep convolutional neural networksAlex Krizhevsky et al. — 2017-05-24
  53. 85JournalAppropriate number and allocation of ReLUs in convolutional neural networksVadim Romanuke — 2017
  54. 86Deep sparse rectifier neural networksXavier Glorot et al. — 2011
  55. 88BookICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)Antonio H. Ribeiro et al. — 2021
  56. 89BookArtificial Intelligence ResearchJohannes C. Myburgh et al. — Springer International Publishing — 2020
  57. 90BookMaking Convolutional Networks Shift-Invariant AgainZhang Richard — 2019-04-25
  58. 91JournalSpatial Transformer NetworksMax Jadeberg et al. — 2015
  59. 92BookDynamic Routing Between CapsulesSara Sabour et al. — 2017-10-26
  60. 94JournalDeep Learning With Conformal Prediction for Hierarchical Analysis of Large-Scale Whole-Slide Tissue ImagesHåkan Wieslander et al. — February 2021
  61. 97Stochastic Pooling for Regularization of Deep Convolutional Neural NetworksMatthew D. Zeiler et al. — 2013-01-15
  62. 99Improving neural networks by preventing co-adaptation of feature detectorsGeoffrey E. Hinton et al. — 2012
  63. 104JournalFace Recognition: A Convolutional Neural Network ApproachSteve Lawrence — 1997
  64. 107IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2015, Boston, MA, USA, June 7–12, 2015Christian Szegedy et al. — IEEE Computer Society — 2015
  65. 108Image Net Large Scale Visual Recognition ChallengeOlga Russakovsky et al. — 2014
  66. 110BookHuman Behavior UnterstandingMoez Baccouche et al. — Springer Berlin Heidelberg — 2011-11-16
  67. 111Journal3D Convolutional Neural Networks for Human Action RecognitionShuiwang Ji et al. — 2013-01-01
  68. 112Video-based Sign Language Recognition without Temporal SegmentationJie Huang et al. — 2018
  69. 114Two-Stream Convolutional Networks for Action Recognition in VideosKaren Simonyan et al. — 2014
  70. 1162018 25th IEEE International Conference on Image Processing (ICIP)Xuhuan Duan et al. — 25th IEEE International Conference on Image Processing (ICIP) — 2018
  71. 117Convolutional Learning of Spatio-temporal FeaturesGraham W. Taylor et al. — Springer-Verlag — 2010-01-01
  72. 118BookCVPR 2011Q. V. Le et al. — IEEE Computer Society — 2011-01-01
  73. 119A Deep Architecture for Semantic ParsingEdward Grefenstette et al. — 2014-04-29
  74. 121A Convolutional Neural Network for Modelling SentencesNal Kalchbrenner et al. — 2014-04-08
  75. 122Convolutional Neural Networks for Sentence ClassificationYoon Kim — 2014-08-25
  76. 124Natural Language Processing (almost) from ScratchRonan Collobert et al. — 2011-03-02
  77. 125Comparative study of CNN and RNN for natural language processingW Yin et al. — 2017-03-02
  78. 126An empirical evaluation of generic convolutional and recurrent networks for sequence modelingS. Bai et al. — 2018
  79. 127JournalDetecting dynamics of action in text with a recurrent neural networkN. Gruber — 2021
  80. 128JournalApproximation Theory of Convolutional Architectures for Time Series ModellingJ. Haotian et al. — 2021
  81. 129JournalDeepEthogram, a machine learning pipeline for supervised behavior classification from raw pixelsJames P Bohnslav et al. — 2021-09-02
  82. 130JournalAutomated monitoring of honey bees with barcodes and artificial intelligence reveals two distinct social networks from a single affiliative behaviorTim Gernat et al. — 2023-01-27
  83. 131JournalAutomatically identifying, counting, and describing wild animals in camera-trap images with deep learningMohammad Sadegh Norouzzadeh et al. — 2018-06-19
  84. 132A General Method for Detection and Segmentation of Terrestrial Arthropods in ImagesAsger Svenning et al. — 2025-04-14
  85. 133New idtracker.ai: rethinking multi-animal tracking as a representation learning problem to increase accuracy and reduce tracking timesJordi Torrents et al. — 2025-06-02
  86. 135JournalDeepPoseKit, a software toolkit for fast and robust animal pose estimation using deep learningJacob M Graving et al. — 2019-10-01
  87. 136JournalPublisher Correction: SLEAP: A deep learning system for multi-animal pose trackingTalmo D. Pereira et al. — May 2022
  88. 137JournalDeepBehavior: A Deep Learning Toolbox for Automated Analysis of Animal and Human Behavior Imaging DataAhmet Arac et al. — 2019-05-07
  89. 138Time-Series Anomaly Detection Service at Microsoft Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data MiningHansheng Ren et al. — 2019
  90. 139AtomNet: A Deep Convolutional Neural Network for Bioactivity Prediction in Structure-based Drug DiscoveryIzhar Wallach et al. — 2015-10-09
  91. 140Understanding Neural Networks Through Deep VisualizationJason Yosinski et al. — 2015-06-22
  92. 143JournalEvolving neural networks to play checkers without relying on expert knowledgeK Chellapilla et al. — 1999
  93. 144JournalEvolving an expert checkers playing program without using human expertiseK. Chellapilla et al. — 2001
  94. 145BookBlondie24: Playing at the Edge of AIDavid Fogel — Morgan Kaufmann — 2001
  95. 146Teaching Deep Convolutional Neural Networks to Play GoChristopher Clark et al. — 2014
  96. 147Move Evaluation in Go Using Deep Convolutional Neural NetworksChris J. Maddison et al. — 2014
  97. 149An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence ModelingShaojie Bai et al. — 2018-04-19
  98. 150Conditional Time Series Forecasting with Convolutional Neural NetworksAnastasia Borovykh et al. — 2018-09-17
  99. 151Time-series modeling with undecimated fully convolutional neural networksRoni Mittelman — 2015-08-03
  100. 152Probabilistic Forecasting with Temporal Convolutional Neural NetworkYitian Chen et al. — 2019-06-11
  101. 153JournalConvolutional neural networks for time series classiBendong Zhao et al. — 2017-02-01
  102. 154QCNN: Quantile Convolutional Neural NetworkGábor Petneházi — 2019-08-21
  103. 156NIPS 20172017-10-20
  104. 157BookArtificial Intelligence Applications and InnovationsJinliang Zang et al. — Springer International Publishing — 2018
  105. 159Distributed Deep Q-LearningHao Yi Ong et al. — 2015-08-18
  106. 161JournalSelf-segmentation of sequences: automatic formation of hierarchies of sequential behaviorsR. Sun et al. — June 2000
  107. 163BookProceedings of the 26th Annual International Conference on Machine LearningHonglak Lee et al. — ACM — 1 January 2009
  108. 164BookHierarchical Neural Networks for Image InterpretationSven Behnke — Springer — 2003
  109. 166Proceedings of the 15th International Conference on Document Analysis and Recognition (ICDAR)2019
  110. 167HeiCuBeDa Hilprecht – Heidelberg Cuneiform Benchmark Dataset for the Hilprecht CollectionheiDATA – institutional repository for research data of Heidelberg University — 2019-06-07
  111. 168Period Classification of 3D Cuneiform Tablets with Geometric Neural NetworksBartosz Bogacz et al. — 2020