Neural network (machine learning)
Neural networks have become the engine behind voice assistants, cancer diagnoses, deepfake videos, and the large language models now reshaping how people interact with machines. But the idea behind them is older, stranger, and more biologically grounded than most people realize. At its core, a neural network is a computational model inspired by the structure of the brain itself. Neurons, synapses, signals, weights: the vocabulary of neuroscience was consciously transplanted into mathematics and software. How did that transplant take hold? What kept it from working for decades? And what finally unlocked its extraordinary power?
Warren McCulloch and Walter Pitts, in 1943, published the first mathematical model of an artificial neuron capable of representing logical functions. Their paper considered neural networks that contain cycles and noted that the current activity of such networks could be affected by activity indefinitely far in the past. That single observation planted the seed for what would later become recurrent networks.
The connection to biological structure was not merely metaphorical. Each artificial neuron receives numerical inputs, weights them, sums them with a bias term, and then passes the result through a nonlinear activation function to produce an output. The edges connecting neurons model the synapses in the brain. The layered structure, with an input layer, hidden layers, and an output layer, mirrors the way biological neural circuits process information from the sensory periphery inward.
D. O. Hebb, in the late 1940s, proposed a learning hypothesis based on neural plasticity. That principle, known as Hebbian learning, held that connections between neurons strengthen when those neurons fire together. It became the basis for early experiments including Rosenblatt's perceptron and the Hopfield network. The brain's architecture was not just an inspiration; it was an active blueprint that researchers kept returning to, even as the field grew increasingly mathematical.
Frank Rosenblatt, a psychologist, described the perceptron in 1958. It was one of the first implemented neural networks and was funded by the United States Office of Naval Research. The announcement triggered public excitement, and the US government drastically increased funding for the field. Researchers spoke optimistically about perceptrons eventually emulating human intelligence, a period later called "the Golden Age of AI".
That confidence collapsed. Marvin Minsky and Seymour Papert published a book called Perceptrons that highlighted a fundamental flaw: single-layer perceptrons could only solve linearly separable problems. They specifically emphasized that perceptrons were incapable of processing the exclusive-or circuit. The book deflated interest through the late 1960s and into the 1970s.
What the critics missed is that work on deeper networks had already been underway. Alexey Ivakhnenko and Valentin Lapa published the first working deep learning algorithm in the Soviet Union in 1965, a method they regarded as a form of polynomial regression generalizing Rosenblatt's perceptron. A 1971 paper described a network with eight layers trained by this method. Shun'ichi Amari published the first deep learning multilayer perceptron trained by stochastic gradient descent in 1967. The Minsky-Papert critique was simply irrelevant to these deeper architectures, but the damage to funding and public confidence had already been done.
Interest in neural networks revived during the 1980s because of a new algorithm: backpropagation. The algorithm works by propagating error gradients backward from the output layer through to the input layer, adjusting the weights at each step. This made it possible to train multi-layer networks efficiently for the first time.
The history of the algorithm is tangled. Henry J. Kelley developed a precursor in 1960 within control theory. Seppo Linnainmaa published the modern form in his master's thesis in 1970. G. M. Ostrovski et al. republished it in 1971. Paul Werbos applied it explicitly to neural networks in 1982. David E. Rumelhart and colleagues popularized it in 1986 but did not cite the original work. The terminology "back-propagating errors" had actually been introduced by Rosenblatt himself as early as 1962, though he did not describe how to implement it. Backpropagation is mathematically an efficient application of the chain rule, which Gottfried Wilhelm Leibniz derived in 1673.
Once widely adopted, backpropagation opened the door to training the multi-layer architectures that single-layer critics had dismissed. It also set the stage for the GPU-accelerated training pipelines that would follow. From 1991 to 2015, computing power delivered by GPUs increased around a million-fold, enabling backpropagation to scale to networks of previously unimaginable depth.
Kunihiko Fukushima introduced the neocognitron in 1979, a deep convolutional architecture with convolutional layers, downsampling layers, weight replication, and max pooling. That structure became the foundation for convolutional neural networks, which proved transformative for computer vision.
Yann LeCun and colleagues built LeNet in 1989 to recognize handwritten ZIP codes on mail; training required three days. By 1998, LeNet-5, a seven-level convolutional network, was being applied by banks to recognize numbers on checks digitized in 32 by 32 pixel images. In October 2012, AlexNet, built by Alex Krizhevsky, Ilya Sutskever, and Geoffrey Hinton, won the large-scale ImageNet competition by a significant margin over shallow machine learning methods. In 2011, DanNet had already achieved superhuman performance in a visual pattern recognition contest, outperforming traditional methods by a factor of three.
Long short-term memory, or LSTM, addressed a different problem. Sepp Hochreiter's diploma thesis in 1991 identified the vanishing gradient problem that crippled deep recurrent networks. He and Jürgen Schmidhuber introduced LSTM, which set accuracy records across multiple domains. The forget gate, added in 1999, completed the architecture and made it the default choice for recurrent network design for years.
Then in 2017, the paper "Attention Is All You Need" introduced the transformer architecture, which uses attention mechanisms to model long-range dependencies in data. Modern large language models including GPT, Gemini, Grok, DeepSeek, and Qwen are all built on this foundation.
Training a neural network means adjusting the weights of its connections until the network's outputs match desired targets closely enough to be useful. The process relies on a loss function that measures the degree of error between what the network predicts and what it should predict. As long as the loss function continues to decline, the network is improving. Training ends when additional observations no longer usefully reduce the cost.
A hyperparameter is any configurable part of the network whose value is set before training begins: learning rate, batch size, number of layers, and number of nodes are all examples. The learning rate controls how large each corrective step is. A high learning rate shortens training time but sacrifices accuracy; a lower rate takes longer but allows greater precision. Adaptive learning rates, which increase or decrease as appropriate, help avoid oscillating between high and low weight values.
Networks face a structural hazard called overfitting: when the network is so large relative to the training data that it memorizes the training examples rather than learning to generalize. Cross-validation and regularization techniques address this. In probabilistic frameworks, selecting a larger prior probability over simpler models reduces the tendency to overfit. Neural networks also require far more sample inputs than biological brains need to reach a given level of function. As of 2026, training a commercial large language model typically required hundreds of thousands of computers and cost tens of millions of dollars.
From 1988 onward, neural networks transformed the field of protein structure prediction, particularly when cascading networks began training on profiles produced by multiple sequence alignments. That application arrived well before the modern deep learning wave.
In medicine, neural networks have been used to diagnose cancers and to distinguish highly invasive cancer cell lines from less invasive ones using only cell shape data. They have accelerated reliability analysis of infrastructure subject to natural disasters and predict settling in building foundations. In cybersecurity, they classify Android malware, identify domains belonging to threat actors, and detect URLs posing a security risk.
Generative adversarial networks, introduced by Ian Goodfellow and colleagues in 2014, drove state-of-the-art image generation through 2018. Nvidia's StyleGAN, released in 2018 and based on the Progressive GAN by Tero Karras and colleagues, achieved excellent image quality and provoked widespread discussion about deepfakes. Diffusion models, emerging in 2015, eventually surpassed GANs in generative modeling, with systems such as DALL-E 2 and Stable Diffusion both arriving in 2022.
DALL-E was trained on 650 million pairs of images and texts and can create artworks from user text descriptions. Companies including AIVA and Jukedeck have applied transformer architectures to generate original music. Film production companies have used neural networks to analyze the likely financial success of a film. In 2018, Amazon scrapped a recruiting tool after discovering the model favored men over women for software engineering roles because the training data itself reflected a workforce that was predominantly male; the system penalized resumes containing the word "woman" or the name of a women's college.
The multilayer perceptron is a universal function approximator, as established by the universal approximation theorem. A recurrent architecture with rational-valued weights can match the power of a universal Turing machine using a finite number of neurons and linear connections. Using irrational values for weights produces a machine with what researchers describe as super-Turing power.
Despite this theoretical reach, neural networks face real constraints. The brain operates on roughly 20 watts of power. Training commercial transformer models in 2026 required data centers drawing hundreds of megawatts. The statistical properties of real-world data shift over time, a phenomenon called concept drift, which can cause a network's accuracy in deployment to diverge substantially from what was measured during training. Neural networks are also "black box" systems: their internal decision-making processes remain difficult to interpret, and they are vulnerable to adversarial examples that can cause incorrect predictions through deliberately crafted inputs.
These concerns have spurred research into explainable artificial intelligence and hybrid models that combine neural learning with symbolic reasoning. Neuromorphic engineering has pursued a different path entirely, building non-von-Neumann chips designed to implement neural networks directly in circuitry; Alphabet introduced custom chips called Tensor Processing Units as one example of this direction. The vanishing gradient problem that Hochreiter identified in 1991 took years to solve; the interpretability problem has not yet found its equivalent breakthrough.
Common questions
What is a neural network in machine learning?
A neural network is a computational model inspired by the structure and functions of biological neural networks. It consists of connected artificial neurons arranged in layers, where signals travel from an input layer through hidden layers to an output layer, with weights on each connection adjusted during training to improve accuracy.
Who invented the first neural network?
Warren McCulloch and Walter Pitts published the first mathematical model of artificial neurons in 1943, capable of representing logical functions. Frank Rosenblatt described the perceptron in 1958, one of the first implemented neural networks, funded by the United States Office of Naval Research.
What is backpropagation and who developed it?
Backpropagation is an algorithm that trains neural networks by propagating error gradients backward from the output layer to the input layer to adjust connection weights. Seppo Linnainmaa published the modern form in his master's thesis in 1970; David E. Rumelhart and colleagues popularized it in 1986.
What is the difference between a convolutional neural network and a recurrent neural network?
Convolutional neural networks, whose deep architecture originated with Kunihiko Fukushima's neocognitron in 1979, are specialized for spatially structured data like images and are the essential tool for computer vision. Recurrent neural networks allow connections between neurons in the same or previous layers, making them suited for sequential data such as speech and time series; the LSTM architecture, introduced by Sepp Hochreiter and Jürgen Schmidhuber, became the default recurrent design after the forget gate was added in 1999.
What is the transformer architecture and why does it matter?
The transformer architecture was introduced in 2017 in the paper "Attention Is All You Need" and uses attention mechanisms to model long-range dependencies in data. Large language models including GPT, Gemini, Grok, DeepSeek, and Qwen are all built on this architecture.
How much does it cost to train a large language model in 2026?
As of 2026, training a commercial large language model such as GPT, Grok, or Gemini typically required hundreds of thousands of computers and cost tens of millions of dollars. Training these models requires data centers drawing hundreds of megawatts of power, compared to the roughly 20 watts the human brain consumes.
All sources
234 references cited across the entry
- 1Explained: Neural networksLarry Hardesty — MIT News Office — 14 April 2017
- 2BookComprehensive Biomedical PhysicsZ.R. Yang et al. — Elsevier — 2014
- 3BookPattern Recognition and Machine LearningChristopher M. Bishop — Springer — 17 August 2006
- 4JournalAttention Is All You NeedAshish Vaswani et al. — 2017
- 5BookNeural Networks for BabiesFerrie, C. — Sourcebooks — 2019
- 6JournalGauss and the Invention of Least SquaresStephen M. Stigler — 1981
- 7Annotated History of Modern AI and Deep LearningJuergen Schmidhuber — 2025-12-29
- 8BookThe History of Statistics: The Measurement of Uncertainty before 1900Stephen M. Stigler — Harvard — 1986
- 9JournalA logical calculus of the ideas immanent in nervous activityWarren S. McCulloch et al. — December 1943
- 10NewsRepresentation of Events in Nerve Nets and Finite AutomataS.C. Kleene — Princeton University Press — 1956
- 11BookThe Organization of BehaviorDonald Hebb — Taylor & Francis — 2005
- 12JournalSimulation of Self-Organizing Systems by Digital ComputerB.G. Farley — 1954
- 13JournalTests on a cell assembly theory of the action of the brain, using a large digital computerN. Rochester — 1956
- 14JournalThe Perceptron: A Probabilistic Model For Information Storage And Organization in the BrainF. Rosenblatt — 1958
- 15BookBeyond Regression: New Tools for Prediction and Analysis in the Behavioral SciencesP.J. Werbos — 1975
- 16JournalThe Perceptron—a perceiving and recognizing automatonFrank Rosenblatt — Cornell Aeronautical Laboratory — 1957
- 17JournalA Sociological Study of the Official History of the Perceptrons ControversyMikel Olazaran — 1996
- 18ThesisContributions to perceptron theoryR. D. Joseph — Cornell University — 1961
- 19BookPrinciples of Neurodynamics: Perceptrons and the Theory of Brain MechanismsFrank Rosenblatt — Spartan Books — 1962
- 20BookArtificial Intelligence A Modern ApproachRussel — Pearson Education — 2010
- 21A Proposal for the Dartmouth Summer Research Project on Artificial IntelligenceJohn McCarthy et al. — 1955
- 22BookPerceptrons: An Introduction to Computational GeometryMarvin Minsky et al. — MIT Press — 1969
- 23BookCybernetics and Forecasting TechniquesAlexey G. Ivakhnenko et al. — American Elsevier Publishing Co. — 1967
- 24JournalHeuristic self-organization in problems of engineering cyberneticsA.G. Ivakhnenko — March 1970
- 25JournalPolynomial theory of complex systemsAlexey Ivakhnenko — 1971
- 26JournalA Stochastic Approximation MethodH. Robbins et al. — 1951
- 27JournalA theory of adaptive pattern classifiersShun'ichi Amari — 1967
- 28JournalVisual feature extraction by a multilayered network of analog threshold elementsK. Fukushima — 1969
- 29JournalNeural network with unbounded activation functions is universal approximatorSho Sonoda et al. — 2017
- 30Searching for Activation FunctionsPrajit Ramachandran et al. — 16 October 2017
- 31BookPerceptrons: An Introduction to Computational GeometryMarvin Minsky et al. — MIT Press — 1969
- 32JournalReminder of the First Paper on Transfer Learning in Neural Networks, 1976Stevo Bozinovski — 2020-09-15
- 33JournalLearning representations by back-propagating errorsDavid Rumelhart et al. — 1986
- 34BookThe Early Mathematical Manuscripts of Leibniz: Translated from the Latin Texts Published by Carl Immanuel Gerhardt with Critical and Historical Notes (Leibniz published the chain rule in a 1676 memoir)Gottfried Wilhelm Freiherr von Leibniz — Open court publishing Company — 1920
- 35JournalGradient theory of optimal flight pathsHenry J. Kelley — 1960
- 36ThesisThe representation of the cumulative rounding error of an algorithm as a Taylor expansion of the local rounding errorsSeppo Linnainmaa — University of Helsinki — 1970
- 37JournalTaylor expansion of the accumulated rounding errorSeppo Linnainmaa — 1976
- 38Who Invented Backpropagation?Juergen Schmidhuber — IDSIA, Switzerland — 25 October 2014
- 39Applications of advances in nonlinear sensitivity analysisPaul J. Werbos — Springer-Verlag — 1982
- 40BookTalking Nets: An Oral History of Neural NetworksThe MIT Press — 2000
- 41BookThe Roots of Backpropagation: From Ordered Derivatives to Neural Networks and Political ForecastingPaul J. Werbos — John Wiley & Sons — 1994
- 42JournalLearning representations by back-propagating errorsDavid E. Rumelhart et al. — October 1986
- 43JournalNeural network model for a mechanism of pattern recognition unaffected by shift in position—NeocognitronK. Fukushima — 1979
- 45JournalDeep Learning in Neural Networks: An OverviewJ. Schmidhuber — 2015
- 46JournalNeocognitron: A new algorithm for pattern recognition tolerant of deformations and shifts in positionKunihiko Fukushima et al. — 1 January 1982
- 47Phoneme Recognition Using Time-Delay Neural NetworksAlex Waibel — December 1987
- 49JournalShift-invariant pattern recognition neural network and its optical architectureWei Zhang — 1988
- 51JournalImage processing of human corneal endothelium based on a learning networkWei Zhang — 1991
- 53JournalGradient-based learning applied to document recognitionYann LeCun — 1998
- 54JournalLearning Patterns and Pattern Sequences by Self-Organizing Nets of Threshold ElementsS.-I. Amari — November 1972
- 55JournalNeural networks and physical systems with emergent collective computational abilitiesJ. J. Hopfield — 1982
- 56JournalThe Importance of Cajal's and Lorente de Nó's Neuroscience to the Birth of CyberneticsJuan Manuel Espinosa-Sanchez et al. — 5 July 2023
- 59JournalFeeling and thinking: Preferences need no inferencesR. Zajonc — 1980
- 60JournalThoughts on the relations between emotion and cognition.Richard S. Lazarus — November 1982
- 61JournalModeling Mechanisms of Cognition-emotion Interaction in Artificial Neural Networks, since 1981Stevo Bozinovski — 2014
- 62JournalNeural Sequence ChunkersJürgen Schmidhuber — April 1991
- 63JournalLearning complex, extended sequences using the principle of history compression (based on TR FKI-148, 1991)Jürgen Schmidhuber — 1992
- 64BookHabilitation thesis: System modeling and optimizationJürgen Schmidhuber — 1993
- 66BookA Field Guide to Dynamical Recurrent NetworksS. Hochreiter — John Wiley & Sons — 15 January 2001
- 67JournalLong Short-Term MemorySepp Hochreiter et al. — 1 November 1997
- 68Book9th International Conference on Artificial Neural Networks: ICANN '99Felix Gers et al. — 1999
- 69JournalA learning algorithm for boltzmann machinesDavid H. Ackley et al. — 1 January 1985
- 70BookParallel Distributed Processing: Explorations in the Microstructure of Cognition, Volume 1: FoundationsPaul Smolensky — MIT Press — 1986
- 71JournalThe Helmholtz machine.Dayan Peter et al. — 1995
- 72JournalThe wake-sleep algorithm for unsupervised neural networksGeoffrey E. Hinton et al. — 26 May 1995
- 73How bio-inspired deep learning keeps winning competitions « the Kurzweil LibraryJuergen Schmidhuber
- 75JournalDeep, Big, Simple Neural Nets for Handwritten Digit RecognitionDan Claudiu Cireşan et al. — 21 September 2010
- 76JournalFlexible, High Performance Convolutional Neural Networks for Image ClassificationD. C. Ciresan et al. — 2011
- 77BookAdvances in Neural Information Processing Systems 25Dan Ciresan et al. — Curran Associates, Inc. — 2012
- 78BookMedical Image Computing and Computer-Assisted Intervention – MICCAI 2013D. Ciresan et al. — 2013
- 79Book2012 IEEE Conference on Computer Vision and Pattern RecognitionD. Ciresan et al. — 2012
- 80JournalImageNet Classification with Deep Convolutional Neural NetworksAlex Krizhevsky et al. — 2012
- 81Very Deep Convolution Networks for Large Scale Image RecognitionKaren Simonyan et al. — 2014
- 82JournalGoing deeper with convolutionsChristian Szegedy — 2015
- 83Building High-level Features Using Large Scale Unsupervised LearningAndrew Ng et al. — 2012
- 84BookDeep LearningIan Goodfellow and Yoshua Bengio and Aaron Courville — MIT Press — 2016
- 85BookNonlinear System Identification: NARMAX Methods in the Time, Frequency, and Spatio-Temporal DomainsS. A. Billings — Wiley — 2013
- 86Generative Adversarial NetworksIan Goodfellow et al. — 2014
- 87A possibility for implementing curiosity and boredom in model-building neural controllersJürgen Schmidhuber — MIT Press/Bradford Books — 1991
- 88JournalGenerative Adversarial Networks are Special Cases of Artificial Curiosity (1990) and also Closely Related to Predictability Minimization (1991)Jürgen Schmidhuber — 2020
- 89GAN 2.0: NVIDIA's Hyperrealistic Face Generator14 December 2018
- 90Progressive Growing of GANs for Improved Quality, Stability, and VariationT. Karras et al. — 26 February 2018
- 91Prepare, Don't Panic: Synthetic Media and Deepfakeswitness.org
- 92JournalDeep Unsupervised Learning using Nonequilibrium ThermodynamicsJascha Sohl-Dickstein et al. — PMLR — 1 June 2015
- 93Very Deep Convolutional Networks for Large-Scale Image RecognitionKaren Simonyan et al. — 10 April 2015
- 94Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet ClassificationKaiming He et al. — 2016
- 95Deep Residual Learning for Image RecognitionKaiming He et al. — 10 December 2015
- 96Highway NetworksRupesh Kumar Srivastava et al. — 2 May 2015
- 97Book2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR)Kaiming He et al. — IEEE — 2016
- 98Microsoft researchers win ImageNet computer vision challengeAllison Linn — 10 December 2015
- 99JournalLearning to control fast-weight memories: an alternative to recurrent nets.Jürgen Schmidhuber — 1992
- 100Transformers are RNNs: Fast autoregressive Transformers with linear attentionAngelos Katharopoulos et al. — PMLR — 2020
- 101Linear Transformers Are Secretly Fast Weight ProgrammersImanol Schlag et al. — Springer — 2021
- 102BookProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System DemonstrationsThomas Wolf et al. — 2020
- 103JournalApplication of Artificial Intelligence to the Management of Urological CancerMaysam F. Abbod — 2007
- 104JournalAn artificial neural network approach to rainfall-runoff modellingChristian W. Dawson — 1998
- 106JournalWhat Is Machine Learning, Artificial Neural Networks and Deep Learning?-Examples of Practical Applications in MedicineJakub Kufel et al. — 2023-08-03
- 107JournalA brief review of feed-forward neural networksMurat H. Sazlı — 2006-01-01
- 108BookArtificial intelligenceAddison-Wesley Pub. Co — 1992
- 109BookSimulation neuronaler NetzeAndreas Zell — Addison-Wesley — 2003
- 110JournalA Survey of Partially Connected Neural NetworksD. Elizondo et al. — October 1997
- 111JournalFlexible, High Performance Convolutional Neural Networks for Image ClassificationDan Ciresan — 2011
- 112Lecture Notes: Neural Network ArchitecturesEvelyn Herberg — 2023-04-18
- 113BookSimulation Neuronaler NetzeAndreas Zell — Addison-Wesley — 1994
- 114JournalComparative analysis of Recurrent and Finite Impulse Response Neural Networks in Time Series PredictionMilos Miljanovic — February–March 2012
- 115BookThe nature of statistical learning theoryVladimir N. Vapnik et al. — Springer — 1998
- 116BookFundamentals of machine learning for predictive data analytics: algorithms, worked examples, and case studiesJohn D. Kelleher et al. — The MIT Press — 2020
- 117JournalA comprehensive survey of loss functions and metrics in deep learningJuan Terven et al. — 2025-04-11
- 118JournalExtreme learning machine: theory and applicationsGuang-Bin Huang et al. — 2006
- 119JournalThe no-prop algorithm: A new learning algorithm for multilayer neural networksBernard Widrow — 2013
- 120Training recurrent networks without backtrackingYann Ollivier et al. — 2015
- 121JournalA Practical Guide to Training Restricted Boltzmann MachinesG. E. Hinton — 2010
- 122What Is Hyperparameter Tuning? IBM23 July 2024
- 123BookDeep learningIan Goodfellow et al. — The MIT press — 2016
- 124JournalHyperparameter optimization: Foundations, algorithms, best practices, and open challengesBernd Bischl et al. — March 2023
- 125Forget the Learning Rate, Decay LossJiakai Wei — 26 April 2019
- 126Book2009 International Conference on Computational Intelligence and Natural ComputingY. Li et al. — 1 June 2009
- 127BookIntroduction to machine learningEtienne Bernard — Wolfram Media — 2021
- 128BookIntroduction to Machine LearningEtienne Bernard — Wolfram Media Inc — 2021
- 129JournalMetaheuristic design of feedforward neural networks: A review of two decades of researchVarun Kumar Ojha et al. — 1 April 2017
- 130Genetic reinforcement learning for neural networksDominic, S. — IEEE — July 1991
- 131JournalProcess control via artificial neural networks and reinforcement learningJ.C. Hoskins — 1992
- 132BookNeuro-dynamic programmingD.P. Bertsekas et al. — Athena Scientific — 1996
- 133JournalComparing neuro-dynamic programming algorithms for the vehicle routing problem with stochastic demandsNicola Secomandi — 2000
- 134Neuro-dynamic programming for the efficient management of reservoir networksde Rigo, D. — Modelling and Simulation Society of Australia and New Zealand — 2001
- 135Genetic algorithms and neuro-dynamic programming: application to water supply networksDamas, M. — IEEE — 2000
- 136BookOptimization in MedicineGeng Deng — 2008
- 138JournalSelf-learning agents: A connectionist theory of emotion based on crossbar value judgmentStevo Bozinovski et al. — 2001
- 139Evolution Strategies as a Scalable Alternative to Reinforcement LearningTim Salimans et al. — 7 September 2017
- 140Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforcement LearningFelipe Petroski Such et al. — 20 April 2018
- 141NewsArtificial intelligence can 'evolve' to solve problems10 January 2018
- 142JournalA Learning Algorithm for Boltzmann MachinesDavid H. Ackley et al. — 1985
- 143BookStochastic Models of Neural NetworksClaudio Turchetti — IOS Press — 2004
- 144MagazineHands-On Bayesian Neural Networks—A Tutorial for Deep Learning UsersLaurent Valentin Jospin et al. — 2022
- 145JournalTopologyNet: Topology based deep convolutional and multi-task neural networks for biomolecular property predictionsZixuan Cang et al. — 27 July 2017
- 146A selective improvement technique for fastening Neuro-Dynamic Programming in Water Resources Network Managementde Rigo, D. et al. — IFAC — January 2005
- 147BookApplied Soft Computing Technologies: The Challenge of ComplexityC. Ferreira — Springer-Verlag — 2006
- 148An improved PSO-based ANN with simulated annealing techniqueDa, Y. — Elsevier — July 2005
- 150JournalA learning algorithm of CMAC based on RLSTing Qin et al. — 2004
- 151JournalContinuous CMAC-QRLS and its systolic arrayTing Qin et al. — 2005
- 152JournalTunability: Importance of Hyperparameters of Machine Learning AlgorithmsPhilipp Probst et al. — 26 February 2018
- 153Neural Architecture Search with Reinforcement LearningBarret Zoph et al. — 4 November 2016
- 154JournalAuto-keras: An efficient neural architecture search systemHaifeng Jin et al. — ACM — 2019
- 155Hyperparameter Search in Machine LearningMarc Claesen et al. — 2015
- 156JournalTuring computability with neural netsH.T. Siegelmann et al. — 1991
- 157NewsAnalog computer trumps Turing modelSunny Bains — 3 November 1998
- 158JournalComputational Power of Neural Networks: A Kolmogorov Complexity CharacterizationJosé Balcázar — July 1997
- 159BookProceedings of the 27th ACM International Conference on MultimediaFriedland Gerald — ACM — 2019
- 161BookInformation Theory, Inference, and Learning AlgorithmsDavid J.C. MacKay — Cambridge University Press — 2003
- 162JournalWide neural networks of any depth evolve as linear models under gradient descentJaehoon Lee et al. — 2020
- 163Neural Tangent Kernel: Convergence and Generalization in Neural NetworksArthur Jacot et al. — 2018
- 164BookNeural Information ProcessingXu ZJ, Zhang Y, Xiao Y — Springer, Cham — 2019
- 165JournalOn the Spectral Bias of Neural NetworksNasim Rahaman et al. — 2019
- 166Theory of the Frequency Principle for General Deep Neural NetworksTao Luo et al. — 2019
- 167JournalDeep Frequency Principle Towards Understanding Why Deeper Learning is FasterZhiqin John Xu et al. — 18 May 2021
- 168BookHandbook of Applied MathematicsRobin Esch — Springer US — 1990
- 169BookA Concise Guide to Market ResearchMarko Sarstedt et al. — Springer Berlin Heidelberg — 2019
- 170Book2016 IEEE Symposium Series on Computational Intelligence (SSCI)Jie Tian et al. — December 2016
- 171BookDynamic Data Assimilation – Beating the UncertaintiesWesam Salah Alaloul et al. — 2019
- 172Book2013 International Conference Oriental COCOSDA held jointly with 2013 Conference on Asian Spoken Language Research and Evaluation (O-COCOSDA/CASLRE)Madhab Pal et al. — IEEE — 2013
- 173JournalA cloud based architecture capable of perceiving and predicting multiple vessel behaviourDimitrios Zissis — October 2015
- 174JournalLung sound classification using cepstral-based statistical featuresNandini Sengupta — August 2016
- 1753D-R2N2: A Unified Approach for Single and Multi-view 3D Object ReconstructionChristopher B. Choy et al. — 2016
- 176JournalIntroduction to Neural Net Machine VisionTurek, Fred D. — March 2007
- 177Book2015 13th International Conference on Document Analysis and Recognition (ICDAR)Durjoy S. Maitra et al. — August 2015
- 178JournalSensor for food analysis applying impedance spectroscopy and artificial neural networksJosef Gessler — August 2021
- 179JournalThe time traveller's CAPMJordan French — 2016
- 180JournalNeural network approach to quantum-chemistry data: Accurate prediction of density functional theory energiesRoman M. Balabin et al. — 2009
- 181JournalMastering the game of Go with deep neural networks and tree searchDavid Silver — 2016
- 182NewsArtificial Intelligence Glossary: Neural Networks and Other Terms ExplainedAdam Pasick — 27 March 2023
- 183NewsFacebook Boosts A.I. to Block Terrorist PropagandaSam Schechner — 15 June 2017
- 184BookIntroduction to Artificial Intelligence: from data analysis to generative AIAlberto Ciaramella et al. — Intellisemantic Editions — 2024
- 185Book2026 IEEE Madhya Pradesh Section Conference (MPCON)Abhishek Maity — March 2026
- 186JournalApplication of Neural Networks in Diagnosing Cancer Disease Using Demographic DataN Ganesan — 2010
- 187JournalArtificial Neural Networks Applied to Outcome Prediction for Colorectal Cancer Patients in Separate InstitutionsLeonardo Bottaci — The Lancet — 1997
- 188JournalMeasuring systematic changes in invasive cancer cell shape using Zernike momentsElaheh Alizadeh et al. — 2016
- 189JournalChanges in cell shape are correlated with metastatic potential in murineSamanthe Lyons — 2016
- 190JournalDeep Learning for Accelerated Reliability Analysis of Infrastructure NetworksMohammad Amin Nabian et al. — 28 August 2017
- 191JournalAccelerating Stochastic Assessment of Post-Earthquake Transportation Network Connectivity via Machine-Learning-Based SurrogatesMohammad Amin Nabian et al. — 2018
- 192JournalUse of artificial neural networks to predict 3-D elastic settlement of foundations on soils with inclined bedrockE. Díaz et al. — September 2018
- 193JournalArtificial Neural Network for Modelling Rainfall-RunoffA. Tayebiyan et al.
- 194JournalArtificial Neural Networks in Hydrology. I: Preliminary ConceptsRao S. Govindaraju — 1 April 2000
- 195JournalArtificial Neural Networks in Hydrology. II: Hydrologic ApplicationsRao S. Govindaraju — 1 April 2000
- 196JournalSignificant wave height record extension by neural networks and reanalysis wind dataD. J. Peres et al. — 1 October 2015
- 197JournalReview on Applications of Neural Network in Coastal EngineeringG. S. Dwarakish et al. — 2013
- 198JournalArtificial Neural Networks applied to landslide susceptibility assessmentLeonardo Ermini et al. — 1 March 2005
- 199Book2017 International Joint Conference on Neural Networks (IJCNN)R. Nix et al. — May 2017
- 201BookCyber Threat IntelligenceSajad Homayoun et al. — Springer International Publishing — 2018
- 202BookProceedings of the Twenty-Seventh Hawaii International Conference on System Sciences HICSS-94Ghosh et al. — January 1994
- 203Latest Neural Nets Solve World's Hardest Equations Faster Than Ever BeforeAnil Ananthaswamy — 19 April 2021
- 205Caltech Open-Sources AI for Solving Partial Differential EquationsAnthony Alford
- 206JournalVariational Quantum Monte Carlo Method with a Neural-Network Ansatz for Open Quantum SystemsAlexandra Nagy — 28 June 2019
- 207JournalConstructing neural stationary states for open quantum many-body systemsNobuyuki Yoshioka et al. — 28 June 2019
- 208JournalNeural-Network Approach to Dissipative Quantum Many-Body DynamicsMichael J. Hartmann et al. — 28 June 2019
- 209JournalSimulation of alcohol action upon a detailed Purkinje neuron model and a simpler surrogate model that runs >400 times fasterForrest MD — April 2015
- 210JournalSemantic Image-Based Profiling of Users' Interests with Neural NetworksSzymon Wieczorek et al. — 2018
- 211JournalScaling deep learning for materials discoveryAmil Merchant et al. — December 2023
- 212JournalAdvances in Artificial Neural Networks – Methodological Development and ApplicationYanbo Huang — 2009
- 213JournalExploring the Advancements and Future Research Directions of Artificial Neural Networks: A Text Mining ApproachElham Kariri et al. — 2023
- 214JournalDeep neural networks for speech enhancement and speech recognition: A systematic reviewSureshkumar Natarajan et al. — 2025-07-01
- 215JournalDeep Learning for Intelligent Human–Computer InteractionZhihan Lv et al. — 2022-11-11
- 216What Is an Artificial Neural Network, and Why Does It Matter for AI?Coursera Staff — 2024-09-26
- 217JournalDeep networks for system identification: A surveyGianluigi Pillonetto et al. — 2025-01-01
- 218JournalA Survey of Data Mining and Machine Learning Methods for Cyber Security Intrusion DetectionAnna Buczak et al. — 2016
- 219JournalGenerative AI and ChatGPT: Applications, challenges, and AI-human collaborationFiona Fui-Hoon Nah et al. — 3 July 2023
- 221JournalFrom artificial neural networks to deep learning for music generation: history, concepts and trendsJean-Pierre Briot — January 2021
- 222JournalGhost in the (Hollywood) machine: Emergent applications of artificial intelligence in the film industryPei-Sze Chow — 6 July 2020
- 223BookThe 3rd International Conference on Information Sciences and Interaction SciencesXinrui Yu et al. — IEEE — June 2010
- 224JournalContinual lifelong learning with neural networks: A reviewGerman I. Parisi et al. — 1 May 2019
- 225JournalGrowing pains for deep learningChris Edwards — 25 June 2015
- 228NewsGoogle Built Its Very Own Chips to Power Its AI BotsCade Metz — 18 May 2016
- 229JournalStatistical process monitoring of artificial neural networksAnna Malinovskaya et al. — January 2024
- 230JournalAddressing bias in big data and AI for health care: A call for open scienceNatalia Norori et al. — October 2021
- 231JournalFailing at Face Value: The Effect of Biased Facial Recognition Technology on Racial Discrimination in Criminal JusticeWang Carina — 27 October 2022
- 232JournalGender Bias in Hiring: An Analysis of the Impact of Amazon's Recruiting AlgorithmXinyu Chang — 13 September 2023
- 233Book2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)Adam Kortylewski et al. — IEEE — June 2019
- 234JournalA hybrid neural networks-fuzzy logic-genetic algorithm for grade estimationTahmasebi et al. — 2012