Skip to main content
ScienceVerse
Artificial IntelligenceDifficulty 2-4

Neural networks

An artificial neural network is a mathematical function built from many simple units arranged in layers: each unit multiplies its inputs by weights, adds a bias and passes the total through an activation function. Stacking layers lets the network represent patterns — such as XOR — that no single linear unit can. The weights are not written by hand but learned from data, usually by back-propagation and gradient descent; the networks are inspired by the brain but are not realistic models of it.

Sign in to save this concept.

Representation and training

A feedforward network composes affine maps and element-wise non-linearities. ReLU, g(z) = max(0, z), is the default hidden-unit choice in modern practice; softmax outputs convert a vector of scores into a probability distribution over classes.

softmax(z)i=ezi∑jezj\mathrm{softmax}(\mathbf{z})_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

Softmax maps scores z to probabilities that are positive and sum to 1.

The universal approximation theorem (Cybenko, 1989; Hornik et al., 1989, as presented by Goodfellow et al.) guarantees expressiveness for networks with a hidden layer and sufficient width, but not learnability: optimisation may fail to find the parameters, and the learned function may overfit. Back-propagation applies the chain rule layer by layer to obtain gradients efficiently; an optimiser such as stochastic gradient descent then updates the weights.

Common misconception: Interpreting individual hidden units as clean, human-readable concepts is unreliable: Szegedy et al. (2013) found no distinction between individual high-level units and random combinations of them under several analysis methods.
Full explanation — the complete reference version every reading depth is based on

One artificial neuron

A single unit (often called a neuron) takes some input numbers, multiplies each by a weight, adds them up together with a bias, and passes the result through an activation function. The weights say how much each input matters; the bias shifts the point at which the unit 'switches on'.

z=∑iwixi+b,a=g(z)z = \sum_{i} w_i x_i + b, \qquad a = g(z)

Weighted sum z of the inputs x with weights w and bias b, followed by the activation function g.

σ(z)=11+e−z\sigma(z) = \frac{1}{1 + e^{-z}}

The logistic sigmoid, a common activation function, squashes any number into the range 0 to 1.

ReLU(z)=max⁡(0,z)\mathrm{ReLU}(z) = \max(0, z)

The rectified linear unit, today's default for hidden units, turns negative values into 0 and passes positive values through.

Layers and why they matter

Units are arranged in layers. The input layer holds the raw numbers, one or more hidden layers transform them, and the output layer gives the answer — for a classifier, often a softmax layer that turns scores into probabilities that add up to 1. A single linear unit can only separate classes with a straight line, so it cannot compute XOR ('one or the other, but not both'). A network with one hidden layer of just two units can. In general, the universal approximation theorem says a network with at least one hidden layer can approximate a very broad class of functions as closely as you like, given enough hidden units — though it says nothing about whether training will find the right weights.

Worked example: a neuron that computes AND

Give a sigmoid neuron weights 20 and 20 and a bias of −30, and feed it two inputs that are each 0 or 1. For (1, 1): z = 20 + 20 − 30 = 10 and σ(10) ≈ 0.99995. For (1, 0) or (0, 1): z = −10 and σ(−10) ≈ 0.000045. For (0, 0): z = −30, even closer to 0. With a decision threshold of 0.5, the neuron says 'true' only when both inputs are 1 — it computes logical AND. This is exactly the AND example in the neural-network playground, where you can change the weights and watch the decision boundary move.

The 3D neural-network scene shows the next step up: a 3-4-2 network with a ReLU hidden layer and a softmax output layer. Its weights are fixed illustrative values, not the result of training. With inputs 1.0, 0.5 and −0.2, one hidden unit's weighted sum is −0.34, so ReLU switches it off (it outputs 0), while the other three pass on 0.7, 0.27 and 0.51. The two output scores, 0.237 and −0.186, go through the softmax and become probabilities of about 0.60 and 0.40, which add up to 1.

  1. XOR in two layers: a hidden unit computing OR (weights 20, 20; bias −10) and one computing NAND (weights −20, −20; bias 30) feed an output unit computing AND (weights 20, 20; bias −30).
  2. Input (1, 0): OR ≈ 1, NAND ≈ 1, so the output AND ≈ 1 — true.
  3. Input (1, 1): OR ≈ 1, NAND ≈ 0, so the output AND ≈ 0 — false. That is XOR.

How networks learn their weights

The idea is old and came from biology: the Nobel committee explains that artificial neural networks were originally inspired by the structure of the brain, with nodes standing in for neurons and adjustable connections likened to synapses. McCulloch and Pitts published 'A logical calculus of the ideas immanent in nervous activity' in 1943, and Frank Rosenblatt's paper on the perceptron appeared in 1958. Hand-picking weights works for AND and XOR, but real networks have millions of them. In 1986 Rumelhart, Hinton and Williams described back-propagation: a procedure that repeatedly adjusts every connection weight to reduce the difference between the network's output and the desired output. Their striking finding was that hidden units then come to represent useful features of the task on their own. Back-propagation works out which direction to change each weight; gradient descent takes the step.

Deep learning is the use of many such layers, each learning a representation of the data at a higher level of abstraction. In 2024 the Nobel Prize in Physics was awarded to John Hopfield and Geoffrey Hinton for foundational discoveries and inventions that enable machine learning with artificial neural networks.

Common misconception: Misconception: 'A neural network is a computer copy of a brain.' Correct idea: artificial neural networks are engineered systems loosely inspired by neurons, but they are generally not designed to be realistic models of biological function. Each unit is just a weighted sum and a simple function.
Common misconception: Misconception: 'If a network can represent a function, it will learn it.' Correct idea: the universal approximation theorem guarantees a large enough network can represent the function, not that training will find the right weights or that the result will generalise.
Info: Where this connects: gradient descent explains how the weights are adjusted; computer vision and language models are the two biggest applications; transformers are a particular way of wiring layers together.

Ask ScienceVerse

Still curious about Neural networks? Ask a question, get hints, take a short lesson or try a challenge. The tutor answers only from this concept's approved sources, and says so when it has none.

Ask the tutor about this concept on the full tutor page.

Connections

Prerequisites

Understand these first:

Guided learning path

See everything to learn before this, in order, with your progress:

Leads to

This concept is a building block for:

Related concepts

Explore in 3D

See this concept as an interactive 3D scene:

Try the simulation

Explore this concept interactively, with real adjustable parameters:

Try the experiment

Put this concept into practice with a hands-on activity (each shows its supervision requirement first):

Check your understanding

Take a quick check of two to five questions, with an explanation for every answer:

See the neighbourhood of Neural networks in the Knowledge Galaxy

Sources and methodology

  • Neural networks used for machine learning are engineered systems inspired by the biological brain, but they are generally not designed to be realistic models of biological function. (awaiting scientific review)
  • The Royal Swedish Academy of Sciences explains that machine learning with artificial neural networks was originally inspired by the structure of the brain: the brain's neurons are represented by nodes, and the connections between nodes can be likened to synapses that can be made stronger or weaker. (awaiting scientific review)
  • Warren McCulloch and Walter Pitts published 'A logical calculus of the ideas immanent in nervous activity' in The Bulletin of Mathematical Biophysics in 1943. (awaiting scientific review)
    • A logical calculus of the ideas immanent in nervous activity — Peer-reviewed paper
  • Frank Rosenblatt's paper 'The perceptron: A probabilistic model for information storage and organization in the brain' was published in Psychological Review in 1958. (awaiting scientific review)
    • The perceptron: A probabilistic model for information storage and organization in the brain — Peer-reviewed paper
  • Rumelhart, Hinton and Williams (1986) described back-propagation, a learning procedure that repeatedly adjusts the weights of the connections in a network so as to minimise the difference between the network's actual output and the desired output. (awaiting scientific review)
  • Rumelhart, Hinton and Williams reported that, through back-propagation's weight adjustments, internal 'hidden' units that are neither inputs nor outputs come to represent important features of the task — an ability that distinguished it from earlier, simpler methods such as the perceptron-convergence procedure. (awaiting scientific review)
  • In modern neural networks the default recommendation for hidden units is the rectified linear unit (ReLU), defined by the activation function g(z) = max(0, z). (awaiting scientific review)
  • A linear model cannot represent the XOR function, but a feedforward network with a single hidden layer containing two hidden units can. (awaiting scientific review)
  • The universal approximation theorem says a feedforward network with a linear output layer and at least one hidden layer with a suitable activation function can approximate a very broad class of functions to any desired non-zero error, given enough hidden units — but it does not guarantee that a training algorithm will find those weights. (awaiting scientific review)
  • Deep learning uses computational models composed of multiple processing layers that learn representations of data with multiple levels of abstraction, with the backpropagation algorithm indicating how each layer's internal parameters should change. (awaiting scientific review)
  • The 2024 Nobel Prize in Physics was awarded to John J. Hopfield and Geoffrey Hinton 'for foundational discoveries and inventions that enable machine learning with artificial neural networks'. (awaiting scientific review)
  • Worked calculation (author's own, matching the neural-network playground's AND-gate example): a sigmoid neuron with weights 20 and 20 and bias −30 gives σ(−10) ≈ 0.000045 for inputs (1, 0) and σ(10) ≈ 0.99995 for inputs (1, 1), so with a 0.5 threshold it outputs 'true' only when both inputs are 1. (awaiting scientific review)

Claims marked “awaiting scientific review” cite the sources listed but have not yet been signed off by a scientific reviewer.

Content status: published 1 October 2026.

  • Scientific review: this version has not yet been signed off by a scientific reviewer.
  • The Advanced explanation has not yet been reviewed for age suitability.