Neural Network definition
A neural network is a machine learning model made of connected layers of simple computing units, called neurons, that pass numbers to one another through weighted connections. By adjusting those weights during training, the network learns to map inputs such as pixels, words or sensor readings to outputs such as labels, scores or predictions.
How does a neural network work?
Data enters through an input layer, passes through one or more hidden layers and leaves through an output layer. Each neuron computes a weighted sum of its inputs, adds a bias and applies an activation function such as ReLU or sigmoid, which lets the network model nonlinear relationships. In a classic network that reads handwritten digits, the input layer receives 784 pixel values from a 28 by 28 image and the output layer produces ten scores, one per digit.
Size is described by layers and parameters. A small network for tabular data might have a few thousand weights, while large language models have billions. More parameters give more capacity to learn, but they also need more data, memory and compute, so the right size is usually the smallest network that reaches the target accuracy on validation data.
How neural networks learn
Training is a loop that runs thousands or millions of times. The network starts with random weights and gradually improves as it sees labeled examples, a process known as gradient descent. Each pass over the full dataset is called an epoch, and data is fed in small batches to fit in GPU memory.
- Forward pass: inputs flow through the layers to produce a prediction.
- Loss: a function such as cross-entropy measures the error against the correct answer.
- Backpropagation: the chain rule computes how much each weight contributed to that error.
- Update: an optimizer such as stochastic gradient descent or Adam adjusts the weights.
- Repeat: the cycle runs over many batches until validation accuracy stops improving.
Types of neural networks
Feedforward networks, also called multilayer perceptrons, handle fixed-size inputs such as tabular features. Convolutional networks share weights across an image to detect local patterns wherever they appear. Recurrent networks process sequences step by step, while transformers use attention to look at a whole sequence at once and now power most language models. Graph neural networks operate on connected data such as social networks, road maps or molecules.
Strengths and limitations
Neural networks can approximate very complex functions and tend to improve as data grows, which makes them the default for vision, speech and language. They also have weak points. They need large labeled datasets unless you start from a pretrained model, they can memorize training data instead of generalizing, and their internal reasoning is hard to inspect.
Techniques such as dropout, weight decay and early stopping reduce overfitting, and explainability tools like SHAP or saliency maps show which inputs drove a prediction. None of these remove the need for careful evaluation on data that looks like real production traffic, including the edge cases and noisy inputs that clean benchmark datasets rarely contain.
Neural network examples
- Image recognition: tagging products, reading license plates and spotting defects.
- Speech: transcribing calls and powering voice assistants.
- Language: translation, summarization, search and chatbots.
- Forecasting: energy demand, sales and sensor readings.
- Recommendations: ranking products, videos or articles for each user.