Transformer Model definition
A transformer model is a neural network architecture that processes sequences, such as text, using a mechanism called self-attention, which lets every element weigh its relationship to every other element at once. Introduced by Google researchers in 2017, transformers underpin large language models and are widely used in translation, speech recognition and computer vision.
How does a transformer work?
Input text is split into tokens, each turned into an embedding, and positional information is added so the model knows word order. The sequence then passes through a stack of identical layers. Each layer has a self-attention block, where every token gathers information from the others, followed by a feed-forward network that processes each token individually. Residual connections and normalization keep training stable as layers are stacked deep.
In self-attention, each token produces a query, a key and a value. Comparing one token's query with every key decides how much attention it pays to each other token, and the result is a weighted mix of their values. In "the animal didn't cross the street because it was too tired", attention lets "it" link strongly to "animal". Multi-head attention runs several of these comparisons in parallel to capture different relationships.
Types of transformer architectures
Most teams never build a transformer from scratch. They choose a pretrained model of the right type and size from a provider or a hub such as Hugging Face, then adapt it through prompting, fine-tuning or a small task-specific layer added on top.
- Encoder-only (BERT, RoBERTa): read the whole input at once, used for classification, search and embeddings.
- Decoder-only (the GPT, Llama and Claude families): generate text one token at a time.
- Encoder-decoder (T5 and the original 2017 translation model): map one sequence to another.
- Vision transformers (ViT): treat image patches as tokens.
- Mixture-of-experts transformers: route each token to a subset of specialist layers to save compute.
Transformers vs RNNs and CNNs
Recurrent neural networks read sequences one step at a time, which made them slow to train and prone to forgetting information from far back in a sequence. Transformers process all positions in parallel on GPUs and connect distant words directly, which let researchers train far larger models on far more data. That scalability, more than any single trick, explains their dominance.
The main weakness is cost on long inputs, since standard attention grows with the square of sequence length. Efficient attention methods, caching and alternative designs such as state space models aim to reduce this. Convolutional networks remain competitive for many vision tasks, especially on small devices, and hybrids that combine convolution and attention are common.
Where transformers are used
Transformers began in machine translation and now appear across AI. The original paper, "Attention Is All You Need", described a translation model, but the same architecture proved general enough to handle almost any data that can be expressed as a sequence of tokens, from words to image patches to amino acids.
- Chat assistants, coding tools and other LLM applications.
- Search ranking and semantic embeddings.
- Speech recognition, such as OpenAI's Whisper.
- Image classification and detection with vision transformers.
- Image and video generation, where transformers encode prompts or act as the denoising network.
- Biology, where attention-based models help predict protein structures.
- Recommendation systems that model sequences of user actions.