LLM definition
A large language model (LLM) is a deep learning model, usually based on the transformer architecture, trained on vast amounts of text to predict the next token in a sequence. That simple objective, applied at large scale, gives LLMs the ability to answer questions, summarize, translate, write code and follow instructions written in natural language.
How do large language models work?
Text is split into tokens, and each token becomes an embedding vector. The vectors pass through dozens of transformer layers, where attention lets every token draw information from the others, so a word is understood in context. The final layer outputs a probability for every token in the vocabulary. The model samples one, appends it to the text and repeats, generating a response one token at a time.
Training happens in stages. Pretraining on large text collections teaches language, facts and patterns. Supervised fine-tuning on example conversations teaches the model to follow instructions. Preference tuning, using methods such as RLHF or DPO, nudges it toward answers people rate as helpful and safe. Most of the cost sits in pretraining, which is why only a handful of organizations build frontier models from scratch.
Examples of large language models
- Hosted proprietary families: OpenAI GPT, Anthropic Claude, Google Gemini.
- Open-weight families: DeepSeek, Alibaba Qwen, Google Gemma, OpenAI gpt-oss, Mistral and Meta Llama, which can be self-hosted.
- Code-specialized models used in developer tools and coding agents.
- Small language models built for low-cost, on-device or high-volume tasks.
- Reasoning-focused variants that spend extra computation thinking before they answer.
- Domain models tuned for areas such as law, medicine or finance.
What LLMs are good and bad at
LLMs excel at transforming language: drafting, summarizing, rewriting, translating, extracting fields and classifying text. Given relevant context in the prompt, they reason over it well and can write and explain code. They are weaker at things outside the text they see. Their knowledge stops at a training cutoff, they can state false information confidently, and they are unreliable at exact arithmetic or lookups unless connected to tools.
They also have practical limits: a finite context window, latency that grows with output length, and costs billed per token. Good LLM applications are designed around these traits, giving the model the facts it needs through retrieval, letting it call calculators or APIs for precise work, and checking outputs before they reach users.
How businesses use LLMs
Most products combine a few standard patterns rather than training a new model. The first decision is usually whether to call a hosted API or self-host an open-weight model. APIs give the strongest models with no infrastructure. Self-hosting gives full control over data location and predictable cost at high volume, at the price of running GPU infrastructure.
- Prompting: instructions and examples in the prompt shape behavior.
- Retrieval-augmented generation: company documents supplied at query time.
- Function calling: the model triggers APIs, searches or database queries.
- Fine-tuning: adjusting a model for a specific format, tone or task.
How to choose an LLM
Public leaderboards are a starting point, not an answer. Build a test set of 50 to 200 real examples from your use case and compare candidate models on quality, latency, cost per request, context length, license terms and data handling. Nexzem runs this comparison early in LLM projects, and often finds a smaller, cheaper model meets the bar for most traffic.