Skip to content

What is Tokens (in LLMs)?

Generative AI & LLMs, explained by the engineers who build it. Definition, how it works, use cases and common questions.

in LLMs definition

Tokens are the small units of text that a large language model reads and writes, typically whole words, parts of words, punctuation marks or spaces. Models convert text into tokens before processing it, and LLM pricing, speed and context window limits are all measured in tokens, so token counts directly affect cost and performance.

How does tokenization work?

Before a model sees text, a tokenizer splits it into pieces from a fixed vocabulary, often tens of thousands to a few hundred thousand entries. Most modern tokenizers use subword methods such as byte-pair encoding, WordPiece or SentencePiece. Common words become a single token, while rare words are split into fragments, so "unbelievable" might become "un", "believ" and "able". Each token maps to an ID, and each ID maps to an embedding inside the model.

A widely quoted rule of thumb for English is about four characters, or roughly three quarters of a word, per token, but this varies by tokenizer. Many other languages, including Hindi and other non-Latin scripts, often need more tokens for the same meaning, which raises cost. Tokenization also explains some odd failures, such as models miscounting letters in a word they see as two fragments.

Why tokens matter

Tokens are also the unit of thinking for reasoning models, which generate internal reasoning tokens before the visible answer. Those tokens are usually billed as output, so a short reply to a hard question can cost more than its length suggests, and budgets should allow for it.

  • Cost: providers bill per input and output token, with output usually priced higher.
  • Context window: the maximum tokens a model can handle per request, input and output combined.
  • Speed: output tokens are generated one after another, so long answers take longer.
  • Rate limits: APIs cap tokens per minute as well as requests.
  • Caching: some providers discount repeated prompt prefixes, rewarding stable prompt design.

Example: estimating token cost

A support assistant sends a 1,200-token system prompt, 1,500 tokens of retrieved articles and a 100-token question, then receives a 300-token answer. That is 2,800 input and 300 output tokens per request. At 10,000 requests a day, the feature processes 28 million input and 3 million output tokens daily. Multiplying by the provider's per-million-token prices gives the daily bill, and shows that retrieved context is the biggest lever.

How to reduce token usage

  • Trim system prompts and remove instructions the model follows anyway.
  • Retrieve fewer, better chunks by improving search and reranking.
  • Summarize long conversation history instead of resending it in full.
  • Use prompt caching for long, unchanging prompt prefixes.
  • Set a maximum output length and ask for concise formats.
  • Route simple steps to smaller, cheaper models.
  • Request structured output instead of verbose prose when code will parse the result.
  • Strip boilerplate such as HTML markup and repeated headers before sending documents.

How to count tokens

Each model family uses its own tokenizer, so the same text yields different counts on different models. Libraries such as OpenAI's tiktoken and Hugging Face tokenizers count tokens locally, and several providers offer token counting endpoints or return exact usage with every API response. Logging those usage figures per feature and per customer is the simplest way to understand where AI spending actually goes and to catch runaway costs early.

in LLMs: common questions

Something else on your mind? Ask a consultant and get a reply within one business day.

How many words are 1,000 tokens?

For typical English prose, roughly 750 words, based on the common estimate of about three quarters of a word per token. Code, numbers, unusual vocabulary and many non-English languages use more tokens per word, so measure with the actual tokenizer when costs matter.

Why are output tokens more expensive than input tokens?

Input tokens can be processed in parallel in one pass, while output tokens are generated one at a time, each requiring a full pass through the model. Generation therefore uses more compute per token, which providers reflect in higher output prices.

How can a business control LLM token costs?

Track token usage per feature, set budgets and alerts, shorten prompts, cache stable context, cap response length and use smaller models where quality allows. Nexzem adds per-feature token dashboards to LLM applications so cost trends are visible before the monthly invoice arrives.

Keep exploring the generative ai & llms glossary

Need in LLMs in your product?

A solutions consultant replies within one business day with next steps, a rough estimate and a suggested team.