Embeddings definition
Embeddings are lists of numbers, called vectors, that represent the meaning of words, sentences, images or other data so that similar items end up with similar vectors. Produced by machine learning models, embeddings let software compare meaning mathematically, which powers semantic search, recommendations, clustering and retrieval-augmented generation.
How do embeddings work?
An embedding model is trained so that items used in similar ways or with similar meanings land near each other in a high-dimensional space, typically with hundreds to a few thousand dimensions. Early word embeddings such as word2vec showed the idea vividly: the vector for "king" minus "man" plus "woman" lands close to "queen". Modern models embed whole sentences and documents, capturing context rather than single words.
Similarity between two embeddings is usually measured with cosine similarity, the angle between their vectors. "I can't log in to my account" and "password reset not working" share few words but produce close vectors, while "log in" in a forestry document lands somewhere else entirely. That property is what makes embeddings the backbone of semantic search. Dot product or Euclidean distance give similar rankings when vectors are normalized.
Types of embeddings
Embeddings can also be fine-tuned. When a general model confuses terms that matter in your domain, such as two similar part numbers or related legal concepts, training it on pairs of related and unrelated texts from your own data can noticeably improve retrieval quality without changing the rest of the system.
- Word embeddings: word2vec, GloVe and fastText, one vector per word.
- Sentence and document embeddings: Sentence-BERT models and hosted embedding APIs from providers such as OpenAI, Cohere, Google and Voyage AI.
- Image and multimodal embeddings: CLIP-style models place images and text in one shared space.
- User and item embeddings: learned by recommendation systems from behavior.
- Graph embeddings: represent nodes in a network, such as customers and products.
What embeddings are used for
Worked example: a legal team embeds every clause in its contract archive. A lawyer pastes a new indemnity clause and instantly sees the twenty most similar clauses from past deals, with notes on how each was negotiated. No keyword query could do this, because similar clauses are worded in many different ways across firms and years.
- Retrieval for RAG and semantic search.
- Recommendations and "more like this" features.
- Clustering documents, tickets or feedback by topic.
- Classification with a light model trained on embeddings.
- Deduplication and anomaly detection.
- Matching resumes to job descriptions, or questions to FAQ answers.
- Input features for downstream machine learning models.
How to choose an embedding model
Start with the public MTEB benchmark to shortlist models, then test on your own data, since rankings vary by domain and language. Check supported languages, maximum input length, vector size, which drives storage cost and search speed, price per token and whether the model can be self-hosted. Some models support shorter vectors with little quality loss, which helps at large scale.
Vectors from different models are not compatible. Switching models means re-embedding the whole collection, so record which model and version produced each vector. Nexzem builds re-embedding into the ingestion pipeline from the start, which turns a future model upgrade into a scheduled job instead of a migration project.