An embedding turns something discrete, like a word, a product or a whole paragraph, into a list of numbers, typically a few hundred to a few thousand long. Nobody picks the numbers by hand. They’re learned, so that items used in similar ways end up with similar vectors: “cat” lands near “kitten,” “invoice” near “receipt.”
Once meaning is geometry, comparing meaning becomes arithmetic. The angle or distance between two vectors gives a semantic similarity score.
Where they come from
Early language systems gave each word a one-hot vector: as long as the vocabulary, all zeros except a single 1. Every word was equally far from every other, so nothing learned about “cat” carried over to “kitten.” In 2003, Bengio and colleagues trained a language model that learned a dense vector for each word alongside everything else. A sentence it had never seen could still get a sensible probability if its words had vectors close to those in sentences it had seen. Word2vec, in 2013, made training word vectors cheap enough to run over more than a billion words in under a day.
Inside a large language model, the first layer is an embedding table: each token looks up its vector, position information is added, and the transformer layers take it from there.
What they’re used for
- Semantic search: embed your documents and the query, then return the nearest documents. That’s the retrieval step in RAG, and vector databases exist to make nearest-neighbour lookups fast across millions of vectors.
- Clustering and deduplication: near-identical vectors flag near-identical content.
- Recommendations: put users and items in the same space and suggest what’s nearby.
Sentence-level embedding models made this practical at scale. Sentence-BERT (2019) showed that comparing precomputed sentence vectors could find the most similar pair among 10,000 sentences in about 5 seconds, a job that took around 65 hours when every pair went through BERT.
The catch
“Similar” means similar in how the training text used things, which isn’t always what you need. “Hot” and “cold” appear in the same kinds of sentences, so they can sit close together despite meaning opposite things. Embeddings also absorb the associations and biases of their training text. And vectors from different models aren’t comparable: switch embedding models and you re-embed everything.