Landmark 6 stops to get here · leads to 13

Embedding

A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

Your route here

6 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding · you are here ✓ understood

Picture it

ANIMALScatdogwolfkittenVEHICLEScartruckbusbikeCOLORSredbluegreenILLUSTRATIVE · SIMILAR MEANINGS SIT CLOSE TOGETHER
Notice how words with similar meanings land near each other, so distance in the space tracks semantic similarity.

An embedding turns something discrete, like a word, a product or a whole paragraph, into a list of numbers, typically a few hundred to a few thousand long. Nobody picks the numbers by hand. They’re learned, so that items used in similar ways end up with similar vectors: “cat” lands near “kitten,” “invoice” near “receipt.”

Once meaning is geometry, comparing meaning becomes arithmetic. The angle or distance between two vectors gives a semantic similarity score.

Where they come from

Early language systems gave each word a one-hot vector: as long as the vocabulary, all zeros except a single 1. Every word was equally far from every other, so nothing learned about “cat” carried over to “kitten.” In 2003, Bengio and colleagues trained a language model that learned a dense vector for each word alongside everything else. A sentence it had never seen could still get a sensible probability if its words had vectors close to those in sentences it had seen. Word2vec, in 2013, made training word vectors cheap enough to run over more than a billion words in under a day.

Inside a large language model, the first layer is an embedding table: each token looks up its vector, position information is added, and the transformer layers take it from there.

What they’re used for

  • Semantic search: embed your documents and the query, then return the nearest documents. That’s the retrieval step in RAG, and vector databases exist to make nearest-neighbour lookups fast across millions of vectors.
  • Clustering and deduplication: near-identical vectors flag near-identical content.
  • Recommendations: put users and items in the same space and suggest what’s nearby.

Sentence-level embedding models made this practical at scale. Sentence-BERT (2019) showed that comparing precomputed sentence vectors could find the most similar pair among 10,000 sentences in about 5 seconds, a job that took around 65 hours when every pair went through BERT.

The catch

“Similar” means similar in how the training text used things, which isn’t always what you need. “Hot” and “cold” appear in the same kinds of sentences, so they can sit close together despite meaning opposite things. Embeddings also absorb the associations and biases of their training text. And vectors from different models aren’t comparable: switch embedding models and you re-embed everything.

Where it sits

Explore nearby

In the research

All papers →

5 papers that build on Embedding .

Canon · 2013 Efficient Estimation of Word Representations in Vector Space It made embeddings, meaning as geometry, the starting point of modern NLP. Canon · 2020 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks It named and defined the pattern behind most production LLM apps that answer from your own documents. Canon · 2021 Learning Transferable Visual Models From Natural Language Supervision CLIP is the bridge between words and pictures inside text-to-image models and many vision-language systems. Frontier · Jan 2026 · 167 citations Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking Good multimodal embeddings are what make RAG work over screenshots, PDFs and video. Frontier · Jan 2026 · 63 citations Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models It proposes memory lookup as a new scaling axis for LLMs, separate from compute.

Sources

  1. Bengio, Ducharme, Vincent and Jauvin, "A Neural Probabilistic Language Model" . Journal of Machine Learning Research, 2003
  2. Mikolov et al., "Efficient Estimation of Word Representations in Vector Space" . 2013 (word2vec)
  3. Reimers and Gurevych, "Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks" . 2019