Language & LLMs Jan 2013

Efficient Estimation of Word Representations in Vector Space

Tomas Mikolov et al. · ICLR 2013 workshop

arXiv:1301.3781

In short

Two very simple models learn a vector for every word by predicting words from their neighbours, fast enough to train on billions of words in under a day. Words used in similar contexts end up close together, and directions in the space capture relationships.

Why it matters

It made embeddings, meaning as geometry, the starting point of modern NLP.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 9 ideas · basics first
  1. Natural Language Processing · read first ✓ understood

    The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.

  2. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  3. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  4. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  5. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  6. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  7. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  8. Embedding · read first ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  9. Word2Vec · read first ✓ understood

    A technique for learning word embeddings that capture semantic relationships (Skip-gram and CBOW models).

Nearby papers

Summary in our own words; read the paper for the details. ← All papers