Vision & Multimodal Feb 2021

Learning Transferable Visual Models From Natural Language Supervision

Alec Radford et al. · ICML 2021

arXiv:2103.00020

In short

CLIP trains an image encoder and a text encoder together on 400 million image–caption pairs, learning to match each image with its caption. The shared embedding space lets it classify images into categories it was never trained on, just from their names.

Why it matters

CLIP is the bridge between words and pictures inside text-to-image models and many vision-language systems.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 17 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding · read first ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  8. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  9. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  10. Contrastive Learning · read first ✓ understood

    A self-supervised learning approach that learns representations by contrasting similar and dissimilar examples.

  11. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  12. Pre-training ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  13. Transfer Learning ✓ understood

    Leveraging knowledge learned from one task/domain to improve performance on a related task with less data.

  14. Zero-Shot Learning · read first ✓ understood

    A model's ability to perform tasks it wasn't explicitly trained on, using only instructions or descriptions.

  15. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  16. Transformer ✓ understood

    A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.

  17. CLIP · read first ✓ understood

    Contrastive Language-Image Pre-training - a model jointly trained on images and text, enabling zero-shot image classification.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers