Vision & Multimodal Oct 2020

An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Alexey Dosovitskiy et al. · ICLR 2021

arXiv:2010.11929

In short

Cut an image into 16×16 patches, treat each patch like a word, and feed the sequence to a standard transformer. Pre-trained on enough data, this Vision Transformer matches or beats the best convolutional networks for less training compute.

Why it matters

It brought vision and language onto one architecture, which made today’s multimodal models possible.

Read first

The 3 Field Guide ideas this paper leans on.

Starting from scratch? The full route 14 ideas · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  8. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  9. Transformer · read first ✓ understood

    A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.

  10. Supervised Learning ✓ understood

    Learning from examples paired with the correct answer, so a model can predict answers for new inputs it hasn't seen.

  11. Classification ✓ understood

    A supervised learning task where the model assigns each input to one of a fixed set of categories, such as spam or not spam.

  12. Computer Vision ✓ understood

    The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.

  13. Image Classification · read first ✓ understood

    Assigning a single label or category to an entire image, a fundamental computer vision task.

  14. Vision Transformer · read first ✓ understood

    Applying the transformer architecture to computer vision by treating image patches as tokens, achieving state-of-the-art results.

Nearby papers

Summary in our own words; read the paper for the details. ← All papers