Landmark 2 stops to get here · leads to 2
Self-Supervised Learning
Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).
Your route here
2 stops · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Unsupervised Learning ✓ understood
Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.
- Self-Supervised Learning · you are here ✓ understood
Picture it
- 01 Unlabeled data e.g. raw text scraped at scale
- 02 Hide part of it "The cat [MASK] on the mat"
- 03 Predict the hidden part The data itself supplies the answer: "sat"
- 04 Learned representations Reused for downstream tasks via fine-tuning
Where it sits
Before this
Unsupervised Learning Self-Supervised Learning
Explore nearby
Training Pre-training Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning. Training Contrastive Learning A self-supervised learning approach that learns representations by contrasting similar and dissimilar examples. Language & LLMs Masked Language Modeling A pre-training objective where random tokens are masked and the model learns to predict them from context. Neural Networks Representation Learning Learning useful features or representations of data automatically, rather than hand-crafting them. Foundations Semi-Supervised Learning Learning from a combination of labeled and unlabeled data, leveraging abundant unlabeled data to improve performance.
In the research
All papers →A paper that builds on Self-Supervised Learning .