Landmark 6 stops to get here · leads to 5

Fine-Tuning

The process of further training a pre-trained model on a specific dataset to adapt it for a particular task or domain.

Your route here

6 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  5. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  6. Pre-training ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  7. Fine-Tuning · you are here ✓ understood

Picture it

  1. 01 Pre-trained base model General knowledge from broad data
  2. 02 Task-specific dataset Smaller, focused examples
  3. 03 Continue training Usually a lower learning rate
  4. 04 Evaluate on validation Check gains, watch for forgetting
  5. 05 Specialized model
Notice that training starts from existing weights, not scratch, so far less data and compute are needed.

Fine-tuning leverages a model’s pre-trained knowledge and adapts it to specific tasks, domains, or behaviors. This is far more efficient than training from scratch.

Process

  1. Start with a pre-trained base model
  2. Prepare task-specific training data
  3. Continue training with a lower learning rate
  4. Evaluate on validation data
  5. Deploy the fine-tuned model

Types

  • Full Fine-Tuning: Update all model parameters
  • Parameter-Efficient Fine-Tuning (PEFT): Update only a subset (LoRA, adapters)
  • Instruction Fine-Tuning: Train on instruction-following data
  • RLHF: Reinforce learning from human feedback

Benefits

  • Faster than training from scratch
  • Requires less data
  • Improves task-specific performance
  • Enables domain adaptation

Where it sits

Explore nearby

In the research

All papers →

7 papers that build on Fine-Tuning ; showing 5, canon first.

Canon · 2018 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding It established “pre-train once, fine-tune everywhere”, and still powers much of search and classification. Canon · 2021 LoRA: Low-Rank Adaptation of Large Language Models It made adapting big models cheap enough for everyone, and is why fine-tunes are shared as small adapter files. Canon · 2022 Training Language Models to Follow Instructions with Human Feedback This recipe, RLHF, is what turned raw language models into assistants like ChatGPT. Frontier · Sep 2025 · 114 citations Fast-dLLM v2: Efficient Block-Diffusion LLM A cheap path to faster inference from models you already have. Frontier · Aug 2025 · 112 citations On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification A tiny, theory-backed change to the most common training step in the stack.