Fine-tuning takes a model that has already been pretrained on massive general-purpose data. Training then continues on a much smaller, curated dataset to specialize the model's behavior.
Fine-tuning starts from learned weights and updates them with a much smaller, task-focused dataset. It needs far less data and compute than pretraining.
The practical differences are:
- Scale: trillions of tokens versus thousands of curated examples.
- Objective: usually the same next-token prediction loss, but applied to task-specific input-output pairs.
- Learning rate: much lower than pretraining, so behavior changes without overwriting broad capabilities.
- Goal: pretraining builds raw capability; fine-tuning shapes behavior such as style, output format, domain vocabulary, and task reliability.
Almost no application team pretrains from scratch. In practice, customization means prompting, retrieval-augmented generation (RAG), or fine-tuning, and fine-tuning itself is usually parameter-efficient (LoRA-style) rather than updating every weight.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓