Fine-tuning continues the training of a model on your own examples so its weights change. Afterward, the model is more likely to follow your format, vocabulary, or task without you repeating long instructions every time.
What actually changes
A prompt changes the input for one request. Retrieval adds documents for one request. Fine-tuning updates parameters, so the new behavior shows up on later requests. You still need a clean evaluation set. A fine-tune can memorize a bad pattern just as easily as a good one.
Common forms
| Form | What you train on | Typical goal |
|---|---|---|
| Supervised fine-tuning | Pairs of input and the ideal output | A fixed format, such as extracting fields from a letter |
| Preference tuning | A better answer and a worse answer for the same prompt | Tone, safety, or which style people prefer |
| Parameter-efficient tuning (LoRA and similar) | The same pairs, but only a small set of added weights | A cheaper adaptation you can swap in and out |
When to fine-tune, and when not to
- Fine-tune when the task is stable and you have hundreds of good input-output examples.
- Do not fine-tune to “teach” a fact that will change next month. Put that fact in a document and use RAG.
- Try a stronger prompt and a few examples first. Fine-tune when that still fails in a consistent way.
A practical example
A team classifies inbound complaints into a fixed set of reason codes and writes a one-line summary in house style. They collect 800 reviewed examples. A small fine-tune learns the codes and the summary shape. The complaint text itself still comes from the customer; the model is not the system of record.
What to remember
- Fine-tuning changes weights. Prompting and RAG do not.
- It is for skills and formats, not for a database of changing facts.
- Quality of examples matters more than quantity of unreviewed text.