Fine-Tuning vs. Prompting
Fine-tuning adapts a pre-trained model to your specific domain or task. Understanding when it is worth the investment — and what alternatives exist — is essential for production AI decisions.
What fine-tuning is
Fine-tuning continues training a pre-trained model on a smaller, domain-specific dataset. The model updates its weights to better reflect the patterns in your data, improving performance on your specific task distribution. Unlike prompting, fine-tuned behaviour is baked into the model weights rather than provided at inference time.
When fine-tuning makes sense
- You have a narrow, well-defined task where format and style consistency are critical
- You have enough high-quality training examples (typically hundreds to thousands)
- You need the model to reliably follow a specific output schema that prompting alone does not achieve
- Latency or cost constraints prevent including extensive examples in every prompt
- You need the model to internalise proprietary terminology or domain-specific knowledge
When prompting (or RAG) is better
Fine-tuning is not a solution for hallucinations — the model cannot reliably internalise factual knowledge through fine-tuning and retrieve it accurately. For tasks requiring accurate recall of specific facts or documents, retrieval-augmented generation (RAG) is almost always superior.
Practical costs
Fine-tuning on proprietary APIs (OpenAI, Anthropic) costs money and requires careful data preparation. The operational overhead — maintaining training datasets, tracking model versions, managing deployment — should be weighed against the performance gains versus a well-prompted base model.
Fine-tuning should come after prompting has been optimised, not instead of it. Many problems attributed to needing fine-tuning can be solved with better prompts and few-shot examples.