Fine-Tuning definition
Fine-tuning is the process of taking a pretrained machine learning model, such as a large language model, and training it further on a smaller, task-specific dataset so it performs better on that task. It adjusts the model's weights to teach a consistent format, tone, vocabulary or skill that prompting alone does not reliably achieve.
How does fine-tuning work?
You prepare a dataset of examples showing the input and the ideal output, often a few hundred to a few thousand for a focused task. Training then runs for a small number of passes with a low learning rate, nudging the pretrained weights toward your examples without erasing what the model already knows. A held-out set of examples is used to check that quality actually improved over the base model with a good prompt.
Full fine-tuning updates every weight, which needs significant GPU memory for large models. Parameter-efficient fine-tuning instead trains a small set of added weights. LoRA, the most common method, adds low-rank matrices to selected layers, and QLoRA does the same on a quantized model to fit on smaller GPUs. Hosted services from OpenAI, Google's Gemini Enterprise Agent Platform and Amazon Bedrock offer fine-tuning without managing hardware.
Types of fine-tuning
- Supervised fine-tuning (SFT): training on input and ideal output pairs.
- Instruction tuning: SFT on many varied tasks so a model follows instructions in general.
- Preference tuning: RLHF or DPO, using ranked answers to steer style and safety.
- Parameter-efficient methods: LoRA, QLoRA and adapters.
- Continued pretraining: more training on raw domain text, such as legal or medical documents.
- Distillation: training a smaller model to copy a larger model's outputs.
- Domain adaptation: tuning speech or vision models on specialized recordings or images.
When to fine-tune, and when not to
Fine-tune when you need consistent behavior that prompts cannot hold: a strict output format, a specific brand voice, a specialized classification or extraction task, or the same quality from a smaller, cheaper and faster model. Fine-tuning also shortens prompts, since the instructions and examples no longer need to be sent with every request.
Do not fine-tune to add knowledge that changes, such as prices, policies or product documentation. The model may blend facts unpredictably, and every update would need retraining. Retrieval-augmented generation handles changing knowledge better and can cite sources. Always test strong prompting and retrieval first, because they are cheaper to build and maintain.
Example: a document extraction model
An insurer extracts a dozen fields from claim forms. A large model with a detailed prompt works but is slow and costly at volume. The team collects a few thousand human-verified extractions and fine-tunes a smaller open model with LoRA. The tuned model matches the large model on the evaluation set for this narrow task, runs on the insurer's own servers and answers faster. Low-confidence extractions still go to a human reviewer, and the team retrains when claim forms change.
Risks and costs
Fine-tuning quality depends on data quality: inconsistent or wrong examples teach inconsistent or wrong behavior. Over-training can cause the model to lose general abilities, known as catastrophic forgetting, or to memorize sensitive training data. Tuned models also need maintenance, since moving to a newer base model means retraining. Nexzem treats fine-tuning datasets as versioned assets and compares every tuned model against a prompted baseline before deployment.