Optimize generative AI systems and model performance
Advanced Fine-Tuning and Model Customization
CoreChoose supervised, preference, or reinforcement fine-tuning, prepare and generate training data, read training metrics to catch overfitting, and move a fine-tuned model from evaluation to production deployment.
Aligned to the live AI-300 guide, which publishes no skills-measured date; guide and product behavior verified October 10, 2026.
Why this matters
Fine-tuning changes model behavior permanently for a deployment, at real training and hosting cost. Picking the method that matches the data you have, validating it against the base model, and deploying it deliberately keeps customization from becoming an expensive regression.
Must Know
- Supervised fine-tuning (SFT) trains on prompts with desired responses; direct preference optimization (DPO) trains on preferred and rejected responses to the same input; reinforcement fine-tuning (RFT) trains on prompts with a grader that scores generated responses.
- SFT data is JSONL with one chat example of system, user, and assistant messages per line. DPO data uses input, preferred_output, and non_preferred_output fields.
- Provide a validation file so training reports validation loss and token accuracy next to training loss. Validation loss that rises while training loss falls indicates overfitting.
- For DPO, token accuracy does not measure whether outputs match preferences, so evaluate preference alignment separately.
- Synthetic data generation in Foundry (preview) can produce simple question-and-answer or tool-use training data from reference files, which should be reviewed and filtered before training.
- Developer deployments have no hourly hosting fee and are removed after 24 hours, which suits evaluating a fine-tuned model before committing to a production deployment type.
- SFT and DPO cost is training tokens multiplied by epochs, with no charge for queue time or failed jobs. RFT is billed by training time plus any model-grader tokens, and a canceled RFT run is charged for the training it completed.
Compare and Distinguish
- Prompt engineering versus RAG versus fine-tuning: instructions versus fresh knowledge at query time versus learned behavior and format.
- SFT versus DPO: imitating demonstrations versus learning which of two responses is preferred.
- DPO versus RFT: fixed preference pairs versus graded exploration, often for reasoning tasks.
- Developer deployment versus Standard or Global deployment: short-lived evaluation versus durable production serving.
Scenario examples
- Scenario: A model must always answer in a strict JSON structure learned from thousands of examples. Think: supervised fine-tuning.
- Scenario: Subject experts have pairs of better and worse answers for the same prompts. Think: direct preference optimization.
- Scenario: Validation loss climbs after the second epoch while training loss keeps dropping. Think: overfitting, so reduce epochs or add diverse data.
Exam traps
- Fine-tuning is not the right tool for frequently changing knowledge; use RAG for fresh facts.
- A Developer deployment is not a production option because it expires after 24 hours.
- Synthetic data used without review can amplify errors and bias.
- Falling training loss alone does not prove the fine-tuned model generalizes.
Key takeaways
- Match the fine-tuning method to the data: demonstrations, preferences, or a grader.
- Watch validation metrics and compare against the base model before promotion.
- Evaluate on a Developer deployment, then deploy to a production type.
How it works
- Fine-tuning adapts model weights, commonly with low-rank adaptation, using the training file and reports metrics per step.
- The resulting model is deployed like other models, with a deployment type chosen for cost, residency, and throughput.
Objects and administrative surfaces
- Foundry fine-tuning jobs, training and validation files, and result metrics.
- Synthetic data generation in the Foundry portal.
- Deployments of fine-tuned models with Developer, Standard, or Global deployment types.
When to use it
- Use fine-tuning for consistent style, format, or domain behavior that prompting cannot reliably achieve.
- Use RFT when correctness can be graded automatically, such as structured reasoning tasks.
Security and governance implications
- Remove sensitive data from training files and generated synthetic data.
- Require evaluation and safety checks before promoting a fine-tuned model to production.
Troubleshooting signals
- A fine-tuned model that underperforms the base model on general prompts may be overfit to narrow training data.
- Validation metrics missing from results usually mean no validation file was supplied.
More detail
- Design and implement SFT, DPO, and RFT fine-tuning.
- Create and manage synthetic training data.
- Monitor and optimize fine-tuned model performance.
- Manage a fine-tuned model from development through production deployment.
Ready for the quiz?
- What data does DPO need?
- What pattern in loss curves shows overfitting?
- Why is a Developer deployment unsuitable for production?
- When is RAG a better choice than fine-tuning?
Related objectives
- D5.2.S1 — Design and implement advanced fine-tuning methods
- D5.2.S2 — Create and manage synthetic data for fine-tuning
- D5.2.S3 — Monitor and optimize fine-tuned model performance
- D5.2.S4 — Manage a fine-tuned model from development through production deployment