Fine-Tuning in Microsoft Foundry: What LoRA, SFT, DPO, and RFT Actually Do to Your Model's Weights
Your prompt is 2,400 tokens long. It has a system message with eleven bullet-pointed rules, four few-shot examples, and a disclaimer about edge cases you added after the third production incident. It works — most of the time. But it's expensive, it's fragile to prompt drift, and every new edge case means another paragraph bolted onto an already-brittle instruction block.
At some point, the right engineering move stops being "write a better prompt" and starts being "change the weights." That's fine-tuning. And in Microsoft Foundry, fine-tuning isn't a side feature bolted onto Azure OpenAI — it's a first-class, multi-technique customization pipeline that spans Supervised Fine-Tuning (SFT), Direct Preference Optimization (DPO), and Reinforcement Fine-Tuning (RFT), built on Low-Rank Adaptation (LoRA) under the hood, with a job lifecycle, checkpoint system, and deployment model that most developers never look at closely enough to use correctly.
This article is that closer look.
Why This Matters
Prompt engineering and fine-tuning solve overlapping but distinct problems, and conflating them is one of the most expensive mistakes a team can make in production GenAI systems. Prompt engineering is runtime configuration — it costs tokens on every single call, it competes for context window space, and it's only as reliable as the model's ability to follow instructions buried in a wall of text. Fine-tuning is a one-time training cost that changes what the model does by default, which means:
- Shorter prompts per call (lower latency, lower token cost at scale)
- More reliable adherence to format, tone, and domain conventions
- The ability to teach behaviors that are hard to specify declaratively (style, disambiguation judgment calls, routing decisions in multi-agent systems)
- Consistent behavior that doesn't degrade when someone edits the system prompt six months later
The catch is that fine-tuning is also easy to misuse. It's not a substitute for RAG when the problem is "the model doesn't know this fact." It's not a safety net for bad base-model selection. And if your dataset is garbage, LoRA will faithfully learn the garbage. Understanding what's happening mechanically — at the matrix-math level, not just the portal-button level — is what separates teams that use fine-tuning well from teams that burn a training budget on a model that performs worse than the base model they started with.
Table of Contents
- Core Concepts: SFT, DPO, and RFT
- What LoRA Actually Does to the Model
- Microsoft Foundry's Fine-Tuning Architecture
- Training Tiers: Standard, Global, and Developer
- Preparing a Dataset That Won't Waste Your Budget
- Running a Fine-Tuning Job: Python SDK and REST
- Hyperparameters That Actually Move the Needle
- Checkpoints, Pausing, and Continuous Fine-Tuning
- Deployment: PTU, Standard, and Automatic Deployment
- A Real-World Developer Scenario: Structured Extraction at Scale
- Production Considerations
- Security and Governance
- Cost Considerations
- Common Mistakes and Pitfalls
- Fine-Tuning vs. Alternatives
- Practical Recommendations
- Conclusion
- References
1. Core Concepts: SFT, DPO, and RFT
Microsoft Foundry exposes three distinct customization techniques, and picking the wrong one for your problem is the single most common reason fine-tuning projects disappoint.
Supervised Fine-Tuning (SFT) is the default. You provide input/output pairs — a prompt and the exact response you want — and the model is trained to reproduce that mapping. This is straightforward imitation learning: no reward function, no comparative judgment, just "given this input, produce this output." SFT is the right starting point for the overwhelming majority of use cases: domain specialization, task-specific formatting, tone and style adaptation, and teaching smaller models to imitate the behavior of a larger one (distillation).
Direct Preference Optimization (DPO) trains on pairs of responses to the same prompt — one preferred, one rejected — without requiring a separate reward model. This matters because some qualities are much easier to compare than to specify. It's hard to write a single "ideal" sarcastic response to a factual question, but it's trivial to say "this response is more sarcastic than that one." DPO is the right tool when your quality signal is comparative and subjective: tone, helpfulness, safety posture, brand voice.
Reinforcement Fine-Tuning (RFT) uses reward signals from grader models to optimize for objectives that are hard to express as either a fixed target output or a pairwise preference. RFT is suited to domains with an evaluable correctness criterion — math, code correctness, structured reasoning — where "lucky guessing is difficult" and an automated or model-based grader can consistently judge quality. It requires more ML maturity to do well: you need a reliable grader, a clear reward signal, and patience for a more exploratory training process. At the time of writing, RFT in Foundry is available on reasoning models like o4-mini, with gpt-5 RFT support GA but invitation-gated (verify current model availability before planning a project around RFT — this list changes frequently).
The practical decision rule: start with SFT unless you have a specific reason not to. Reach for DPO when you have comparative preference data and a style/alignment objective. Reach for RFT only when you have an objective, graders can score it, and you have the ML expertise to debug a reinforcement-learning-style training run that doesn't converge cleanly.
2. What LoRA Actually Does to the Model
Here's the part most "how to fine-tune" tutorials skip entirely, and it's the part that actually explains why Foundry fine-tuning is fast, affordable, and doesn't require you to provision a cluster of A100s.
Full fine-tuning updates every parameter in a model. For a model with tens of billions of parameters, that means storing full-precision gradients and optimizer states for every weight matrix — a memory and compute bill that scales with total parameter count, not with how much you're actually trying to change.
Low-Rank Adaptation (LoRA) starts from an observation: when you adapt a large pretrained model to a narrower task, the change in weights needed is typically low-rank — it lives in a much smaller subspace than the full weight matrix. So instead of updating the full weight matrix W (shape d × d), LoRA freezes W entirely and injects two small trainable matrices, A (shape d × r) and B (shape r × d), where r — the rank — is small (commonly 8 to 64, versus d which might be in the thousands).
During training, only A and B are updated. The effective weight used in the forward pass is W + BA. Since r ≪ d, the number of trainable parameters is a tiny fraction of the full matrix — often under 1% of total model parameters — which means:
- Training requires far less GPU memory (no need to hold optimizer state for the full weight matrix)
- Training is faster and cheaper per step
- The resulting adapter is small and portable — you can swap adapters without reloading the entire base model
- The frozen base model's general capabilities are preserved, because you never touch
W
This is precisely why Foundry's serverless fine-tuning tier doesn't ask you for GPU quota: LoRA-based training has a fundamentally smaller resource footprint than full fine-tuning, which is what makes a consumption-priced, multi-tenant fine-tuning service economically viable in the first place. It's also why continuous fine-tuning — taking an already fine-tuned model and tuning it again on new data — is cheap and fast: you're composing small adapter updates, not re-training a dense model from scratch.
3. Microsoft Foundry's Fine-Tuning Architecture
At a system level, a Foundry fine-tuning job moves through a consistent pipeline regardless of which technique (SFT/DPO/RFT) or product surface (serverless OpenAI fine-tuning, serverless Foundry Models fine-tuning, or managed compute) you use:
- Data ingestion and validation — your JSONL training (and optional validation) files are uploaded to the project, checked for UTF-8 + BOM encoding, schema conformance to the chat-completions message format, and size limits (under 512 MB per file).
- Job submission and queuing — the job is submitted with a base model reference, a training tier, optional hyperparameters, and an optional seed for reproducibility. Jobs queue behind other tenants' jobs sharing the same regional capacity pool.
-
Training compute — the platform trains LoRA adapters against the frozen base model, producing per-epoch checkpoints and streaming
train_loss,full_valid_loss,train_mean_token_accuracy, andfull_valid_mean_token_accuracymetrics. - Safety evaluation — before a checkpoint becomes deployable, it passes through a safety evaluation pass. This is also what happens when you pause a job mid-training: Foundry doesn't just freeze the process, it runs the safety check against the current checkpoint so you have a usable artifact even from an incomplete run.
- Checkpoint retention — the three most recent checkpoints are retained and deployable after a job completes, letting you deploy an earlier epoch if the final epoch overfit.
- Deployment — a deployable checkpoint becomes a named custom model deployment, callable through the same Chat Completions-compatible interface as any other Foundry model deployment.
The key architectural decision Microsoft made here is to treat fine-tuning as a managed, multi-tenant job system layered on top of shared capacity, not as a bring-your-own-compute workload. That's a deliberate trade-off (more on this in Section 5), and it's why "Standard," "Global," and "Developer" exist as distinct tiers — they're different ways of allocating that shared capacity to your job.
4. Training Tiers: Standard, Global, and Developer
This is an underappreciated lever. Foundry's serverless fine-tuning offers three training tiers that trade off cost, latency, and data residency — and picking the wrong one either costs you more than necessary or breaks a compliance requirement silently.
| Tier | Data residency | Cost | Queue time | When to use |
|---|---|---|---|---|
| Standard | Training stays in your resource's region | Baseline | Normal | Regulated workloads where data must not leave a specific region |
| Global | Data and weights copied to wherever capacity is available | Lower | Faster | No data-residency constraint; you want the cheapest, fastest path |
| Developer | No residency guarantee | Lowest | Variable — jobs may be preempted and resumed | Experimentation, iteration on hyperparameters, price-sensitive non-production training |
The Developer tier deserves a specific callout: it uses idle capacity, which means your job can be paused and resumed by the platform without warning, and there's no SLA. That's a fine trade for a team iterating through ten small-batch experiments to find the right hyperparameters, and a bad trade for a job gating a release deadline.
5. Preparing a Dataset That Won't Waste Your Budget
Every fine-tuning failure mode traces back to the dataset, so it's worth being precise about format and quality requirements rather than treating this as a formality.
Format. Training data must be JSON Lines (JSONL), UTF-8 encoded with a byte-order mark (BOM), each file under 512 MB, using the conversational message schema:
{"messages": [
{"role": "system", "content": "You are a support-ticket triage classifier. Respond with only a JSON object: {\"category\": str, \"priority\": \"low\"|\"medium\"|\"high\"}."},
{"role": "user", "content": "Our checkout page has been returning 500 errors for the last 20 minutes and we're losing orders."},
{"role": "assistant", "content": "{\"category\": \"incident\", \"priority\": \"high\"}"}
]}
Multi-turn conversations are supported in a single line, and you can mark specific assistant turns with "weight": 0 to exclude them from the loss calculation — useful when you want the model to learn from the later turn in a conversation but need the earlier turns present for context:
{"messages": [
{"role": "system", "content": "Support classifier, terse mode."},
{"role": "user", "content": "Checkout is down."},
{"role": "assistant", "content": "{\"category\": \"incident\", \"priority\": \"medium\"}", "weight": 0},
{"role": "user", "content": "It's affecting all customers, not just some."},
{"role": "assistant", "content": "{\"category\": \"incident\", \"priority\": \"high\"}", "weight": 1}
]}
Size. Jobs technically run with as few as 10 examples, but that's not enough to reliably shift model behavior. The practical floor is 50 high-quality examples for initial validation, with production-grade customization typically needing several hundred to a few thousand. Dataset quality matters more than quantity: a large set of mediocre or inconsistent examples will actively degrade output quality relative to the base model, because the model is faithfully learning your inconsistency.
The system message trap. Whatever system message you use during training must be used at inference time. This is not optional guidance — it's a hard behavioral dependency. If you fine-tune with a specific system prompt and then swap it out at call time (a very easy mistake when system prompts live in application config separate from training data), the fine-tuned behavior degrades unpredictably because the model learned the input distribution including that system message, not just the output distribution.
Vision fine-tuning. For multimodal models like GPT-4o and GPT-4.1, training examples can include image_url content blocks alongside text, useful for chart interpretation, document processing, and visual quality assessment tasks — the same JSONL conversational schema, just with richer content arrays.
6. Running a Fine-Tuning Job: Python SDK and REST
Here's a realistic, end-to-end SFT job using the OpenAI Python SDK against a Foundry resource (the same client works for Azure OpenAI-hosted models; parameters differ slightly for open-source models fine-tuned through the Foundry Models path).
# pip install openai azure-identity
import time
from openai import AzureOpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
# Production pattern: use Entra ID managed identity instead of API keys
token_provider = get_bearer_token_provider(
DefaultAzureCredential(), "https://cognitiveservices.azure.com/.default"
)
client = AzureOpenAI(
azure_endpoint="https://<your-foundry-resource>.openai.azure.com/",
azure_ad_token_provider=token_provider,
api_version="2024-10-21",
)
# Step 1: upload training and validation files
training_file = client.files.create(
file=open("support_triage_train.jsonl", "rb"),
purpose="fine-tune",
)
validation_file = client.files.create(
file=open("support_triage_valid.jsonl", "rb"),
purpose="fine-tune",
)
# Step 2: submit the fine-tuning job
job = client.fine_tuning.jobs.create(
training_file=training_file.id,
validation_file=validation_file.id,
model="gpt-4o-mini-2024-07-18",
suffix="triage-v3", # becomes part of the deployed model name
seed=42, # reproducibility across repeated runs
hyperparameters={
"n_epochs": 3,
"batch_size": -1, # -1 = platform picks ~0.2% of training set size
"learning_rate_multiplier": 0.1,
},
# extra_body carries tier selection on Foundry-specific API surfaces
extra_body={"trainingType": "Standard"},
)
print(f"Job submitted: {job.id}, status: {job.status}")
# Step 3: poll for completion and stream metrics
while True:
job = client.fine_tuning.jobs.retrieve(job.id)
print(f"status={job.status}")
if job.status in ("succeeded", "failed", "cancelled"):
break
time.sleep(60)
if job.status == "succeeded":
print(f"Fine-tuned model: {job.fine_tuned_model}")
# e.g. gpt-4o-mini-2024-07-18.ft-a1b2c3d4e5f6-triage-v3
# Step 4: list checkpoints to pick the best epoch, not just the last one
checkpoints = client.fine_tuning.jobs.checkpoints.list(job.id)
for cp in checkpoints.data:
print(cp.id, cp.metrics)
The equivalent REST call for job submission, useful when you're orchestrating training from a pipeline rather than a notebook:
curl -X POST "https://<your-foundry-resource>.openai.azure.com/openai/fine_tuning/jobs?api-version=2024-10-21" \
-H "Authorization: Bearer $AAD_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-4o-mini-2024-07-18",
"training_file": "file-abc123",
"validation_file": "file-def456",
"suffix": "triage-v3",
"hyperparameters": {
"n_epochs": 3,
"learning_rate_multiplier": 0.1
}
}'
Note the authentication pattern: production jobs should authenticate with Microsoft Entra ID tokens via managed identity, not static API keys — the same operational guidance that applies everywhere else in Foundry's access model.
7. Hyperparameters That Actually Move the Needle
Three hyperparameters are exposed, and each has a specific failure signature when misconfigured:
-
batch_size— the number of examples per forward/backward pass. Larger batches suit larger datasets and produce lower-variance updates but update less frequently. Left at-1, Foundry computes it as roughly 0.2% of your training set size, capped at 256. If your dataset is small (under ~500 examples) and you manually set a large batch size, you can end up with fewer effective gradient updates than the task needs — one symptom istrain_lossthat barely moves. -
learning_rate_multiplier— multiplies the base model's original pretraining learning rate. The recommended experimentation range is 0.02–0.2. Too high, and you'll seetrain_lossdropping fast whilefull_valid_lossclimbs — classic overfitting. Too low, and the model barely shifts from base-model behavior even after multiple epochs. -
n_epochs— full passes through the training set. Left at-1, it's chosen dynamically based on dataset size (smaller datasets generally need more epochs to make an impact; larger ones need fewer to avoid overfitting). Watchfull_valid_mean_token_accuracy: if it plateaus or regresses whiletrain_mean_token_accuracykeeps climbing, you've gone one or two epochs too far — which is exactly what checkpoints are for.
The diagnostic loop that actually works in practice: run a baseline job with defaults, inspect the loss/accuracy curves per epoch, and only then start adjusting — ideally one hyperparameter at a time, logged with a seed value so a rerun is directly comparable rather than confounded by random initialization differences.
8. Checkpoints, Pausing, and Continuous Fine-Tuning
A checkpoint is produced at the end of every training epoch, and it's a fully usable, independently deployable model — not just a debugging artifact. This matters for two reasons.
First, the final epoch is not always the best epoch. If validation loss starts climbing at epoch 3 while training loss keeps improving, the epoch-2 checkpoint is probably the better model to deploy, and Foundry retains the three most recent checkpoints specifically so you have that choice without re-running the job.
Second, you can pause a running job and still get a deployable artifact. When you pause, Foundry runs the safety evaluation against the current checkpoint before making it available — so if metrics are visibly diverging mid-run, you don't have to choose between "let it finish and waste compute" and "cancel and get nothing."
Continuous fine-tuning — treating an already fine-tuned model as the base model for a subsequent job — is how most production fine-tuning actually evolves over time. Rather than retraining from the foundation model every time you get new labeled data, you reference the fine-tuned model ID (which looks like gpt-4o-2024-08-06.ft-d93dda6110004b4da3472d96f4dd4777-ft) as the base model for the next round. Because LoRA adapters are cheap to train, this iterative pattern — ship, collect new examples from production misses, retrain on top of the current model, redeploy — is both fast and inexpensive compared to what continuous full fine-tuning would cost.
9. Deployment: PTU, Standard, and Automatic Deployment
Once you have a checkpoint you trust, deployment follows the same model as any other Foundry deployment: standard pay-per-token, or Provisioned Throughput Units (PTU) if you need guaranteed latency and throughput at scale. Fine-tuned models deploy as named custom models, callable through the same Chat Completions-compatible surface your application already uses for the base model — which means swapping from a prompted base model to a fine-tuned custom model is often a one-line deployment-name change in your application code, not a rewrite.
Automatic deployment — enabling a flag so a successful training job deploys itself without a manual step — is available for OpenAI models and requires the Foundry Owner role (or a custom role with Microsoft.CognitiveServices/accounts/deployments/write). It's a convenience for fast iteration loops, but treat it carefully in shared projects: an automatically deployed checkpoint still consumes deployment quota and starts incurring hosting costs immediately, whether or not someone evaluated it first.
10. A Real-World Developer Scenario: Structured Extraction at Scale
Consider a platform team running a support-ticket triage pipeline in front of a Foundry Agent Service workflow. The original implementation used a large general-purpose model with a long system prompt specifying category taxonomy, priority rules, and six few-shot examples — roughly 1,800 tokens of fixed overhead on every single classification call, at a volume of 400,000 tickets a month.
The fine-tuning play here is almost a textbook SFT case: the task is narrow, the input/output mapping is well-defined (ticket text → JSON classification), and there's a year of historical tickets with human-assigned categories sitting in a support database — a dataset, not a wishlist.
The engineering sequence:
- Export 3,000 historical tickets with validated category/priority labels into the JSONL message format, holding out 300 for validation.
- Run a baseline SFT job on
gpt-4o-miniwith default hyperparameters and the Developer tier (cheap, no SLA needed for an experiment). - Inspect
full_valid_mean_token_accuracyacross epochs, pick the best checkpoint (epoch 2 of 4, since epoch 3–4 showed validation accuracy regressing — classic overfitting on a dataset this size). - Deploy that checkpoint to a Standard (non-PTU) deployment, run it against a held-out production sample side-by-side with the original prompted baseline using Foundry's agentic evaluators (
TaskAdherenceEvaluator, plus a custom grader comparing classification accuracy against ground truth). - Once accuracy parity (or improvement) is confirmed, cut the system prompt down to a single sentence — the examples and rules are now baked into the weights — and redeploy.
The outcome: a 1,800-token system prompt collapses to roughly 60 tokens, latency drops because there's less input to process, and per-call cost drops proportionally to the token reduction at 400,000 calls/month — a saving that compounds specifically because this is a high-volume, narrow-task endpoint, which is exactly the profile where fine-tuning's economics work best.
11. Production Considerations
-
Version your fine-tuned models like code. The
suffixparameter and the resultingft-{jobid}identifier are your only built-in versioning signal — track which dataset version, hyperparameters, and seed produced which deployed model ID in your own change-management system, because Foundry doesn't do this bookkeeping for you across unrelated jobs. - Don't skip the held-out evaluation step. Loss curves tell you whether training converged; they don't tell you whether the model is actually better at your task than the base model with a good prompt. Always run a side-by-side comparison against both the base model and the previous fine-tuned version using the same evaluation set.
- Re-fine-tune on a cadence, not just reactively. Production misses accumulate; a continuous fine-tuning loop that periodically retrains on recent misclassifications (reviewed by a human before being added to the training set) keeps the model current without needing a full dataset rebuild each time.
- Watch for distribution drift between training data and production traffic. A model fine-tuned on last year's ticket categories will quietly underperform once your product adds three new feature areas that generate new, unseen ticket types.
12. Security and Governance
Fine-tuning introduces a data-handling surface that's easy to overlook in a security review: your training data — which may contain customer support transcripts, internal documents, or proprietary business logic — is uploaded to the service and used to produce a model artifact that can, in principle, leak details of its training data through its outputs (a well-documented risk for fine-tuned LLMs generally).
Specific controls to apply:
- Choose the Standard training tier for regulated data. Global tier training explicitly copies data and weights outside your resource's region for cheaper, faster capacity access — a reasonable trade-off for public or synthetic data, a potential compliance violation for regulated data. Default to Standard when data residency has any legal or contractual bearing.
- RBAC matters more here than for inference-only usage. Training and deploying fine-tuned models requires elevated roles (Foundry Owner or equivalent custom roles with deployment-write permissions) — scope these tightly, since a fine-tuned model deployment is effectively a new, persistent artifact derived from potentially sensitive data.
- Scrub PII from training examples before upload, not after. Once data is baked into model weights through training, you can't selectively redact it from the resulting model the way you can delete a row from a database.
- Treat the safety evaluation gate as a backstop, not a substitute for dataset review. The automated safety pass catches gross policy violations; it will not catch a subtly biased labeling pattern in your own historical data (e.g., if past human triage decisions systematically under-prioritized certain ticket sources).
13. Cost Considerations
Fine-tuning cost has three distinct components, and conflating them is how budgets get blown:
- Training cost — billed for the training compute/tokens consumed during the job itself, varying by tier (Global cheapest, Standard baseline, Developer cheapest-but-preemptible) (confirm current per-model, per-tier pricing in the Azure pricing calculator before budgeting — rates change independently of this article).
- Storage cost — training/validation files and retained checkpoints consume storage over time.
- Hosting cost — once deployed, a fine-tuned custom model is billed like any other deployment: pay-per-token for Standard deployments, or the hourly/reserved PTU rate if you need guaranteed throughput. This is usually the dominant long-run cost, not the training job itself.
The net economic case for fine-tuning only closes at sufficient call volume: the training cost and ongoing marginal hosting overhead need to be offset by token savings from shorter prompts (fewer input tokens per call) and/or the ability to use a smaller, cheaper base model that now performs at the level of a larger one for your narrow task. Low-volume, highly varied tasks rarely clear this bar — that's a signal to stay with prompting or retrieval instead.
14. Common Mistakes and Pitfalls
- Fine-tuning to inject facts. If the goal is "the model should know about our Q3 product catalog," that's a retrieval problem (see Foundry IQ), not a fine-tuning problem. Fine-tuning shapes behavior and style, not a live, updatable knowledge base — a fine-tuned model's "knowledge" is frozen at training time and expensive to refresh compared to just updating a search index.
- Inconsistent system messages between training and inference. Covered above, but worth repeating because it's the single most common silent failure mode reported against fine-tuned deployments.
- Treating "more epochs" as strictly better. Past the point where validation metrics plateau or regress, additional epochs actively overfit — memorizing training examples at the cost of generalization.
- Skipping a held-out evaluation set entirely. Training on 100% of available data with no validation split means you have no signal for when the model starts overfitting, and no honest measurement of whether the fine-tuned model actually beats the baseline.
- Deploying the final checkpoint by default instead of the best checkpoint. The UI and SDK make the last epoch the path of least resistance; the metrics frequently say otherwise.
- Choosing RFT or DPO because they sound more sophisticated than SFT. Technique selection should follow from what your data actually looks like (fixed targets vs. preference pairs vs. gradable objectives) — not from technique prestige.
15. Alternatives and Trade-offs
| Approach | Best for | Weak for |
|---|---|---|
| Prompt engineering | Fast iteration, no training data needed, behavior changes instantly | High token overhead at scale, brittle to drift, ceiling on complex behavior change |
| RAG / Foundry IQ | Fresh, updatable factual knowledge; citeable sources | Doesn't change model behavior or style; adds retrieval latency |
| Fine-tuning (SFT/DPO/RFT) | Stable behavior/style/format change, token cost reduction at high volume, teaching narrow task competence | Frozen knowledge, requires a real dataset, added operational surface (versioning, retraining cadence) |
| Agent Optimizer (prompt/tool optimization) | Tuning an existing prompt-based agent's instructions and tool descriptions without training a new model | Doesn't change underlying model weights — ceiling is whatever the base model can do with a better prompt |
| Model Router | Automatically picking the cheapest adequate model per request | Doesn't improve any individual model's task performance |
In most mature Foundry deployments, these aren't mutually exclusive — a fine-tuned model still benefits from RAG for facts, still runs behind an Agent Optimizer pass for its remaining prompt surface, and still sits behind Model Router if there's a mix of easy and hard requests.
16. Practical Recommendations
- Start every fine-tuning project with a baseline SFT job on the smallest capable model, using the Developer tier, before committing to a larger model or a more exotic technique.
- Build your held-out evaluation set before you start iterating on hyperparameters, and reuse it across every job so comparisons are apples-to-apples.
- Keep your training data pipeline reproducible — version the JSONL files alongside your application code, not just in an ad hoc export folder.
- Default to the Standard training tier unless you've explicitly confirmed your data has no residency constraints.
- Treat the deployed fine-tuned model as another artifact in your release process: canary it, evaluate it continuously in production (Foundry's continuous evaluation rules apply to fine-tuned deployments the same as any other), and keep the previous version's deployment warm until the new one is proven.
17. Conclusion
Fine-tuning in Microsoft Foundry is deceptively simple from the portal — upload a file, pick a model, click submit — and deceptively deep underneath: a LoRA-based training architecture that makes the economics work, three distinct techniques (SFT, DPO, RFT) mapped to three distinct problem shapes, a tiered training system trading off cost against data residency, and a checkpoint model that assumes, correctly, that your last epoch is not always your best one. Developers who treat it as a black box tend to end up with a fine-tuned model that's marginally different from the base model and hard to explain. Developers who understand what's happening to the weights — and what the loss curves are actually telling them — end up with models that are measurably cheaper, faster, and more reliable at the one thing they were built to do.
If you're sitting on a bloated system prompt and a year of labeled production data, that's not a prompt-engineering problem anymore. Go run the baseline job.
18. References
- Microsoft Learn — Customize a model with fine-tuning (Microsoft Foundry)
- Microsoft Learn — Fine-tune models with Microsoft Foundry (classic) — concepts and use cases
- Azure-Samples — AIFoundry-Customization-Datasets
- Hu et al., 2021 — LoRA: Low-Rank Adaptation of Large Language Models (foundational paper behind the adaptation technique Foundry's fine-tuning service builds on)
This is Day 22 of the Microsoft Foundry 100 Days / 100 Blogs series — one deep technical article a day on Microsoft Foundry's architecture, APIs, and production patterns.



Top comments (0)