MLOps: models that survive production · Chapter 2 of 12 · AI Engineering · new chapter every Thursday night
By the end of this chapter: Reproduce a training run from six months ago.
The problem
Say you are an engineer who has just been asked to retrain the churn prediction model from Q1. A new enterprise customer needs it deployed in their own environment, but with a slight tweak to the output format. You check out the Git tag release-q1, run the training script, and wait.
The original run logged an F1 score of 0.88. Your new run logs 0.64.
You have not changed a single line of code. The random seeds are pinned. The dependencies are locked in a requirements file. But the data in s3://company-ml-data/churn/latest.csv is from today. Since Q1, the upstream engineering team has added three new columns, changed the definition of an 'active user', and dropped all rows from a deprecated region.
You spend three days trying to manually reconstruct the Q1 dataset from database backups. You fail. You have to explain to your lead that the model currently running in production cannot be rebuilt, audited, or modified, because the data it was trained on no longer exists.
Before you start
You need a terminal, Python, Git, and DVC (Data Version Control).
Install DVC via pip. Do not use Homebrew or apt for this; keeping DVC in your Python environment ensures your CI pipeline and your local machine run the same version.
pip install dvc==3.48.4
Verify the setup. This command must return a version number for both tools:
git --version && dvc --version
Create a new directory for this exercise and move into it:
mkdir churn-model && cd churn-model
Why Git cannot hold your data
The instinct when faced with this problem is to put the data in Git alongside the code. If the code and data are in the same commit, checking out the commit restores both.
Git is designed for text. It tracks changes line by line. If you commit a 5GB CSV file, Git will attempt to compress it and store the delta. It will fail to do this efficiently. Your .git directory will balloon. Cloning the repository will take hours. Eventually, you will push to GitHub, which will reject any file larger than 100MB, and you will have to rewrite your Git history to remove the commit.
Git Large File Storage (LFS) is the standard workaround in software engineering, but it fails in machine learning. Git LFS tightly couples your data storage to your Git provider. If you have 5TB of training data, storing it in GitHub LFS is prohibitively expensive compared to an AWS S3 bucket. Furthermore, Git LFS lacks the semantics for data pipelines; it cannot tell you if a dataset was generated by a specific script, only that the file changed.
You need Git to track the exact state of your data, without Git ever touching the data itself.
The pointer file pattern
To solve this, we use the pointer file pattern. This is the mechanism underlying DVC.
Instead of committing dataset.csv to Git, you hash the file's contents. You store the heavy dataset.csv in a dumb object store (like S3, or just a hidden folder on your laptop) renamed to its hash, for example a1b2c3d4.csv.
Then, you create a tiny text file called dataset.csv.dvc. This file contains nothing but the hash a1b2c3d4 and the file size. You commit this tiny text file to Git.
Git tracks the code and the pointer. DVC tracks the heavy file. When you check out a Git commit from six months ago, Git updates dataset.csv.dvc to contain the old hash. You then tell DVC to read that pointer, find the corresponding heavy file in the object store, and copy it back into your workspace.
Code and data are synchronised, but Git only handles text.
Build it
We are going to create a dataset, version it, overwrite it with new data, and then successfully time-travel back to the original state.
Step 1: Initialise the repositories
Initialise Git, then initialise DVC. DVC will create a .dvc directory to act as your local object store.
git init
dvc init
DVC creates some internal configuration files. Commit them to Git so your repository is clean.
git commit -m "Initialise DVC"
Step 2: Generate the Q1 data
Create a Python script named generate_data.py to simulate our Q1 dataset.
# generate_data.py
import csv
def write_data(filename, rows):
with open(filename, 'w', newline='') as f:
writer = csv.writer(f)
writer.writerow(['user_id', 'active_days', 'churned'])
writer.writerows(rows)
# Q1 Data
write_data('dataset.csv', [
[1, 15, 0],
[2, 3, 1],
[3, 20, 0]
])
print("Wrote Q1 data to dataset.csv")
Run it:
python generate_data.py
Step 3: Track the data with DVC
Tell DVC to track the CSV file.
dvc add dataset.csv
Look at the output. DVC tells you exactly what to do next:
To track the changes with git, run:
git add dataset.csv.dvc .gitignore
When you ran dvc add, DVC did three things:
- It calculated the MD5 hash of
dataset.csv. - It moved
dataset.csvinto.dvc/cache/under its hash name, and put a hard link back in your workspace. - It created
dataset.csv.dvc(the pointer) and updated.gitignoreso you do not accidentally commit the raw CSV to Git.
Step 4: Commit the pointer to Git
Commit the code, the pointer, and the gitignore file. This creates our Q1 snapshot.
git add generate_data.py dataset.csv.dvc .gitignore
git commit -m "Train Q1 model"
git tag release-q1
Step 5: Generate the Q2 data
Time passes. The definition of the data changes. Modify generate_data.py to output the Q2 data. Change the rows and add a new column.
# generate_data.py
import csv
def write_data(filename, rows):
with open(filename, 'w', newline='') as f:
writer = csv.writer(f)
# Schema change: added 'region'
writer.writerow(['user_id', 'active_days', 'region', 'churned'])
writer.writerows(rows)
# Q2 Data
write_data('dataset.csv', [
[4, 30, 'EU', 0],
[5, 2, 'US', 1],
[6, 45, 'EU', 0]
])
print("Wrote Q2 data to dataset.csv")
Run it to overwrite the CSV:
python generate_data.py
Step 6: Track and commit the Q2 data
Update DVC with the new file, then update Git with the new pointer.
dvc add dataset.csv
git add generate_data.py dataset.csv.dvc
git commit -m "Train Q2 model"
If you open dataset.csv now, you will see the Q2 data with the region column.
Step 7: Time travel
You are asked to reproduce the Q1 model. First, check out the old Git commit using the tag we made.
git checkout release-q1
Look at dataset.csv. It has not changed. It still contains the Q2 data.
This is the most common point of confusion. Git only updated the text files it tracks. It updated generate_data.py and it updated the pointer file dataset.csv.dvc. It did not touch the actual CSV.
To synchronise your workspace with the pointer file, run:
dvc checkout
Output:
M dataset.csv
Open dataset.csv again. The region column is gone. The data is exactly as it was in Q1. You can now run your training script and get the exact 0.88 F1 score you got six months ago.
When this breaks
The file is modified but dvc checkout does nothing
You manually edited dataset.csv to fix a typo. You realise you made a mistake, so you run dvc checkout to restore the file to the state of the pointer. Nothing happens. The typo is still there.
DVC uses file timestamps and sizes to optimise checkouts. If you edit a file in place without changing its size enough, DVC might not realise it has changed. To force DVC to verify the hashes and overwrite your local modifications, use:
dvc checkout --force
ERROR: unexpected error - .dvc/cache is not tracked by git
You cloned your repository to a new machine, ran dvc checkout, and got an error about a missing cache or a missing file.
ERROR: failed to pull data from the cloud - Access Denied
The .dvc/cache directory is in your local .gitignore. It does not push to GitHub. When you clone the repository, you get the pointers, but not the heavy files. To get the heavy files, you must configure a DVC remote (like an S3 bucket) using dvc remote add, push the data there from the original machine (dvc push), and pull it on the new machine (dvc pull). If you get Access Denied, your AWS credentials in your terminal session do not have s3:GetObject permissions for that bucket.
Git merge conflicts in .dvc files
You and a colleague both updated the dataset on different branches. When you merge, Git throws a conflict in dataset.csv.dvc.
A .dvc file is just YAML containing a hash. You cannot merge hashes. You must pick one. Accept either your changes or their changes in the Git merge tool. Then, whichever one you picked, run dvc checkout to ensure your local CSV matches the hash you just committed to the merge.
What it costs
This approach costs storage space, and it scales poorly for append-only data.
Because DVC hashes the entire file, it does not do delta compression for tabular data. If you have a 10GB CSV file and you append one row, dvc add will hash the new file, copy the entire 10GB into .dvc/cache, and update the pointer. You are now storing 20GB of data. If you do this every day for a month, you are storing 300GB of redundant data.
For datasets that update frequently, flat-file versioning with DVC is the wrong tool. You are paying the storage cost of full duplication. In those scenarios, you must move to a time-travel database format like Apache Iceberg or Delta Lake, which version data at the row or block level, or rely on a Feature Store.
DVC is best used for unstructured data (images, audio) where files are immutable and added discretely, or for tabular data that is generated as a distinct, immutable artifact for a specific training run.
It also costs developer friction. Every data change now requires two commands (dvc add, then git add). If you forget dvc add, you will commit a stale pointer to Git, and your CI pipeline will train on old data while your Git history claims otherwise.
In the interview
When an interviewer asks, "How do you ensure your training runs are reproducible?", they are checking if you understand that code is only half the state of an ML system.
A weak answer focuses entirely on the environment: "I set the random seeds, pin the requirements, and use Docker." This is weak because it assumes the data is static. In production, data is a moving target. If you run the same Docker container against a database that has changed, the output changes.
A strong answer explicitly names the pointer file pattern. "I version the code, the environment, and the data. I use a tool like DVC to hash the training artifacts, store them in an object store, and commit the pointer files to Git. This guarantees that checking out a Git commit restores the exact bytes of data used for that run."
If you are interviewing for a Senior or Staff role, the interviewer will probe the limits of this. They will ask what happens when the dataset is 500GB and updates daily. A senior candidate must immediately identify the storage duplication trade-off. You should explain that DVC is for artifact versioning, not database versioning. For daily updates at scale, you would advocate for Delta Lake to query the data "as of" a timestamp, or a Feature Store that guarantees point-in-time correctness, rather than hashing flat files.
The follow-up question that separates practitioners from theorists is: "What happens if someone modifies the data file but forgets to run dvc add before committing to Git?" The practitioner knows exactly what happens: the Git commit contains the old hash, the new data is left untracked in the workspace, and anyone pulling the repository will get the old data, silently invalidating the training run.
Your tasks
-
Configure a local remote. By default, DVC stores the heavy files in
.dvc/cache. Create a new directory outside your Git repository (e.g.,mkdir /tmp/dvc-remote). Usedvc remote add -d myremote /tmp/dvc-remoteto configure it. Rundvc push. Delete your local.dvc/cachedirectory, then rundvc pull. Verifydataset.csvis restored. -
Break the sync. Open
dataset.csvand manually change a value. Do not rundvc add. Rundvc status. It will tell you the file is modified. Use the DVC CLI to discard your manual changes and restore the file to match the Git pointer. -
Read the pointer. You do not strictly need the DVC CLI to know what data a Git commit used. Write a three-line Python script that opens
dataset.csv.dvc, parses it as YAML (or just string splits it), and prints the MD5 hash. This is how CI systems verify data without downloading the whole DVC binary.
Your tasks this week
Do the exercises above before the next chapter. Reading a tutorial and doing
one are different activities and only one of them changes what you can build.
Stuck on any of them? Say so — describe what you tried and what happened:
tell me where you got stuck. I read every one, and the questions
that come back more than twice get answered in the next chapter.
MLOps: models that survive production
Chapter 2 of 12. New chapter every Thursday night.
Next: Features: stores, and the simpler thing that usually works.
· The full syllabus and every chapter so far
· Subscribers also get the condensed notes for this chapter, the running
recap of everything the series has covered, and the extended guidance:
subscribe
Written by Amit Chakraborty — founding engineer and senior architect: React Native, AI and RAG systems, production architecture. Portfolio · LinkedIn · GitHub.
Need this built, reviewed or taught to your team? Get in touch or email amit@devamit.co.in. Available for senior and founding engineering roles, consulting and training, remote worldwide.
Top comments (0)