TL;DR
I upgraded a ~70,000-line Django 3.2 monolith to Django 5.2 LTS in six working days, with Claude Code doing most of the mechanical work. The trick wasn't a clever prompt. It was hopping one LTS at a time, turning deprecation warnings into hard errors, and giving the agent a single, boring loop to follow. Here's the exact setup, the three places the agent got it wrong, and five lessons I'd apply to any framework upgrade.
The Problem
Django 3.2 went end-of-life in April 2024. Our app was still on it in 2026. Not because nobody cared, but because every time someone opened the upgrade ticket, they'd look at the changelog for four major releases and quietly close the tab.
Here's what we were dealing with:
- ~70,000 lines of Python across 31 Django apps
- Python 3.9 in production (Django 5.2 needs 3.10+)
- 46 third-party packages, several of which hadn't seen a release in years
- ~2,100 tests, about 78% line coverage, with a handful of known-flaky ones
- One team member who remembered why
USE_TZwas set the way it was (spoiler: they didn't)
The constraint that made it interesting: I had roughly one sprint, and I couldn't freeze feature work. Whatever I did had to land in small, reviewable PRs that the rest of the team could keep merging around.
I'd already been using Claude Code (2.x) for refactors, so the question was: can an AI coding agent grind through a multi-version framework upgrade without me babysitting every file?
Short answer: yes, if you take away its freedom.
How I Solved It
Step 1: Hop LTS to LTS, never skip
The single most important decision was refusing to jump straight from 3.2 to 5.2. Django's deprecation policy is designed around this: a feature deprecated in version X keeps working (with a warning) until X+2. If you jump too far, the warnings that would have told you what to fix are already gone, and you just get ImportErrors.
So the plan looked like this:
flowchart LR
A[Django 3.2<br/>Python 3.9] --> B[Python 3.10<br/>Django 3.2]
B --> C[Django 4.2 LTS<br/>Python 3.10]
C --> D[Python 3.12<br/>Django 4.2]
D --> E[Django 5.2 LTS<br/>Python 3.12]
Each arrow is its own branch, its own green CI run, and its own deploy. Python and Django never moved in the same PR. When something broke in staging, I knew exactly which variable changed.
Step 2: Make deprecation warnings fatal
Django emits RemovedInDjangoXXWarning classes, which subclass Python's DeprecationWarning family. By default those are mostly silent. I flipped them into hard failures in the test run:
# Run the full suite with every deprecation warning raised as an exception
python -W error::DeprecationWarning \
-W error::PendingDeprecationWarning \
manage.py test --parallel 4 2>&1 | tee /tmp/upgrade-run.log
That turned a vague "upgrade Django" task into a concrete, finite list of failing tests with stack traces pointing at the exact line. That list is exactly the kind of thing an agent is good at chewing through.
One gotcha: third-party packages also raise these warnings, and you can't fix their code. I added targeted ignores in pytest.ini-style config (we use the Django runner, so it went in a small warnings.filterwarnings block in the test settings) scoped by module, so our own code stayed strict:
# settings/test.py
import warnings
warnings.filterwarnings("error", category=DeprecationWarning)
# Vendor noise we can't fix yet; tracked in the upgrade checklist
warnings.filterwarnings(
"ignore",
category=DeprecationWarning,
module=r"some_unmaintained_package\..*",
)
Step 3: Let a codemod do the boring 60%
Before the agent touched anything, I ran django-upgrade, a deterministic rewriter for Django-specific patterns:
git ls-files -- '*.py' | xargs django-upgrade --target-version 4.2
It handled the boring, high-volume stuff: url() to re_path(), ugettext to gettext, request.is_ajax() replacements, index_together hints, and a lot more. That one command touched about 340 files.
Why run this before the agent? Because a deterministic tool is cheaper, faster, and never hallucinates. I wanted the agent's attention spent on the judgment calls, not on find-and-replace.
Step 4: Give the agent one loop, and nothing else
This is where Claude Code came in. The instruction I gave it was intentionally narrow. Roughly:
## Upgrade loop (Django 4.2 hop)
1. Run the test command in `scripts/upgrade-test.sh`.
2. Take the FIRST failing test only.
3. Read the full traceback and the relevant Django 4.x release notes section.
4. Make the smallest change that fixes it. Do not refactor nearby code.
5. Re-run that single test, then the app's test module.
6. Commit with message: `upgrade(4.2): <what changed and why>`.
7. Go back to step 1.
Stop and ask me if:
- a fix requires changing a settings default
- a fix touches a migration file
- a third-party package needs a version bump
The "first failing test only" rule did a lot of work. Without it, the agent would see 140 failures, try to fix 30 at once, and produce a diff nobody could review. With it, I got a steady stream of tiny commits, most under 20 lines.
The "stop and ask" list was the other half. Those three categories are where upgrades actually hurt you in production, and I wanted a human looking at every one of them.
Step 5: Triage third-party packages up front
The 46 dependencies were the real schedule risk. I had the agent build a table before touching any code:
package current supports 4.2? supports 5.2? action
----------------------- -------- ------------- ------------- -----------------
django-filter 2.4.0 yes (23.x) yes (25.x) bump
django-storages 1.11 yes (1.14) yes (1.14) bump + STORAGES
some_unmaintained_pkg 0.9.2 unknown no replace / vendor
...
It filled the "supports" columns by reading each package's changelog and classifiers on PyPI. I spot-checked about a third of them. It got two wrong (it trusted a classifier list that was stale), which is why I don't let it do this unsupervised.
Out of 46 packages: 38 just needed a version bump, 5 needed config changes, and 3 had to be replaced or vendored. Knowing that on day one meant no surprises on day five.
Where the Agent Got It Wrong
Three moments where I'm glad the "stop and ask" rule existed.
⚠️ 1. USE_TZ silently flipping
In Django 5.0, the default for USE_TZ changed from False to True. Our settings file never set it explicitly. We'd been relying on the old default for years.
The agent noticed the deprecation warning in the 4.2 hop and proposed adding USE_TZ = True "to match the new default." That would have made every naive datetime we stored start being interpreted as UTC. For a product with scheduled jobs in local time, that's a quiet data-corruption bug.
The right fix was the opposite: pin USE_TZ = False explicitly, ship the upgrade, and then do the timezone migration as its own project. Because this touched a settings default, the agent stopped and asked. ✅
⚠️ 2. CSRF_TRUSTED_ORIGINS needing a scheme
Django 4.0 started requiring a scheme in CSRF_TRUSTED_ORIGINS (https://example.com, not example.com). The agent fixed it correctly in settings, but only for the value in the base settings file. Our staging and production values came from environment variables parsed in a different module.
Tests passed. Staging login broke. 🙃
Lesson there was mine, not the agent's: tests only cover what's in the test settings. I added a startup check that validates every origin has a scheme and fails loudly at boot.
⚠️ 3. Form rendering changed every template
Django 5.0 switched the default form rendering to div-based templates. Our test suite didn't catch it because tests assert on behavior, not markup. But about a dozen pages had CSS that targeted the old table/paragraph structure.
The agent's first suggestion was to rewrite the CSS. I asked it to instead pin the old renderer for the upgrade PR and open a follow-up to migrate templates intentionally. Again: upgrade first, modernize second. Mixing them is how a one-week project becomes a one-month project.
The Numbers
After six working days:
- 4 deploy-sized PRs (Python 3.10, Django 4.2, Python 3.12, Django 5.2), plus ~20 small prep PRs
- 187 agent commits, median size 11 lines changed
- ~340 files rewritten by the codemod, ~160 touched by the agent
- 0 rollbacks in production, 1 staging incident (the CSRF one)
- Test suite runtime dropped from 14m to 11m on Python 3.12, which I didn't expect but will happily take 🚀
I spent most of my time reviewing, answering the agent's "stop and ask" questions, and reading release notes for the edge cases. I wrote maybe 200 lines of code myself.
Lessons Learned
1. Small hops beat big jumps, always. Deprecation warnings are a map. If you skip versions, you throw the map away. Hop LTS to LTS, and keep language and framework upgrades in separate PRs.
2. Turn the upgrade into a failing test list before involving the agent. "Upgrade Django" is a vague goal. "Make these 140 tests pass, one at a time" is a task. Agents are dramatically better at the second one.
3. Use deterministic tools first, agents second. Codemods are free, fast, and correct. Save the agent for the work that needs reading and judgment. My rough rule now: if a regex could do it, the agent shouldn't.
4. Defaults are the most dangerous diff. The scariest changes in this upgrade weren't renamed imports. They were settings whose default changed under us. Any time a framework changes a default, pin the old value explicitly, ship, then migrate on purpose.
5. Write down where the agent must stop. My "stop and ask" list had three items: settings defaults, migrations, and dependency bumps. That list caught every risky change. An agent with no stop conditions will happily "fix" things in the most locally reasonable and globally wrong way.
What's Next
The upgrade is done, but it left a short list of intentional follow-ups:
- Migrate to
USE_TZ = Trueproperly, with a data audit of every naive datetime column - Move our templates to the new
div-based form rendering and drop the pinned renderer - Replace the last vendored package with a maintained alternative
I'm also turning the upgrade loop above into a reusable template, so the next hop (Django 6.x LTS, whenever it lands) is a two-day job instead of a six-day one.
Wrap-up
If you've been staring at an old framework version and dreading the changelog, my honest take: the mechanical part is no longer the hard part. The hard part is deciding what the agent is not allowed to touch.
👉 If this was useful, follow me here on Dev.to. I write build logs like this every week about shipping real software with AI coding agents. And if you've done a big framework upgrade with Claude Code (or failed to), drop your war story in the comments. I'd love to compare stop-and-ask lists. 💡
Top comments (0)