DEV Community

Cover image for We Upgraded to Python 3.15 for the JIT Speedup and Found a Silent Encoding Bug Instead
Rohit Bhadani
Rohit Bhadani Subscriber

Posted on

We Upgraded to Python 3.15 for the JIT Speedup and Found a Silent Encoding Bug Instead

This is a version-upgrade failure mode that's common enough to deserve its own name — walked through the way it actually surfaced, not summarized after the fact.

Python 3.15 landed in October 2026 with a genuinely exciting number attached: an upgraded JIT compiler running meaningfully faster, double digits on some platforms. We run a data pipeline that's CPU-bound enough that a free double-digit speedup was worth bumping the version for on its own. We upgraded a staging worker, ran the usual smoke tests, saw lower CPU time, and scheduled the production rollout for later that week.

Two days into the staging run, a nightly report that summarizes customer activity came back with garbled text in a handful of names — not crashed, not missing, just wrong. A few Cyrillic and accented names had turned into mojibake, the kind of corrupted-looking text that screams "encoding mismatch" to anyone who's seen it before.

The wrong turns

First guess: a database driver issue. We assumed the newer Python version had picked up a newer version of our database driver as a transitive dependency, and that driver had a new default character set. We pinned the driver version explicitly and reran. Same garbled names.

Second guess: a locale difference between staging and the old environment. Different base images sometimes have different system locales, and locale has bitten us before with sorting and string comparison. We checked — identical locale settings on both. Not this either.

Third guess: the report generation code had an actual bug that had always existed and we were only now seeing it by coincidence. We ran the exact same report against the exact same data on the old Python version, on the old worker. Clean output, correct names. Same code, same data, same libraries — different Python version, different result.

What had actually changed

Python 3.15 changed the default text encoding assumption in a specific set of cases involving file I/O without an explicit encoding argument — formalizing a years-long effort to make UTF-8 the default rather than relying on locale-dependent encoding in places where it used to be ambiguous. Our report-writing code opened an output file like this:

with open(report_path, "w") as f:
    f.write(formatted_report)
Enter fullscreen mode Exit fullscreen mode

No encoding= argument. On the old Python version and our old container's locale configuration, this happened to resolve to an encoding that matched the actual byte content of names already stored as UTF-8 in the database. On Python 3.15, the same code path now resolves differently in a way that exposed a mismatch that had technically always been latent — the code had never explicitly specified an encoding, it had just been accidentally correct for years because of how the implicit default happened to line up with our data.

This is the specific shape of bug that's worth naming: not a new bug introduced by the upgrade, but an old, dormant bug whose two incorrect assumptions had been canceling each other out, until a version upgrade changed one of them and the cancellation stopped working.

The fix

1. Make every encoding assumption explicit, everywhere, immediately — not just in the one file that broke.

with open(report_path, "w", encoding="utf-8") as f:
    f.write(formatted_report)
Enter fullscreen mode Exit fullscreen mode

2. We added a lint rule that fails CI on any open() call without an explicit encoding argument, project-wide, to find every other place the same dormant assumption was hiding.

# .flake8 or equivalent
[flake8]
select = W1514
# W1514: Using open without explicitly specifying an encoding
Enter fullscreen mode Exit fullscreen mode

3. Before any future interpreter upgrade touches production, it now runs first against a disposable clone of production with real production-shaped data — not synthetic test fixtures, which happened to be pure ASCII and would never have caught this. I'm the founder of Krova Cloud, and this is exactly why we built fast, disposable VM cloning the way we did: spinning up a clone with a real data snapshot, running the full pipeline against it on the new interpreter version, and diffing the output against the current production output catches exactly this category of "technically always broken, never triggered" bug before it reaches anything that matters. It takes a few minutes and costs close to nothing, against the alternative of corrupted customer-facing reports in production.

Lessons

  • A language version upgrade with an exciting performance number attached is still a behavior change, not just a speed change. Read the changelog for default-behavior shifts, not just the headline feature.
  • Code that's "always worked" without an explicit encoding, timezone, or locale argument is frequently two silent incorrect assumptions canceling out, not one correct one. The cancellation can break the moment either side of it changes.
  • Synthetic test fixtures that happen to be pure ASCII, UTC-only, or otherwise unrepresentative of real data will never catch this class of bug. Testing against a real, representative data snapshot is the only way to find it before production does.
  • A lint rule that enforces explicit encoding (or timezone, or locale) arguments project-wide is cheap insurance against an entire category of future upgrade surprises, not just this one.

If you're planning a Python 3.15 upgrade for the performance win, it's worth a project-wide search for open( calls without an explicit encoding before you schedule it.


I'm Rohit, founder of Krova Cloud — disposable VM clones from real production snapshots, built for testing upgrades and migrations against real data before they touch anything that matters. If you want more deep debugging stories like this one, I write regularly over at debugly.dev too.

Top comments (1)

Collapse
 
suppdevbot profile image
DEV SUPPORTS •

You need to verify your account.

Enter fullscreen mode Exit fullscreen mode

tr.ee/dev-to