DEV Community

DataDriven
DataDriven

Posted on

TRM Labs' 400 TB Migration: Lessons for Data Engineer Interviews

I've been asked some version of "tell me about a large migration you led" in more interview loops than I can count. I've given good answers and I've given answers so bad that I watched the interviewer quietly stop taking notes. The bad ones all had the same problem: I told a story about moving data. The good ones covered how I proved the plan, how I found out it was wrong, and what I changed.

That's why I've been sending people the post TRM Labs published on June 30, 2026. It's a candid write-up of moving roughly 400 TB of Postgres onto StarRocks on Apache Iceberg, and it includes the slipped deadline and the latency regressions that reached customers. Almost nobody publishes that part, which makes it one of the best public case studies a Data Engineer can use to prepare for a migration question in a system design or behavioral interview.

I think interviewers barely care about your cutover. They want to see how you de-risk, how you report bad news, and whether you can back your numbers. The TRM post covers all 3, including the uncomfortable parts.

What TRM Labs migrated, and what the headline leaves out

The facts, straight from TRM's engineering post by Amit Plaha. The source was 2 Postgres clusters in a blue/green setup, around 400 TB covering 70+ blockchains. The target is StarRocks on Apache Iceberg, served through NGQF (NextGen Query Federation). The scope was 30+ data-intensive API routes, with business logic moved out of Postgres functions and into a TypeScript layer. The table that started it all is Address Transfers, which shows how a blockchain address has sent and received funds; it was the first table they ever created, and it was timing out on Postgres.

The post claims "annualized cloud savings north of USD 1.2 million." The schedule was a full quarter of proof work and about 1.5 quarters of route migration, and the target moved from July 30 to the first week of September.

That's a nice résumé bullet. Now read it the way a skeptical interviewer will, because they'll poke at exactly these spots.

Start with the target, which is a hybrid. Real-time data now lives in a managed AlloyDB deployment behind NGQF, and StarRocks-on-Iceberg serves batch. If you summarize this as "they replaced Postgres with a lakehouse," you've told an interviewer you skim. Every migration I've ever worked on ended up as a hybrid in some way. Something always stays behind, and someone always maintains it forever.

Then there's the $1.2M. The post says "a GCP contract renegotiation tied to a committed-use discount (CUD) purchase locked the cost savings to a calendar date." So part of that number comes from contract terms, and the post doesn't split it out or explain how it was calculated. That's a fine choice for an engineering blog. In an interview, an unbacked dollar figure is where the follow-up questions start.

The scale numbers also depend on what's being measured. A companion post on partitioning puts the table at over 500 billion rows and more than 42 TB of compressed Parquet, on sharded Postgres that cost over $100,000 a month in SSD storage alone. 400 TB and 42 TB can both be accurate; they describe different things. If you say "400 TB" in an interview, be ready to explain what that 400 measures.

If you can't explain where your migration's headline number came from, the interviewer will assume you didn't calculate it.

I learned this one the hard way. In one loop I said a migration "cut costs by about half," and the interviewer asked half of what, whether compute, storage or licensing, and over what period. I had no answer and didn't get the offer, which was fair enough.

TRM spent a full quarter proving Iceberg before moving a single route

This section is the one to study. Before they moved a single route, TRM's platform team "spent a full quarter de-risking the execution" and "ran 9+ experiments answering specific questions about Iceberg's behavior under TRM's query patterns." The post names the question areas:

  1. Could the joins the APIs actually run hit single-digit-second latency?
  2. How do partitioning and bucket sizing affect tail latency?
  3. How much throughput does compaction cost?
  4. Does it hold up at customer-facing concurrency?

Look at how specific those are. They test this workload on this engine at this concurrency, which is exactly the shape a pipeline-architecture interview answer should have. Skip the generic system design playbook; DEs don't care about load balancers. When someone asks you to design a migration, the strongest move is to say: "Before I commit, here are the 4 questions that would kill this plan, and here's the experiment for each one."

The best part of the proof phase is what it found. Andrew Fisher's companion post, Why We Stopped Partitioning by Time and What We Did Instead, says a daily-partitioned Iceberg table "broke at every point on our distribution." For the heaviest address keys, the daily layout touched about 25,000 files. Bucketing by address got that down to about 300.

The cause is skew. Around 97% of addresses hold 8% of the volume, and about 10,000 heavy-tail addresses hold 42%. Partitioning event data by time is everyone's default instinct, mine included. It falls apart when your read pattern is "give me everything for this one address," and a handful of addresses are monsters. Fisher's conclusion, paraphrased, is that file fan-out dominated read performance and partition granularity mattered much less.

The fix came from data modeling: match the physical layout to the access pattern, and plan for the skew that breaks averages. Iceberg happened to be the table format TRM applied it in. The same reasoning works in Delta, Hudi, BigQuery clustering, or whatever we're all migrating to in 2029.

The post-cutover numbers from the Fisher post: 7 days of production traffic, 280,000+ paginated queries, P50 around 1.0s and P95 around 2.5s, against a 3.5s P95 target. That's a clean result, measured on real traffic, against a target set ahead of time. Notice the structure, because you'll copy it in a minute.

The slip, the regressions, and your behavioral interview

Now the part most companies would edit out.

Start with the slip. "The original target was July 30; but that timeline shifted to the first week of September." In the middle of the project, the real-time layer behind NGQF moved from Postgres to managed AlloyDB, and that work landed on the same 30+ routes. Plaha writes that "the conventional answer was migrate routes twice," and argues against it: doubled customer-facing risk, real-time data pushed out by months, doubled engineering cost. So they migrated each route once and took the later date.

The post reports the AlloyDB decision and the date change together, and the author judged that absorbing the scope beat doing 2 passes. The case for that judgment is the argument Plaha lays out, since the post compares the 2 paths in reasoning and gives numbers for neither. In an interview, present it as a judgment call made under uncertainty, because that's what it was. Interviewers will trust that more than an invented counterfactual.

Then this line: "We re-baselined visibly with our CTO." I've watched people hide a slip until the deadline passed, and I've done it myself early in my career, and it never goes well. Plaha puts it better than I can: "Early bad news is actionable; late bad news is only reportable." Put that sentence in your behavioral prep and keep it there.

Then the regressions. "Our pre-migration benchmarks did not always predict real-world load and query shapes." And: "Latency regressions on specific routes reached customers before we caught them." A handful of routes also saw brief errors during cutover.

So a full quarter of careful proof still missed production query shapes. I don't count that against the team. A proof phase answers the questions you thought to ask, and production traffic asks the rest. What matters is what came next: "We have since invested in synthetic load testing and shadow-traffic tooling," plus staged rollouts behind feature flags and "automatic rollback triggers as standard practice." On top of that, "the engineers running the migration had the authority to call a rollback without escalating."

Compare that to Stripe's write-up on online migrations, where they used GitHub's Scientist to read from both the old and new tables and got alerted "as soon as a single piece of data was inconsistent in production." Stripe was checking correctness on production reads before switching. TRM's problem was latency on production routes. Those are different failure modes, and a strong migration answer accounts for both. If an interviewer asks "how would you validate this cutover?", the answer they're hoping for is dual reads plus shadow traffic at real concurrency, with rollback criteria written down before day 1.

Here's how I'd turn all of this into a behavioral answer from your own work. Don't claim TRM's story as yours. Cite it as "a public TRM write-up" if it helps you explain how you think, but your behavioral examples have to come from your own career.

  • Open with context in 2 sentences covering what moved, how big, and why. Have a number for scale and be able to say what it measures.
  • Walk through the proof, meaning the 3 or 4 questions that could have killed the plan and how you tested each one.
  • Own the miss, the thing that broke anyway, and your part in it. Interviewers can tell when you're blaming the upstream team without saying so.
  • Explain the re-baseline: who you told, when, and what changed.
  • End on the process change, the specific thing you do differently now. TRM's is shadow traffic and automatic rollback triggers. Yours needs to be just as concrete.

I've sat on panels where a candidate described a perfect migration with zero problems, and every interviewer in the room wrote "not credible" in their feedback. I've also seen candidates get strong hires for a story that included a customer-facing incident, because they could explain exactly why it happened and what they built so it couldn't happen again. Plaha has a line for the opposite pattern: "Re-running a broken approach with cosmetic tweaks is how most migrations stall." That applies to interview prep too.

How a data engineer should practice the migration question

Every migration interview question I've been asked boils down to 3 skills: modeling the data for the new access pattern, designing experiments that could prove you wrong, and communicating when the plan changes. The tools in the story barely matter. Swap StarRocks for Snowflake and AlloyDB for Aurora and the answer stays the same.

So here's the study plan.

Start with the modeling. Take a skewed dataset and explain out loud why time partitioning fails on a per-entity read pattern, and what you'd do about it. If you can't explain the 25,000-files-versus-300 result from first principles, start there. When I want structured practice on that side of the loop, data modeling interview prep is what datadriven.io is good for, and the bucketing-versus-partitioning tradeoff is exactly the kind of problem worth drilling until the explanation comes out cleanly.

For proof design, pick any migration you've lived through and write the 4 questions you wish you'd tested first, then the experiment for each one. That becomes your system design answer.

Know every figure in your story: the baseline, the method, and what's included. If part of your savings came from a contract renegotiation, say so before the interviewer asks.

Last, rehearse the bad news. Tell your slip story out loud. Who did you tell? How early? If the honest answer is "late," say that and explain what you do now. Plenty of people have a slip in their history; the ones who get hired can explain theirs without flinching.

I've been through enough platform migrations to know the lakehouse won't be the last one any of us run. The engine will change, the table format will change, and the vendor will rename everything twice. Skew, unpredictable production query shapes, and deadlines tied to a finance contract will still be there in 10 years.

So, the question for the comments: what's the worst thing your proof phase failed to catch, and did it come up in your next interview loop?

Top comments (0)