Right now, while a lot of software engineering is moving up the abstraction layer (with LLMs writing routine product code), the physical and architectural realities of where and how that code runs have never been more critical.
My Backstory — The Product Era (The Abstraction)
I started where most people do — focusing on user value, feature delivery, and rapid iteration. But eventually, the magic wears off when you realize you’re constantly building on top of systems you don’t fully control or understand. Everything is abstracted under the hood — application frameworks, libraries, methods, databases, memory management, etc. It kind of felt repetitive to me just calling functions/utils/methods to process data for input/output and return it back to customer.
Especially after AI era, it didn’t feel like I’m actually solving any hardcore engineering problem, I sort of became more like a business/product person rather than calling myself an “engineer”. There was no engineering left, it was just prompting, gathering requirements, doing CRUD, reviewing, testing, merging code and that’s all. And all this because we programmers were already working at the most abstracted part of a software — very far from hardware.
AI made me realize, the job of a programmer was essentially not the hardest part in the past let’s say 10 years because everything was already solved and abstracted. Our job was to call functions/methods, wait for data to get fetched, perform some math on it and return it back. There was no more thinking left to dig down hardware and question why and how to optimize it better. Everything felt repetitive to me, same APIs, same business requirements, same code, same principles, microservices patterns, I became more like a machine programmed to write code rather than a human to do some critical thinking and being curious.
And to be honest, this is exactly what AI is so good at — writing application code, but miserable at building large-scale distributed systems that requires combination of human + systems thinking.
The Revenge of the Hardware
In the early days of SaaS, the mantra was “hardware is a commodity; code is king.” Cloud abstractions made us forget about bytes, disks, and network packets.
The twist is AI and massive scale changed that. When you are provisioning clusters, managing GPUs, or optimizing Kubernetes nodes for high-throughput workloads, you suddenly care deeply about memory bandwidth, network latency, and compute density. Every single line of code matters alot.
Platform engineering demands system thinking, for e.g:
- How does a failure in the database pool ripple through the ingress controller?
- How do we scale a cluster dynamically without cascading timeouts of millions of real-time users?
- If we suddenly spin up 20 more instances, can the underlying database handle 20x more concurrent connection pools?
- Will our NAT gateway handle the massive surge in outbound traffic?
The AI Era
AI boom shattered the “infrastructure is invisible” paradigm.
Today running large-scale systems and cloud infrastructure today requires a deep understanding of low-level constraints. If you don’t respect the machine, the scale will break you.
For the last decade, engineering glory was found at the top of the stack. It was about building slick UIs, optimizing user conversion funnels, and shipping features at lightning speed. But the AI era has completely inverted this dynamic.
The Death of Routine Product Code (Upstairs)
At the top of the stack, abstractions have become so high that code is becoming heavily commoditized.
- With LLMs, generating a React component, scaffolding a CRUD API, or writing glue code for a product feature is fast, cheap, and increasingly automated.
- The “problems” at the top are becoming less about deep technical complexity and more about product design, prompt engineering, and stitching together pre-existing services. It’s a space where complexity is being managed away.
The Explosion of Physical Reality (Downstairs)
Meanwhile, at the bottom of the stack, the problems have become intensely complex, fascinating, and deeply tied to physical constraints. You can’t “prompt engineer” your way out of a networking bottleneck, a noisy neighbor saturating the CPU cache, or a memory leak under massive concurrent loads.
When a product app breaks, you check the logs and fix a null pointer. When a platform breaks at scale, you might be dealing with Linux kernel OOM (Out Of Memory) killers renegading through your pods, network packet drops at the NAT gateway, or storage IOPS throttling. It requires genuine detective work.
It’s Now More Fun
AI models, LLMs, and massive data pipelines are the ultimate test of systems thinking and I’m really enjoying it learning, studying, building and practicing around.
- How fast can data move from Cloud Storage to GPU memory (VRAM)?
- Are we bottlenecked by PCIe bandwidth?
- How do we orchestrate distributed training across multiple nodes without the network switches becoming a massive choking point?
Training or running inference on large models requires moving terrifying amounts of data. The bottle-neck isn’t how fast the code executes; it’s how fast data can cross the PCIe bus or the network switch.
How do you dynamically provision Kubernetes nodes with massive GPU attachments right when a burst of heavy compute hits?
I didn’t move to platform engineering to escape the AI revolution; I moved because that was always my core foundation — love for hardware and code, it’s a beautiful blend of systems, physics and humans.

Top comments (0)