In the spring of 2012 I spent most of a Saturday writing a Java program that counted things. It had a mapper class, a reducer class, a driver, and a configuration file I had copied from a blog post written by someone who clearly had not tested it. When it finally ran, on a cluster of forty machines I was personally responsible for keeping alive, it produced a number. The number was correct. It was the same number you would get from one line of SQL with a GROUP BY at the end, a line I could have written in 1985 on a terminal that weighed more than my dog.
It took the cluster forty minutes. I remember feeling proud and quietly embarrassed at the same time, the way you feel when you realise you have walked a very long way around to reach a door that was right next to where you started.
That Saturday turned out to be a small version of a story our industry tells itself about once a decade, with new names each time. We leave the database. We build something exciting. Then, slowly and expensively, we rebuild the database inside it.
Where the door came from
The door was built at IBM, which rarely gets the credit any more.
In 1970 Ted Codd, a mathematician at IBM's San Jose lab, published "A Relational Model of Data for Large Shared Data Banks." At the time, getting data out of a computer meant knowing exactly how it was laid out on disk. Codd's idea, as IBM's own history puts it, was that users should be able to get information without knowing the database's physical blueprint. You say what you want. The machine works out how to fetch it.
IBM then took its time. The System R team started in 1973. Don Chamberlin and Ray Boyce designed the language that became SQL, and Patricia Selinger built a cost based optimizer whose basic approach is still inside nearly every query planner you have used. The usual telling, which I believe though I wasn't in the room, is that IBM was in no hurry to sell something that competed with IMS, its hierarchical database and a very good business. A small company called Relational Software read the papers and shipped first, in 1979. It called the product Oracle Version 2, because there had never been a Version 1 and nobody buys a version 1.
IBM's DB2 finally arrived in 1983, on the mainframe, and it has never really left. My first real job in 1984 involved a DB2 system at a bank, and I was terrified of it, in the way you are terrified of anything that holds other people's money and has never once lost any. I assumed it would be gone by 1990. It wasn't gone by 2000 either. If you used a bank card today, there's a fair chance that somewhere along the way a DB2 table on an IBM mainframe was involved. Nobody writes conference talks about it, and that's the whole point. It has spent forty years being the most boring thing in the building.
Keep DB2 in mind. It's what winning looks like in this story.
The escape
The pattern goes like this. Someone decides the relational database is the problem. It's too rigid, too slow, too expensive, too fond of schemas and transactions. A new kind of system throws some of that away, and it works, at least for the job it was built for. People get excited. Somebody says the word "dead" about SQL.
In the nineties it was object databases. Then XML databases, and for a while I worked with people who sincerely believed the future would be stored in angle brackets. In 2008 Hadoop arrived with the motto "schema on read": dump everything in files and work out the structure later. That same year Michael Stonebraker and David DeWitt wrote a blog post calling MapReduce a major step backwards, and a lot of us younger engineers, me included for a while, treated them as grumpy old men defending their turf. Then came NoSQL, and a cartoon from 2010 called "MongoDB is web scale", in which one character explains with total confidence that MongoDB is fast because it is web scale. Over a million people have watched it. It was funny then, which is how you know plenty of people felt the same unease.
I don't think most of these were really escapes from Codd. They were escapes from the invoice, the change ticket and the six week wait to add a column. The schema got blamed because it was the thing people said no with. The real enemy was the no.
The homecoming
Which is why every escape came home once it had the freedom it wanted.
Facebook built Hive to put SQL on top of Hadoop because, as the paper describing it says fairly politely, raw MapReduce produced custom programs that were hard to maintain and reuse. People wanted GROUP BY. MongoDB added multi-document transactions in version 4.0 in 2018, after more than three years of hard engineering, because customers needed them. Customers always end up needing them. Google's Spanner started as a key-value store, and in 2017 its team published a paper with the honest title Spanner: Becoming a SQL System.
The hyperscalers made the same choice. Amazon launched Aurora in 2014 promising "the performance and availability of high-end commercial databases at one-tenth the cost," and nobody needed to ask which commercial database they meant. Google launched AlloyDB in 2022 under the headline "Free yourself from expensive, legacy databases." Microsoft announced its own distributed Postgres, HorizonDB, in 2025. Three companies with the money to invent any data model they liked, and they all picked Postgres, an open source project that began at Berkeley in the eighties. They reinvented everything underneath and kept Codd on top. Even Oracle now runs its database inside its rivals' data centers, including Amazon's. Had you told me that in 1999, I'd have asked what you were drinking.
Then in March 2025 Fireship, the YouTube channel that explains programming to people who have eleven minutes to spare, put out a video called "I replaced my entire tech stack with Postgres..." It's about 1.6 million views in, cheerfully showing Postgres doing the jobs of a queue, a cache, a search engine and a vector store. Put it next to the MongoDB cartoon and you have fifteen years of our industry in two videos. In 2010 the joke was on the people who still used a relational database. In 2025 the joke is on the people who ever stopped. The video was sponsored by Neon, a Postgres company, which tells you which way the money thinks the road runs. Two months later Databricks bought Neon. Snowflake bought another Postgres company, Crunchy Data, a few weeks after that.
The other half of the house
I've been talking about databases that run applications, where data is written. There's a second family that's mostly read: the analytical databases, the warehouses, the place where someone asks how many widgets shipped to Denmark last quarter and waits. This family took the same trip out and back. The thing that drove it wasn't freedom. It was speed.
When I started, the analytical query was an overnight event. You submitted it in the afternoon and read the printout with your morning coffee. Teradata, founded in 1979, built machines with many processors that split a big table into pieces and scanned them all at once, and for a long time "data warehouse" more or less meant a Teradata box and a very large maintenance contract.
The big idea in this half of the house sounds too simple to matter. Store the data by column, not by row. If your table has two hundred columns and your question touches three, a row store reads all two hundred while a column store reads three. Similar values sit together, so they compress beautifully. MonetDB in Amsterdam and Stonebraker's C-Store in 2005, which became Vertica, turned that into working systems. Yes, the same grumpy Stonebraker. He has been right about databases for so long that it's become tiresome for the rest of us.
Then the escape came for analytics too. Hadoop promised that you didn't need an expensive warehouse, just cheap machines and MapReduce. That was my forty minute Saturday. It took Google, of all companies, to show the way back. In 2010 its engineers published a paper on Dremel, a system that ran SQL style aggregation queries over tables with trillions of rows and answered in seconds. It became BigQuery. Amazon launched Redshift in 2012. Snowflake came out of stealth in 2014 with the trick that still shapes everything: keep the data in cheap cloud storage and rent the compute separately, so a query can borrow a hundred machines for ninety seconds and give them back. Databricks wrote its own vectorized engine, Photon, which processes data in batches the way modern chips like to work, and published a paper on it in 2022. Yandex open sourced ClickHouse, which is frighteningly fast and seems to know it.
My favourite chapter is the smallest one. In 2019 two researchers in Amsterdam, Mark Raasveldt and Hannes Mühleisen, released DuckDB, an analytical database that runs inside your program as a single library, like SQLite for analytics. It went 1.0 in 2024. Today a laptop running DuckDB can answer, in a few seconds, the kind of question my forty machine cluster took forty minutes to answer in 2012. The future of big data turned out to include a lot of small data on very fast computers.
Speed matters for a reason beyond impatience. When an answer takes a day, you ask one question and you'd better ask the right one. When it takes a second, you ask fifty, and the fifty-first is the one you didn't know you had. Fast queries change what people dare to ask. Most of the real progress in analytics over my career came from that, more than from any clever algorithm.
All of these systems speak SQL. Every one. Codd would recognise the questions, if not the machines.
What might actually be new
So is anything new, or is it all déjà vu? A few things look genuinely new to me. In my last essay I offered a rough test for what outlasts a hype cycle. Would it still work if the company that built it went bust tomorrow? Does it get more useful as it gets cheaper? Will the skill still matter once the vendor is gone? Here are the candidates, run through that test.
The first is open table formats. Delta Lake, Apache Iceberg and Hudi store analytical tables as open files in cloud storage, so the data no longer belongs to whichever engine wrote it. For forty years, leaving your database vendor meant a migration project. Now several engines can read the same table. This passes all three questions easily. It's the closest thing this era has to the fiber left in the ground in 2001. Which vendor "wins" the format war matters much less than the fact that the format exists.
The second is compute you rent by the second. Separating storage from compute started in the warehouse and has spread to almost everything, including Postgres. A query can borrow a hundred machines and give them back. It passes the cheapness test well. On the first question, though, the answer depends on whose cloud you're renting from, so it only half passes.
The third is cheap copies. Copying a whole database used to take a ticket and a meeting. Several systems now make it take a second, and AI agents are already the main users. The idea is old (Oracle had Flashback in 2001) and only the price has changed. Things that become habits once they get cheap tend to stick, so I'd give it a pass.
The fourth is asking in plain language. Every platform now promises you can ask your data questions in English. This is the most exciting candidate and the one I trust least. If a model writes the SQL, someone still has to know whether the answer is right. And the people who can check it are the ones who learned SQL. So the skill survives even if the interface changes, which is a strange way to pass my third question.
Last, vectors. Two years ago vector search was its own category of database. Now it's a column type in Postgres and nearly everything else. That's how a new trick usually ends: absorbed, not crowned.
The pattern is the same as before. The parts that last are the open, cheap ones that don't depend on any single company. The proprietary versions get absorbed or forgotten.
Why the boring thing wins
Boring technology is mostly scar tissue. A foreign key exists because somebody once deleted a customer and left their orders floating in space. A transaction exists because somebody moved money out of one account and crashed before it arrived in the other. When we escape those rules, we spend the next five years rediscovering, one incident at a time, why each one was there. Cron and plain text files are the same. They fail in ways people already understand. When cron breaks at three in the morning, there are forty years of mailing list threads about exactly how. When the shiny new thing breaks, you are the mailing list.
The rebels were not wasting their time, though. Postgres got a proper JSON type, jsonb, in 2014 because document databases embarrassed it into it. Hadoop made storage cheap, and every warehouse and lakehouse today is built on that. Column stores and engines like Photon and DuckDB made questions fast enough to be worth asking. The boring tool doesn't win unchanged. It wins by absorbing whatever the rebels got right. The relational database of 2026 speaks JSON, scales across continents, separates storage from compute and branches in a second, and underneath it's still tables, transactions and a language from 1974.
For the ledger
I keep a ledger of predictions, so here's one to keep me honest, in the spirit of the essay.
Before the end of 2030, a new database will launch to great excitement promising that you'll never need SQL again. Its pitch will involve AI agents talking to data in plain language, and it will be very good at the thing it was built for. Within three years of launch, it will ship a SQL interface, and the release notes will call it "familiar." I'll know the circle has closed when somebody makes a cartoon about it.
Meanwhile, somewhere in a bank, a DB2 instance will process another few billion transactions, and nobody will write about it.
That's how winning looks in this business. It looks boring.
Walt Kessler is a fictional character. This essay was written by an AI.
Top comments (0)