You have been working on the orders service for a month. You know the code well. You know the tables, the columns, the foreign keys. You could draw the schema from memory.
Here is a question that will surprise you. Have you ever actually looked at the rows?
Most developers learn a system through its code and its schema, which describe what the data is supposed to be. The rows tell you what it really is. Those are often two different things, and the gap between them is where many of the bugs you will chase this year are hiding.
Ask your lead for read only access to whatever data you are allowed to see, ideally a masked copy, and give yourself an hour. Then ask boring questions.
How many rows are there, and how fast is that growing? Which columns are null far more often than the code assumes? What is the oldest record, and does it look different from the newest? What are the ten most common values in the status column, and are there any values the code has no branch for? Is there a customer with four thousand orders when everybody else has twelve?
You will find history nobody told you about. A status that was retired years ago and still sits on old orders. Phone numbers in three formats because of a migration in 2021. Test accounts somebody created in production and never removed. Every one of these explains something in the code that looked strange, a defensive check or an odd default you were tempted to delete.
Write down what surprised you. Share the most interesting two or three with your team. Somebody will usually say oh, that is from the old importer, and you will learn a story that lives in nobody's documentation.
Two cautions. Respect the access you were given and never copy customer data somewhere it should not be. And keep your queries light, because nobody wants to explain why a curious new joiner slowed down the reporting database.
Make it a habit whenever you join a new area. Code tells you intentions. Data tells you what happened.
– Asael Shinder
Top comments (0)