This is part 1 of a series of articles which breaks down my thoughts and conclusions which I decided to call The Black Box Fallacy.
The Black Box Fallacy
Imagine you have a black box in front of you. Now let’s understand the rules in this situation:
- You cannot open or look inside of it, and you don’t know how it works.
- The only things you have access to are an input and an output.
- The box accepts paper strips with text as an input and outputs a paper strip with new text on it.
You have the perception that this box can help your business grow and be profitable, but to get to this, you need to check how this box works.
With that in mind, you decide to input 10 different strips to it and verify the 10 respective outputs. You get a sense that you are starting to understand how this box works and that it could really improve your business.
But then one engineer says:
“We need more evaluations. We can’t trust this box yet.”
So, you decide to really scale your evaluations, now jumping to an outstanding 10,000 strips.
You verify each respective output and now feel more confident. You want to start offering this box to customers, but another engineer says:
“10,000 evals is a fair number. But we still don’t know how the box works. We need more evaluations!”
In order to feel safe, you decide to follow the engineer’s suggestion and now, with a lot of effort, run 1M evaluations. You now seem to know so much about the black box that you can even rebuild it and run it on the same evaluations.
And that’s what you do. One engineer still doesn’t trust it but feels comfortable enough to approve the deployment, and in order to scale your business, you clone the behavior of this box and deploy it across your organization so you can finally profit.
The “banana” problem
Time passes and after 3 months of running your business, someone randomly adds one unexpected input to the box: “banana”...
The system collapses.
All the engineers get together trying to understand what happened, they check the logs and see the strip there.
Rapidly, the team acts and patches the input so it doesn’t accept this string so you can get back to running your business without more disruptions.
Now that everything is working, the team decides that 1M evaluations are not enough to trust the box. The conclusion is clear:
“Let’s scale the evaluations to 10M.”
After some time running the new evaluations, you then find another string that also collapses the system. In one of the evals, the text “apple” also randomly breaks the system.
You feel that the evaluations are successful and you decide to ship a new version of the box with a patch for both “banana” and “apple”.
The compromise
In this fallacy, one of the main questions we can arrive at is:
“Can you compromise?”
In a situation where we don’t know how a box works, when is it time to stop scaling evaluations before we can trust this product being shipped to customers? When can you trust that we won’t find more “banana” problems, without spending all your profits to make sure that you can trust the box 100%?
Let’s say that you have enough budget to run 1B evaluations... Can you trust it now? Can you really be sure you won’t find any new issues in the system? How much value or trust did you get from scaling the evals from 10M to 1B?
Understanding the black box problem in SWE
The situation described in this fallacy might not feel like it, but it has been a big part of how the software engineering process has functioned for many decades.
Trying to deploy a 1-line code change in a project with more than 1M lines of code is essentially a black box situation. You can be in control of your line change, but as a human, you can’t hold 1M lines of code in your mind at the same time. The only option you have is to rely on different levels of evaluation, not just in this project but also in the other projects it connects to.
In the same way, when adding open-source libraries to your project, almost 100% of the time, the team won’t have the time or budget to read all the code available in these libraries.
You install them, run the evals and... trust.
In both situations, the only option you and your team have is to accept the compromise.
You can spend 3 years scaling your evaluations and making sure that you cover all possible parts of the code before shipping it to production, but are you sure there will ever be an end to it? Should you spend 2 more years to make sure you can cover that 0.0001% chance of failure?
The life of an engineering team is made of these decisions every day. We either set a threshold where we accept the risk, or we get stuck and don’t allow the business to move forward.
With AI, the problem scales to an exponential level where now your 1M lines of code project behaves more and more like a Black Box. We can’t maintain control of all the changes coming to our project. The codebase is huge and AI is generating these changes all the time, reaching a point where you can’t really tell whether a change came from a coding agent or a human.
So what do you do?
- Choose to not allow AI-written code and hold back business growth?
- Read every single line of code produced by AI so you can feel in control, but create a new bottleneck in the business?
- Scale evaluations and safeguards across different levels to ensure you can trust this black box to the point where the risk is small enough to compromise?
Top comments (1)
Official Platform Update
Security protocols have been updated for all developer accounts.