The premise was that AI would change the world. It would supercharge business processes, alter industrial mechanics, create entirely new industries, and lead us toward a world of abundant production. Elon Musk has even suggested that money and work might one day become unnecessary, leaving only abundance. It was a bold vision, and it was presented with great confidence.
To a large extent, the promise held up in the GenAI era. We created millions of near-perfect photos and experimented with video creation. ChatGPT answered almost everything we asked, and Claude handled deep workflows that used to take teams days. Cursor became a favorite among coders who had spent their careers writing syntax by hand. This naturally led to rapid commercialization, and the valuations of frontier models rose to extraordinary levels, helped by the American talent for marketing ideas on a grand scale. Given Silicon Valley’s track record of changing the world we live in, many of us were convinced that the world was changing once again, and the fear of missing out spread across every industry, from banks to retailers to manufacturers. Nobody wanted to be left behind.
Then the conversation began to shift toward reliability and determinism. The arrival of open-weight and open-source models challenged the notion, strongly promoted by the frontier LLM companies, that hundreds of billions of dollars were needed to make LLMs near perfect. If a capable model could be downloaded, fine-tuned, and run on your own infrastructure, the competitive moat looked smaller than advertised. So a new race began: the deployment of AI, again led by Silicon Valley through influential gatekeepers like Y Combinator, which shape what gets funded, what gets attention, and which story the market hears next.
Today, two kinds of companies are emerging globally. The first are those deploying AI across almost every industry vertical, working to keep the promise of AI alive. The second are those taking a more scientific approach, trying to make AI more deterministic through guardrails, observability layers, and similar tools. Both are responding to the same challenge: the technology was presented faster than it could be adopted.
We are now at a crossroads. We are trying to work against the very science of probability that gave rise to LLMs and, with them, the whole AI phenomenon. These systems are powerful precisely because they work with probabilities rather than fixed rules. Trying to replace that probability with deterministic outcomes is extremely difficult, if not impossible. It is a bit like asking the ocean to behave like a swimming pool.
Some may say that I am being pessimistic or contrarian, or that I am going against the natural evolution of AI. I would respectfully disagree. First, I am a founder, not a researcher whose role is to offer opposing views. I build, I sell, and I sit across the table from people who have to make these systems work in real businesses. Second, I am pragmatic and fully aligned with the vision that AI can change the world we live in, though perhaps not in the way it is often promised. I believe AI is as significant as the Internet, if not more so. Like the Internet, it will empower us, bring more human creativity, push us to rethink how we work, and help us create more powerful processes. The Internet did not remove humans from the loop; it gave them more leverage. The idea that we can simply sit back and let everything happen autonomously, at scale, is not a realistic expectation.
Before looking at both kinds of efforts, it is worth pausing on a recent development that I read as a shift in messaging from Silicon Valley.
Dario Amodei is preparing for an IPO of a company that is still loss-making, at a valuation of around 2 trillion. There is nothing wrong with that, and I do not question it morally, but it comes with its own dynamics. When a company goes public, its story has to hold up under the scrutiny of public markets, and that naturally changes how the story is told.
He recently wrote an article calling for a slowdown in frontier AI progress and describing it as a challenge for the human race. Musk and Altman have voiced agreement. Not long ago, these same leaders spoke to us about AGI, about a world where money may not be needed, about abundant production, and about the point of singularity. Now they are calling for a slowdown. Why? My view is that this is not about the capabilities of frontier models. It reflects the fact that adoption at scale has not yet followed. Capability is racing ahead, but enterprises are not absorbing it at the same speed, and the gap between what these models can do and what businesses can trust them to do is where the real challenge lies. That gap has led to the efforts of vertical AI companies and of those working to enforce determinism, so let us look at each of them.
Vertical AI Solutions / Application Layers
This is one way, and perhaps the most discussed, most promoted, and most funded way, to deploy AI into processes, industries, and businesses for commercial use. Every sector now has its own group of startups: AI for legal, AI for healthcare, AI for customer support, AI for finance. I personally believe this is the way forward, but there is a fundamental challenge: it sits in tension with the science of machine learning. The issue is not the process itself but the way it is being positioned.
Vertical AI companies in any given sector are promising deterministic outcomes, which runs directly against how machine learning works. I understand that outputs have to be quantified in order to sell them, since no enterprise buyer signs a contract based on “it usually works well.” But that quantification cannot be produced deterministically, because we are no longer in the API era, where 2+2 was always 4, no matter where it was calculated or how many times you asked. We are in the AI era, where AI relies on the most probable next token. The power it brings, and the scale of output it delivers, reach far beyond the API era, but the certainty being promised remains difficult to reconcile with the science it relies on.
Let us take an example. Suppose a vertical AI company promises 40% higher customer satisfaction if you deploy its solution. The question is: how can they guarantee that it will always remain 40%? Being cloud-native (as most of them are) and reliant on a model, how can they promise a deterministic output? Models can change over time through the data they accumulate, the tool calls they make, the errors that occur, the API calls, and much more. The underlying model can be updated, the context can change, the data seen in production can differ from the data used in testing, and a small change in any of these can shift behavior in ways nobody planned for. Predicting a deterministic output therefore runs against the logic of agent behavior drift.
I am not suggesting that 40% is an unfounded claim. It may well be accurate on the day it is measured. My point is only that it cannot be guaranteed, and this is where the challenge for enterprise adoption begins. A number that holds in a pilot may not hold six months later in production, and it may go unnoticed until the impact is felt.
Consider a bank. A credit score is deterministic and derived from multiple factors. The same profile cannot receive different scores each time, because a customer, a regulator, or a court can ask why one person received one score on Monday and another on Friday, and “the model responded differently” is not an acceptable answer. Furthermore, the EU AI Act, which is a much-needed step for the industry, requires compliance with clear rules on transparency, traceability, and human oversight. Systems that cannot explain their own behavior will find it difficult to meet those requirements. This is where the second type of company comes in.
Determinism-Enforcing Companies
These companies understand the science of probability well and are working to manage it with every technique available. They deserve credit for taking the problem seriously, although the results so far have been limited.
For example, some introduced RAG (retrieval-augmented generation). The idea was that if a context layer was provided, the AI’s output would stay within the boundaries of the data provided. In practice, this did not fully work. Even within a given set of data, the science of machine learning remains in play, and the LLM still selects the most probable next answer. The model can still misread a document, give more weight to one source over another, or blend retrieved content with what it already “knows.” The output remains a probability, just a better-informed one.
After that, companies introduced context layers, memory, and hardcoded guardrails. Each of these adds some safety, but each also adds complexity and new points of failure, and none has fully solved the problem so far. At least, I have not yet come across a company that claims full compliance with the EU AI Act. This may be one reason for slower adoption, and possibly for the delay by policymakers in fully implementing the Act. If the industry cannot yet demonstrate that it meets the rules, regulators face a difficult choice between enforcing standards that very few can meet and extending the deadlines.
So what is the solution?
As I mentioned earlier, I see myself as a pragmatic person, and I believe in solutions more than opinions. My conclusion remains the same: AI adoption is a necessity, but not autonomous adoption. I believe in empowering humans with AI, and my answer is to make processes auditable.
If we cannot make probability deterministic, we should be honest about it. Instead of trying to remove uncertainty from the model, we should make that uncertainty visible, measurable, and controllable. The premise on which we are building AI infrastructure at Zizka AI (ZizkaDB) is simple: making AI processes auditable, with human oversight, so that we keep the power AI offers while retaining a human check on it.
In practice, this means three things. First, humans decide what level of agent drift is tolerable, because the acceptable level of risk for a marketing assistant is very different from that for a credit decision. Second, we can understand why an agent does what it does, so its decisions are not a black box but something a person can inspect, question, and explain. Third, we can replay everything in production, so that when something goes wrong, or even when something goes right, we gain deep insight into how it happened and can learn from it.
Essentially, we are trying to reverse the trend of handing everything over to AI, and instead place the power of AI in human hands so that humans can decide and act. AI provides the capability, while humans provide judgment and accountability. I believe this is the version of the future that businesses, regulators, and ordinary people can genuinely trust.
If you are interested in knowing more about ZizkaDB, check our open-source repository here.
Top comments (0)