<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mir Arshad Ali Talpur</title>
    <description>The latest articles on DEV Community by Mir Arshad Ali Talpur (@mir_arshadalitalpur_1b3).</description>
    <link>https://dev.to/mir_arshadalitalpur_1b3</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4063668%2F5b31344f-7d58-40dc-881c-6ae41ec788c3.jpeg</url>
      <title>DEV Community: Mir Arshad Ali Talpur</title>
      <link>https://dev.to/mir_arshadalitalpur_1b3</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mir_arshadalitalpur_1b3"/>
    <language>en</language>
    <item>
      <title>How zizkadb works and making AI agents auditable – open-source</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:52:11 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/how-zizkadb-works-and-making-ai-agents-auditable-open-source-1h88</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/how-zizkadb-works-and-making-ai-agents-auditable-open-source-1h88</guid>
      <description>&lt;p&gt;ZizkaDB is primarily an open-source tool designed to make AI agents more auditable and reliable.&lt;/p&gt;

&lt;p&gt;I believe that in the agentic era, databases should do more, and that’s why we’re building ZizkaDB.&lt;/p&gt;

&lt;p&gt;In this video, I explain the process in detail using live agentic data, covering events, event types, sessions, causal lineage, session replay, activity tracking, agent behavior and drift management, reports, and suggestions.&lt;/p&gt;

&lt;p&gt;In short, it covers what ZizkaDB does as an agentic database.&lt;/p&gt;

&lt;p&gt;I’d love to hear your feedback, especially if you’ve used the database or watched the video. Your thoughts and suggestions would be greatly appreciated.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>zizkadb</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Behavioral Drift: Know When Your Agent Has Changed, Before Your Customers Do</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Wed, 30 Sep 2026 12:27:54 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/behavioral-drift-know-when-your-agent-has-changed-before-your-customers-do-7e6</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/behavioral-drift-know-when-your-agent-has-changed-before-your-customers-do-7e6</guid>
      <description>&lt;p&gt;This is the third article in the series where I go through each function ZizkaDB offers and explain what it means in real business terms. The first covered Causal Lineage, why an agent did something. The second covered Session Replay, what actually happened, in order. This one covers Behavioral Drift, whether your agent is still behaving the way it did last week, last month, or on the day you shipped it.&lt;/p&gt;

&lt;p&gt;The three are meant to work together, but they answer different questions at different moments. Drift tells you something changed. Replay lets you watch the session where it changed. Lineage tells you why that specific step happened the way it did. Most teams will meet drift first, because it is the one that works passively in the background and taps you on the shoulder before you go looking for a problem.&lt;/p&gt;

&lt;p&gt;What behavioral drift is&lt;/p&gt;

&lt;p&gt;Every agent generates a pattern over time, how often it calls tools versus responding directly, how sessions typically flow from one step to the next, how long sessions run, how often things error out. ZizkaDB continuously compares your agent's recent behavior against an established baseline built from its prior sessions, and scores how much that pattern has moved. When the shift crosses a threshold, you get flagged, with the specific events and transitions that moved the most, ranked by how much they changed.&lt;/p&gt;

&lt;p&gt;This matters because agents change behavior far more easily than traditional software does, and usually without anyone deciding to change it. A model provider updates a model behind an API you call. Someone tweaks a prompt to fix one thing and it changes ten others. A tool your agent depends on gets slower or starts failing silently, and the agent starts routing around it in ways nobody planned. None of these show up in a standard uptime or error rate dashboard, because the agent is still running, still responding, still technically healthy. It is just doing something different than it used to.&lt;/p&gt;

&lt;p&gt;Consider a concrete case. Your team updates a prompt to make the agent explain refund decisions more clearly. Two weeks later, support tickets tied to refunds start creeping up, but nobody connects it to the prompt change because the change looked unrelated and the error rate did not move. With drift detection, the shift would have shown up within days of the update: a specific transition, say the agent skipping a verification step it used to take before issuing a refund, moving several points out of its normal range. You would have seen the exact mechanism before the ticket volume told you something was wrong.&lt;/p&gt;

&lt;p&gt;Drift is not always bad news. Sometimes it means the agent got faster or more accurate after a change. But finding that out from a dashboard, on your terms, is very different from finding it out from a customer, on theirs.&lt;/p&gt;

&lt;p&gt;Where you see benefits in the first 30 days&lt;/p&gt;

&lt;p&gt;Bad releases get caught in days, not weeks. A prompt or model change that shifts behavior in a way that matters shows up as soon as enough recent sessions accumulate, instead of surfacing later as a slow climb in tickets or a drop in conversion nobody can immediately explain.&lt;/p&gt;

&lt;p&gt;Silent regressions get caught even when error rates look fine. An agent can have a flat error rate and normal latency while behaving fundamentally differently underneath, leaning on tools it did not used to need, or skipping steps it used to take. Drift is often the only signal that catches this, because nothing else is watching for it.&lt;/p&gt;

&lt;p&gt;Root cause hunting starts with a hypothesis instead of a blank page. Ranked event and transition deltas point directly at what changed, which turns a vague "something feels off" into a specific, testable starting point for an engineer.&lt;/p&gt;

&lt;p&gt;Release reviews become a five minute check instead of a guessing game. After shipping a prompt or model change, a quick look at the drift dashboard tells you whether behavior moved and by how much, before you move on to the next release.&lt;/p&gt;

&lt;p&gt;You get an early warning system for third party dependencies. When a model provider updates a model behind your API calls, that change is often invisible until it shows up in your agent's behavior. Drift detection is one of the few ways to catch this kind of change on your own timeline.&lt;/p&gt;

&lt;p&gt;Modeled numbers&lt;/p&gt;

&lt;p&gt;These are estimates built from assumptions, not measured customer results. Swap in your own figures and the math still holds.&lt;/p&gt;

&lt;p&gt;Outcome Assumption  Before  After&lt;br&gt;
Bad release detection   3 significant releases a month, 100 euro per engineering hour   Detected after 10 days average via ticket volume, 40 hours of cleanup and rollback work at 4,000 euro per incident  Detected within 2 days via drift alert, 6 hours of targeted fix at 600 euro per incident&lt;br&gt;
Support ticket volume from undetected drift 1 undetected regression per quarter, average 150 extra tickets at 40 euro per hour, 25 minutes each 2,500 euro in extra support cost per incident   Largely avoided by catching the shift before it reaches customers&lt;br&gt;
Engineering time on root cause hunting  5 behavior related investigations a month, 100 euro per hour    4 hours each without a starting hypothesis, 2,000 euro  1 hour each starting from ranked deltas, 500 euro&lt;/p&gt;

&lt;p&gt;Taking just the release detection and root cause lines, that is roughly 4,500 euro a month in avoided cleanup and investigation cost for a mid sized team shipping regularly. The support ticket line is harder to put a clean number on every month, since it depends on whether a regression happens at all, but a single undetected drift event that reaches customers before anyone notices tends to cost more than a year of the monitoring that would have caught it.&lt;/p&gt;

&lt;p&gt;The honest caveat&lt;/p&gt;

&lt;p&gt;I would not ask you to trust this table either. Turn on drift detection against your own release cadence for a month, and check whether it flags a real behavior change before your existing signals do. If it only ever agrees with what you already knew from tickets or dashboards, it is not adding much for your case yet.&lt;/p&gt;

&lt;p&gt;It is also worth being clear about what drift detection does not do. It tells you that behavior moved and roughly where, it does not tell you why, and it does not tell you whether the change is good or bad. That judgment is still yours, and pairing the drift score with health metrics like error rate and session length, and then with lineage for the specific sessions involved, is what turns a flag into an actual decision.&lt;/p&gt;

&lt;p&gt;Why it matters beyond savings&lt;/p&gt;

&lt;p&gt;Agents built on LLMs do not come with the guarantee traditional software does, that the same input reliably produces the same category of output over time. That guarantee has to be built on top, through monitoring, and most teams currently only find out behavior changed when a customer, a support queue, or a churn number tells them. Drift detection moves that discovery earlier and puts it in your hands instead of theirs. For a vertical AI company shipping frequently, that is the difference between debugging a regression on your own schedule and doing it under pressure with an angry customer on the line.&lt;/p&gt;

&lt;p&gt;If you tell me your vertical and how often you ship changes to your agent, I can tailor these numbers to your case.&lt;/p&gt;

&lt;p&gt;Want to test ZizkaDB on your agent? Try our open source version here: &lt;a href="https://github.com/ZIZKA-AI-SL/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/ZIZKA-AI-SL/ZizkaDB&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Interested in a design partnership? Fill in the form on our site or reach me directly at &lt;a href="mailto:founder@zizka.ai"&gt;founder@zizka.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agents</category>
      <category>agenticpostgreschallenge</category>
    </item>
    <item>
      <title>Session Replay: See Exactly What Your Agent Did, Step by Step</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Tue, 29 Sep 2026 09:34:40 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/session-replay-see-exactly-what-your-agent-did-step-by-step-4gf4</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/session-replay-see-exactly-what-your-agent-did-step-by-step-4gf4</guid>
      <description>&lt;p&gt;This is the second article in the series where I go through each function ZizkaDB offers and explain what it means in real business terms. The first covered Causal Lineage, why an agent did something. This one covers Session Replay, what actually happened, in order, from the first event to the last.&lt;/p&gt;

&lt;p&gt;The two are related but answer different questions. Lineage tells you why a specific action happened, tracing the inputs, context, and decision points behind one output. Replay lets you walk through an entire session as it unfolded, every user message, every decision, every tool call, every LLM response, in the exact sequence it occurred. One is a causal explanation for a single action, the other is a full playback of everything that happened around it. In practice, teams use replay to find the moment something went wrong, then use lineage to understand why that specific step happened the way it did. They are meant to be used together.&lt;/p&gt;

&lt;p&gt;What session replay is&lt;/p&gt;

&lt;p&gt;Every session your agent runs is logged as a complete event stream, not sampled, not summarized after the fact. Session Replay reconstructs that stream into a timeline you can step through, from the first user message to the final response, including every intermediate decision, tool call, and LLM round trip in between. You are not piecing together a story from scattered logs across five services, correlating timestamps by hand, or trusting a summary someone else wrote. You open one session and see the whole thing in order, exactly as it happened.&lt;/p&gt;

&lt;p&gt;This matters because most agent failures are not single bad actions, they are a sequence of small missteps that only make sense together. A tool that returned a stale result three steps before the final answer. A decision node that misread context that was itself slightly off. A retry that picked the wrong branch after an earlier step nudged it there. Looking at any one event in isolation rarely explains the failure, because the real cause is often two or three steps upstream of where the problem became visible. Watching the session unfold, in order, is usually the fastest way to spot that.&lt;/p&gt;

&lt;p&gt;Consider a concrete case. A customer complains that your agent gave them the wrong refund amount. Without replay, someone pulls application logs, then tool call logs, then LLM request logs, tries to line them up by timestamp, and guesses at the sequence. It usually works eventually, but it takes time and it is easy to misread the order when logs come from different systems with different clocks. With replay, you open that one session and watch it end to end. You see the customer's original message, the decision to check order history, the tool call that returned the order, the LLM call that calculated the refund, and the point where the number went wrong. The difference is not just speed, it is confidence that you are looking at what actually happened rather than a reconstruction of it.&lt;/p&gt;

&lt;p&gt;Where you see benefits in the first 30 days&lt;/p&gt;

&lt;p&gt;Support tickets close faster. When a customer reports a bad outcome, you no longer ask them to describe what happened or dig through separate systems to piece it together. You open the session and watch it yourself, in the order it actually ran.&lt;/p&gt;

&lt;p&gt;Onboarding new engineers gets easier. A new hire can watch ten real sessions and understand how the agent actually behaves in production, including its edge cases and failure modes, faster than reading documentation or asking teammates to explain from memory.&lt;/p&gt;

&lt;p&gt;QA catches issues before customers do. Reviewing a sample of sessions after a release becomes a routine, five minute habit instead of a manual log dig, so regressions get caught in the first day rather than surfacing as a wave of tickets a week later.&lt;/p&gt;

&lt;p&gt;Product and engineering stop arguing from memory. When a product manager asks why the agent did something odd, you both look at the same replay instead of two different guesses reconstructed from incomplete logs, which tends to shorten these conversations considerably.&lt;/p&gt;

&lt;p&gt;Training data and evaluation sets get better. Sessions that reveal edge cases or failure patterns can be pulled directly into your eval suite, since you have the full, ordered context rather than a fragment.&lt;/p&gt;

&lt;p&gt;Handoffs between shifts or teams get cleaner. If an on call engineer picks up an incident partway through, they can watch the full session instead of relying on a handoff note that may miss details.&lt;/p&gt;

&lt;p&gt;Modeled numbers&lt;/p&gt;

&lt;p&gt;These are estimates built from assumptions, not measured customer results. Swap in your own figures and the math still holds.&lt;/p&gt;

&lt;p&gt;Outcome Assumption  Before  After&lt;br&gt;
Support ticket resolution   300 tickets a month involving agent behavior, 40 euro per hour support cost 25 minutes each, 5,000 euro 5 minutes each, 1,000 euro&lt;br&gt;
New engineer ramp time  4 new hires a year, 80 euro per hour loaded cost    3 days to understand production behavior, 1,920 euro each   1 day, 640 euro each&lt;br&gt;
Post release QA 2 releases a month, 100 euro per hour engineering cost  6 hours manual log review, 1,200 euro   1 hour session sampling, 200 euro&lt;/p&gt;

&lt;p&gt;That is roughly 6,000 euro a month in saved effort for a mid sized team. The ramp time line is worth more than it looks on its own, since a team that can bring new engineers up to speed on real production behavior in a day rather than three days ships faster from day one, and that compounds every time you hire.&lt;/p&gt;

&lt;p&gt;There is a fourth category worth naming even without a clean number attached: fewer misdiagnosed incidents. When teams reconstruct sessions from fragmented logs, they sometimes fix the wrong thing because the reconstructed sequence was slightly wrong. Replay reduces that risk simply by removing the reconstruction step. I have not modeled this one because it is hard to estimate honestly without real data, but it is often the benefit engineering teams mention first once they have used it.&lt;/p&gt;

&lt;p&gt;The honest caveat&lt;/p&gt;

&lt;p&gt;I would not ask you to trust this table either. Run Session Replay against your own support queue for two weeks and time how long resolution actually takes with and without it, using the same tickets or a comparable sample. If it does not save real time in that window, it is not the right priority for you yet, and I would rather you find that out in two weeks than take my word for it.&lt;/p&gt;

&lt;p&gt;It is also worth being clear about what replay does not do. It will not tell you why a decision was made, that is what lineage is for. It will not fix a bad prompt or a flaky tool. What it gives you is ground truth about sequence and timing, which is the raw material every other kind of debugging depends on.&lt;/p&gt;

&lt;p&gt;Why it matters beyond savings&lt;/p&gt;

&lt;p&gt;The hardest part of running agents in production is not building them, it is trusting what they did when nobody was watching. Session Replay removes the guesswork from that trust gap. You do not have to reconstruct a session from fragments, correlate logs by hand, or take someone's word for what happened, you watch it, in order, exactly as it ran. For a vertical AI company, that turns "we think this is what went wrong" into "here is exactly what went wrong," and that difference shows up in how fast your team fixes things, how confident your customers are in the fix, and how quickly new engineers become productive on a system that is, by nature, harder to reason about than traditional software.&lt;/p&gt;

&lt;p&gt;If you tell me your vertical and how many sessions you run per month, I can tailor these numbers to your case.&lt;/p&gt;

&lt;p&gt;Want to test ZizkaDB on your agent? Try our open source version here: &lt;a href="https://github.com/ZIZKA-AI-SL/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/ZIZKA-AI-SL/ZizkaDB&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Interested in a design partnership? Fill in the form on our site or reach me directly at &lt;a href="mailto:founder@zizka.ai"&gt;founder@zizka.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>agents</category>
      <category>python</category>
    </item>
    <item>
      <title>Causal Lineage: What It Is and the Business Outcomes It Delivers</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Mon, 28 Sep 2026 08:07:05 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/causal-lineage-what-it-is-and-the-business-outcomes-it-delivers-4cbb</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/causal-lineage-what-it-is-and-the-business-outcomes-it-delivers-4cbb</guid>
      <description>&lt;p&gt;This is the first article in a series where I will go through each function ZizkaDB offers and explain what it means in real business terms. We start with Causal Lineage, a proprietary function of ZizkaDB.&lt;/p&gt;

&lt;p&gt;I have spoken directly with 50 developers building agentic systems, and most of them told me the same thing: failure is a post production problem. Right now they are trying to use AI wherever they can, and they will deal with trust and traceability once something breaks. It is a classic chicken and egg situation.&lt;/p&gt;

&lt;p&gt;My own view has not changed. Database selection and auditability should be settled before production, not after. But I cannot change the market, so I am adapting to it. I also understand the pressure vertical AI companies are under. If you run a multi agent system and were recently funded, your biggest priority is adoption and sales numbers. So instead of arguing with your immediate needs, this series starts from them.&lt;/p&gt;

&lt;p&gt;Here is the short version: when your agent does something a customer questions, you can answer why in minutes instead of days, and you can prove it. For a vertical AI company, that shows up in three places you already feel: engineering time, sales cycles, and customer trust.&lt;/p&gt;

&lt;p&gt;What causal lineage is&lt;br&gt;
Every agent action in ZizkaDB is linked to what caused it: the user input, the retrieved context, the tool results, and the decision points in between. One why() call returns that whole chain. You are no longer stitching together logs from five systems and guessing at the order.&lt;/p&gt;

&lt;p&gt;Where you see benefits in the first 30 days&lt;br&gt;
Debugging drops from hours to minutes. Most teams spend the bulk of an incident just reconstructing what happened. With lineage, that step mostly disappears.&lt;/p&gt;

&lt;p&gt;Customer disputes stop being arguments. When a customer says your agent got it wrong, you show the exact inputs and reasoning behind the action, instead of relying on memory or screenshots.&lt;/p&gt;

&lt;p&gt;Enterprise deals move faster. Buyers in finance, health, legal, and insurance ask how you audit agent decisions. A concrete answer, backed by a demo, shortens security review and removes a common reason for deals stalling in pilot.&lt;/p&gt;

&lt;p&gt;Prompt and model changes get safer. Drift detection shows you within days whether a release changed how your agent behaves, before customers report it.&lt;/p&gt;

&lt;p&gt;Modeled numbers&lt;br&gt;
These are estimates built from assumptions, not measured customer results. Swap in your own figures and the math still holds.&lt;/p&gt;

&lt;p&gt;Outcome&lt;/p&gt;

&lt;p&gt;Assumption&lt;/p&gt;

&lt;p&gt;Before&lt;/p&gt;

&lt;p&gt;After&lt;/p&gt;

&lt;p&gt;Incident investigation&lt;/p&gt;

&lt;p&gt;20 incidents a month, 100 euro per engineering hour&lt;/p&gt;

&lt;p&gt;7 hours each, 14,000 euro&lt;/p&gt;

&lt;p&gt;0.5 hours each, 1,000 euro&lt;/p&gt;

&lt;p&gt;Dispute handling&lt;/p&gt;

&lt;p&gt;200 disputes a month, 60 euro per hour&lt;/p&gt;

&lt;p&gt;45 minutes each, 9,000 euro&lt;/p&gt;

&lt;p&gt;15 minutes each, 3,000 euro&lt;/p&gt;

&lt;p&gt;Security review in sales&lt;/p&gt;

&lt;p&gt;10 enterprise deals a year, 3 weeks lost each on audit questions&lt;/p&gt;

&lt;p&gt;30 weeks of delay&lt;/p&gt;

&lt;p&gt;roughly 10 weeks&lt;/p&gt;

&lt;p&gt;That is around 19,000 euro a month in saved effort for a mid sized team, but the sales line matters more. Cutting even two weeks off each enterprise deal pulls revenue forward, and it can decide whether a deal closes at all.&lt;/p&gt;

&lt;p&gt;The honest caveat&lt;br&gt;
I would not ask you to trust my table. I would ask you to run a two week pilot on one production agent, measure your real hours per incident and per dispute before and after, and judge from your own data. If the reduction is not obvious in that window, lineage is not the right priority for you yet.&lt;/p&gt;

&lt;p&gt;Why it matters beyond savings&lt;br&gt;
Enterprise buyers are not holding back because your agent lacks capability. They hold back because they cannot explain its decisions to their own auditors. Lineage gives them that explanation, and for a vertical AI company selling into regulated industries, that is often the difference between a pilot and a contract.&lt;/p&gt;

&lt;p&gt;If you tell me your vertical and how many agent actions you run per month, I can tailor these numbers to your case.&lt;/p&gt;

&lt;p&gt;Want to test ZizkaDB on your agent? Try our open source version here: &lt;a href="https://github.com/ZIZKA-AI-SL/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/ZIZKA-AI-SL/ZizkaDB&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Interested in a design partnership? Fill in the form on our site or reach me directly at &lt;a href="mailto:founder@zizka.ai"&gt;founder@zizka.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>causal</category>
      <category>opensource</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>How I Reached My First 100 GitHub Stars (and Why It Was Harder Than Fundraising)</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Fri, 25 Sep 2026 10:43:11 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/how-i-reached-my-first-100-github-stars-and-why-it-was-harder-than-fundraising-5mb</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/how-i-reached-my-first-100-github-stars-and-why-it-was-harder-than-fundraising-5mb</guid>
      <description>&lt;p&gt;I’m usually the person who says everything in life is awesome. Honestly, it has been.&lt;/p&gt;

&lt;p&gt;I come from a small, poor village in Sindh, and I was probably the first person from it to go to university. I started in the insurance sector, then moved to Dubai and became an award-winning seller. After that came nine years in the startup world: one success, two grand failures, and four migrationsand ended Up at Station F Paris (World’s Largest Startup Campus). I don’t complain about any of it, and I rarely even talk about it, because all I have is an amazing journey.&lt;/p&gt;

&lt;p&gt;But today isn’t about me. It’s about my first 100 GitHub stars.&lt;/p&gt;

&lt;p&gt;The Product Behind the Stars&lt;br&gt;
I’m building ZizkaDB, an operational database that answers the why behind every agentic action. When an AI agent does something, most systems can tell you what happened. ZizkaDB is built to tell you why.&lt;/p&gt;

&lt;p&gt;That is a notoriously hard problem, and it needs a specialized stack. Finding people who understand that stack, and who have the patience to build it with me, was harder still.&lt;/p&gt;

&lt;p&gt;I got lucky. My founding engineer has seven years of experience at well-funded startups. My advisor is an Applied Mathematics PhD from Germany who has worked in the forward-deployed AI engineering world. Around the same time, I moved myself from Málaga to Station F in Paris. It was a lot of moving parts, but I managed it because I had done versions of it before.&lt;/p&gt;

&lt;p&gt;What I hadn’t done before was get developers to star a repository.&lt;/p&gt;

&lt;p&gt;May 25th: The Quiet Launch&lt;br&gt;
We released our first SDK on May 25th, and getting even a few stars was hard.&lt;/p&gt;

&lt;p&gt;Station F was my first surprise. It is the biggest startup campus on earth, and you’d expect that to mean instant community. In reality, the French working culture is different from what I was used to. People who sit next to each other can go months without knowing each other’s names. Saying hello for no reason can feel like interrupting someone’s work. I had also just moved to France, so I had no local network to lean on.&lt;/p&gt;

&lt;p&gt;So I did the only thing I could: I wrote.&lt;/p&gt;

&lt;p&gt;Phase 1: Writing, Writing, and More Writing&lt;br&gt;
I published almost 17 articles on Medium. I also posted on Hashnode and dev.to, and I shared articles on Hacker News.&lt;/p&gt;

&lt;p&gt;It worked, partially. Downloads climbed. We crossed 4,200 downloads and 295 active system installations. People were using the product.&lt;/p&gt;

&lt;p&gt;But the stars grew one at a time. Downloads and stars measure different things: a download says “I tried it,” and a star says “I’m willing to put my name next to it.” Getting from one to the other was slow.&lt;/p&gt;

&lt;p&gt;Phase 2: Leaning on Networks&lt;br&gt;
My founding engineer shared the project with the developer communities he knows, and that pushed us to around 80 stars.&lt;/p&gt;

&lt;p&gt;Then it stalled.&lt;/p&gt;

&lt;p&gt;I sent personal messages and posted in WhatsApp groups. I tried whatever I could think of. Nothing moved. I started to believe this was the hardest work I had ever done.&lt;/p&gt;

&lt;p&gt;That surprised me. I know fundraising is one of the hardest things in startups, and I’ve done it before. I know what it takes, and that process is moving at full pace for ZizkaDB. But collecting stars from developers was a different beast. In fundraising, I know the playbook. Here, I was learning it in public, one star at a time.&lt;/p&gt;

&lt;p&gt;Phase 3: Doing What I Know Best&lt;br&gt;
This morning I was at 84 stars. I decided to stop waiting and take charge.&lt;/p&gt;

&lt;p&gt;I walked to every table in Station F where founders were sitting. I used the sales skills I’ve built over the years and asked them, face to face, to open GitHub and look at my repository. I pitched ZizkaDB to at least 40 founders, and about 20 of them starred it.&lt;/p&gt;

&lt;p&gt;That got us to 100.&lt;/p&gt;

&lt;p&gt;The lesson is simple, and it’s one I already knew. Online reach helps, but nothing beats a human being looking someone in the eye and asking. Ironically, the culture that made Station F feel closed at first was easy to break through once I stopped waiting for people to notice me.&lt;/p&gt;

&lt;p&gt;“Isn’t 100 Stars Too Early to Celebrate?”&lt;br&gt;
Some people will say so, and they’re right. 100 stars is not a big deal in the grand scheme of open source.&lt;/p&gt;

&lt;p&gt;But I’ve always owned my failures and my successes, and this one feels great to me right now. I’m writing this for three reasons:&lt;/p&gt;

&lt;p&gt;For myself. When ZizkaDB reaches the top, and I believe it will, I want to remember every bit of what it took to get here.&lt;br&gt;
For founders who feel the first part is the hardest. It probably is. But it is worth it.&lt;br&gt;
For honesty. Nobody talks about the boring middle where you send messages, get no replies, and keep going anyway.&lt;br&gt;
The Full List: What I Did to Reach 100&lt;br&gt;
If you want the short version, here it is:&lt;/p&gt;

&lt;p&gt;Wrote detailed articles on Medium, dev.to, and Hashnode&lt;br&gt;
Shared articles on Hacker News&lt;br&gt;
Reached out to people personally&lt;br&gt;
Shared in WhatsApp communities&lt;br&gt;
Walked up to desks and pitched the product directly&lt;br&gt;
And, not to forget, sent 2,500 emails along the way&lt;br&gt;
None of these was a magic bullet. Each one added a few stars, and together they got me over the line.&lt;/p&gt;

&lt;p&gt;Final Thoughts&lt;br&gt;
People like things to come easily, and I understand that. But in my life, and in the lives of many others, things don’t happen quickly. You pay the price first.&lt;/p&gt;

&lt;p&gt;It’s worth it.&lt;/p&gt;

&lt;p&gt;If you’re a founder stuck at your own version of 84 stars, whether that’s users, revenue, or investors, keep going. And if it’s stars specifically, stand up, walk to the nearest table, and ask.&lt;/p&gt;

&lt;p&gt;I’m building ZizkaDB in public. If you’d like to see what an operational database for agentic “why” looks like, the repository is on GitHub, and a star would mean a lot.&lt;/p&gt;

&lt;p&gt;You can access our public repo here&amp;nbsp;: &lt;a href="https://github.com/ZIZKA-AI-SL/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/ZIZKA-AI-SL/ZizkaDB&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The article is originally published in Medium and can be viewed here: &lt;a href="https://medium.com/@MirArshadTalpur/how-i-reached-my-first-100-github-stars-and-why-it-was-harder-than-fundraising-753608ee43ec?postPublishedType=initial" rel="noopener noreferrer"&gt;https://medium.com/@MirArshadTalpur/how-i-reached-my-first-100-github-stars-and-why-it-was-harder-than-fundraising-753608ee43ec?postPublishedType=initial&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>github</category>
      <category>founder</category>
    </item>
    <item>
      <title>Time-Travel Debugging and Drift Measurement: How to Audit an AI Agent.</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Wed, 23 Sep 2026 12:45:47 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/time-travel-debugging-and-drift-measurement-how-to-audit-an-ai-agent-oc0</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/time-travel-debugging-and-drift-measurement-how-to-audit-an-ai-agent-oc0</guid>
      <description>&lt;p&gt;&lt;strong&gt;Time-travel debugging reconstructs what an AI agent knew at a given moment, and drift measurement flags when its behavior departs from its baseline. Together they turn agent auditing from guesswork into a repeatable process. This guide covers both, with an implementation using the open-source ZizkaDB.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;An auditor doesn't ask "is the agent working right now?" They ask "what did it know last Tuesday when it made this decision, and has it been behaving the same way since?" Those are two different questions, and most monitoring setups can only half-answer either of them. A dashboard shows current state. An error log shows what broke. Neither reconstructs a past moment, and neither tells you whether today's normal behavior is quietly becoming tomorrow's incident.&lt;/p&gt;

&lt;p&gt;Time-travel debugging and drift measurement are the two capabilities that close that gap. This post covers what each one does, why agents need both, and how to implement them with the open-source ZizkaDB.&lt;/p&gt;

&lt;p&gt;Two different questions, two different tools&lt;/p&gt;

&lt;p&gt;Time-travel debugging answers: what did the system know, and what state was it in, at a specific point in the past? It's a reconstruction of a single moment, on demand, after the fact.&lt;/p&gt;

&lt;p&gt;Drift measurement answers: has the system's behavior changed compared to how it used to behave? It's a comparison across many moments, watching for a trend rather than inspecting a point.&lt;/p&gt;

&lt;p&gt;They're complementary rather than overlapping. Drift measurement tells you that something changed and roughly when. Time-travel debugging lets you go to that moment and see what changed. An audit that only drifts-detects tells you there's a problem without letting you inspect it. An audit that only replays moments requires you to already know which moment to look at. Together, they form a loop: drift detection flags the window, time-travel debugging examines it.&lt;/p&gt;

&lt;p&gt;Why agents need both, and ordinary logging doesn't cover them&lt;/p&gt;

&lt;p&gt;The state problem. An agent's decision depends on more than its final output: what it retrieved, what tools returned, what was in its context window, which model and prompt version was live. A line in a log shows you the output. It doesn't reconstruct the state the agent was reasoning over. Time-travel debugging is built specifically to answer "what did the agent know at time T," not just "what did it say."&lt;/p&gt;

&lt;p&gt;The moving-target problem. Agents change without anyone touching their code. A model provider ships a silent update. A retrieval index gets new documents. Users discover new phrasing that pushes the agent into behavior nobody tested. None of these show up as a deployment event, so ordinary release-based monitoring misses them entirely. Drift measurement watches behavior itself, not deployments, which is the only way to catch this class of change.&lt;/p&gt;

&lt;p&gt;The audit problem. An audit, whether internal, a customer's due diligence, or a regulator's request, usually starts with a specific incident or a specific date, not with "please show me everything." You need to answer for a particular moment, with confidence that the record wasn't reconstructed from memory or guesswork after the fact. That is exactly what time-travel debugging is for, and exactly what ordinary logs, which sample activity rather than reconstruct state, tend to be too thin to support.&lt;/p&gt;

&lt;p&gt;What ZizkaDB provides&lt;/p&gt;

&lt;p&gt;ZizkaDB is an open-source, self-hosted audit trail database for AI agents built around exactly this pair of problems. Its README lists both functions plainly:&lt;/p&gt;

&lt;p&gt;Function    What it does&lt;br&gt;
db.at() Reconstruct what the agent knew at a timestamp&lt;br&gt;
db.baseline()   Detect when agent behavior drifts vs. past sessions&lt;/p&gt;

&lt;p&gt;Both sit on top of the same underlying record: every agent step logged as an event, linked to the step that caused it via parent_id. That causal chain is what makes db.at() a real reconstruction rather than a snapshot, and what makes db.baseline() a comparison against real history rather than a synthetic test set.&lt;/p&gt;

&lt;p&gt;Time-travel debugging with db.at()&lt;/p&gt;

&lt;p&gt;The basic pattern is to log events as the agent runs, the same way you would for any audit trail:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
import asyncio&lt;br&gt;
from zizkadb import ZizkaDB&lt;/p&gt;

&lt;p&gt;async def main():&lt;br&gt;
    async with ZizkaDB(host="&lt;a href="http://localhost:8000%22" rel="noopener noreferrer"&gt;http://localhost:8000"&lt;/a&gt;) as db:&lt;br&gt;
        user = await db.log(&lt;br&gt;
            agent="support-bot", event="user_message",&lt;br&gt;
            data={"text": "Why was my order delayed?", "order_ref": "ORD-8842"},&lt;br&gt;
        )&lt;br&gt;
        lookup = await db.log(&lt;br&gt;
            agent="support-bot", event="tool_call",&lt;br&gt;
            data={"tool": "lookup_order", "order_id": "ORD-8842"},&lt;br&gt;
            parent_id=user.event_id,&lt;br&gt;
        )&lt;br&gt;
        answer = await db.log(&lt;br&gt;
            agent="support-bot", event="llm_response",&lt;br&gt;
            data={"model": "gpt-4o", "text": "Your order ships tomorrow."},&lt;br&gt;
            parent_id=lookup.event_id,&lt;br&gt;
        )&lt;/p&gt;

&lt;p&gt;asyncio.run(main())&lt;/p&gt;

&lt;p&gt;When a complaint comes in three weeks later, db.at() lets you go back to that moment and see what the agent had in front of it, rather than relying on someone's memory of what the tool probably returned:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
state = await db.at(agent="support-bot", timestamp="2026-09-01T14:32:00Z")&lt;br&gt;
state.print()&lt;/p&gt;

&lt;p&gt;I haven't run this against a live instance, so treat the exact call signature as illustrative; check the current docs for the precise arguments db.at() accepts. The concept, however, is what matters for auditing: you're not trusting a summary written after the fact, you're replaying the actual recorded state.&lt;/p&gt;

&lt;p&gt;Drift measurement with db.baseline()&lt;/p&gt;

&lt;p&gt;Drift measurement needs a "normal" to compare against, which is what db.baseline() establishes from an agent's past sessions:&lt;/p&gt;

&lt;p&gt;python&lt;br&gt;
baseline = await db.baseline(agent="support-bot", window="30d")&lt;/p&gt;

&lt;p&gt;Once a baseline exists, ongoing sessions can be checked against it, and a widening gap is your signal to look closer, using db.why() to trace a specific bad step or db.at() to inspect the state at the moment things started to shift. As with db.at(), I don't have a captured example run to show you, so verify the exact method signature and output format in the repo's docs before using this in production.&lt;/p&gt;

&lt;p&gt;What counts as "drift" is worth being deliberate about. Useful signals include:&lt;/p&gt;

&lt;p&gt;Tool-usage shape. An agent suddenly calling a tool it rarely used, or skipping one it always used, is often the first visible sign that something upstream changed.&lt;br&gt;
Response length or structure. A support agent whose answers double in length, or whose refusal rate climbs, is behaving differently even if no single reply looks wrong.&lt;br&gt;
Escalation and human-override rate. If human reviewers are stepping in more often, that's drift measured through the outcome that matters most.&lt;br&gt;
Decision distribution. For an agent that classifies or recommends, watch whether the split between outcomes shifts over time, not just whether any single decision was correct.&lt;br&gt;
A worked scenario: catching drift before a complaint&lt;/p&gt;

&lt;p&gt;Say support-bot's baseline shows it resolves order-status questions using lookup_order in roughly 90% of sessions. Two weeks after a routine dependency update, db.baseline() shows that share has dropped to 60%, and human-review escalations have crept up.&lt;/p&gt;

&lt;p&gt;That's the signal, not yet the explanation. The investigation looks like this:&lt;/p&gt;

&lt;p&gt;Drift measurement flags the window. Something changed starting around the dependency update.&lt;br&gt;
Time-travel debugging inspects it. Use db.at() on a handful of sessions from just after the update to see what state the agent was reasoning over.&lt;br&gt;
Causal lineage explains the mechanism. Use db.why() on a specific bad answer to trace it back, the same technique covered in our causal lineage guide, which might show, for example, that a tool's response schema changed and the agent started silently falling back to a weaker answer path instead of calling lookup_order.&lt;br&gt;
You fix the cause, not the symptom. Rather than patching individual bad replies, you fix the schema mismatch that's been quietly degrading every session since the update.&lt;/p&gt;

&lt;p&gt;Without drift measurement, you'd likely learn about this from a spike in complaints, weeks after it started, with no easy way to find the first affected session. Without time-travel debugging, you'd know something changed but have no way to inspect the state that produced the change short of trying to reproduce it live, which frequently doesn't work against a non-deterministic system.&lt;/p&gt;

&lt;p&gt;Why this matters for auditing specifically, not just debugging&lt;/p&gt;

&lt;p&gt;Debugging and auditing overlap but aren't identical. Debugging is done by the team that built the agent, usually right after something breaks. Auditing is often done by someone else, on a schedule or in response to an incident, and it has to hold up as evidence, not just as an explanation.&lt;/p&gt;

&lt;p&gt;That changes what you need from the tooling:&lt;/p&gt;

&lt;p&gt;Reproducibility of the record, not the run. An auditor doesnn't need to re-run the agent. They need to trust that what db.at() shows them is what actually happened, not a best guess. That's why ZizkaDB pairs these functions with tamper-evident, checksum-backed event storage, so the record you're time-traveling into is verifiably the original one.&lt;br&gt;
A defensible answer to "how long has this been happening?" This is precisely what db.baseline() is for. It's a materially different exercise from noticing one bad reply and hoping it was a one-off.&lt;br&gt;
Compliance relevance. The EU AI Act's Article 12 asks providers of high-risk systems to keep logs that help identify situations involving risk and support post-market monitoring (Article 72). Drift measurement is close to a direct implementation of ongoing post-market monitoring. Time-travel debugging is close to a direct implementation of the record-keeping needed to investigate a flagged situation. For the full compliance picture, including human oversight and record-keeping design, see our guide to making AI agents EU AI Act compliant.&lt;br&gt;
Human oversight evidence. If reviewers approve or override agent decisions, logging those as events (as covered in the causal lineage post) means db.at() can reconstruct not just what the agent did, but what a human did in response, at the same moment.&lt;br&gt;
How this fits with observability&lt;/p&gt;

&lt;p&gt;If you already run an observability tool for latency, cost, or error rates, you don't need to replace it. Span-tree observability tools are good at how the system performed. Time-travel debugging and drift measurement, as ZizkaDB frames it, are built to answer why a specific decision happened and whether behavior has changed, using an explicit causal record and Postgres-backed storage you control rather than a hosted trace pipeline. Many teams will want both: performance monitoring for operations, and an audit trail for the questions that come from outside the engineering team.&lt;/p&gt;

&lt;p&gt;Getting started&lt;/p&gt;

&lt;p&gt;Requires Docker; the first image pull can take 5 to 10 minutes.&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
curl -fsSL &lt;a href="https://raw.githubusercontent.com/Zizka-ai/ZizkaDB/main/scripts/quickstart-remote.sh" rel="noopener noreferrer"&gt;https://raw.githubusercontent.com/Zizka-ai/ZizkaDB/main/scripts/quickstart-remote.sh&lt;/a&gt; | bash&lt;/p&gt;

&lt;p&gt;Read the script before piping it into a shell, as you would with any installer. Or self-host from a clone:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
git clone &lt;a href="https://github.com/Zizka-ai/ZizkaDB.git" rel="noopener noreferrer"&gt;https://github.com/Zizka-ai/ZizkaDB.git&lt;/a&gt; &amp;amp;&amp;amp; cd ZizkaDB&lt;br&gt;
bash scripts/setup-local.sh&lt;/p&gt;

&lt;p&gt;That gives you the API at localhost:8000, the dashboard at localhost:3001, and Swagger docs at localhost:8000/swagger. From there:&lt;/p&gt;

&lt;p&gt;Instrument one agent's sessions with db.log() and parent_id, as in the causal lineage guide.&lt;br&gt;
Let it run long enough to establish a meaningful baseline, then call db.baseline().&lt;br&gt;
Pick a past session and call db.at() on it, to confirm the reconstruction shows what you expect.&lt;br&gt;
Build a habit of checking drift on a schedule, and treat a widening gap as a trigger to time-travel into the affected sessions, not just a number to note and move past.&lt;/p&gt;

&lt;p&gt;ZizkaDB is open source under AGPL-3.0 (the MCP server is MIT), with native support for Python, TypeScript, LangChain, CrewAI, LiveKit, and MCP, plus an optional managed cloud at db.zizka.ai if you'd rather not run infrastructure. Code and docs are on GitHub.&lt;/p&gt;

&lt;p&gt;Conclusion&lt;/p&gt;

&lt;p&gt;An agent that can't be time-traveled into can't really be audited, only guessed about. An agent that isn't measured for drift can only be audited after someone notices something's wrong. Used together, the two turn "we think it's fine" into "here's what it knew, here's how its behavior has moved, and here's the moment things changed." That's the actual bar an audit has to clear, and it's the reason these two capabilities, more than dashboards or error rates, are what make an agent auditable rather than merely monitored.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>zizkadb</category>
      <category>agents</category>
    </item>
    <item>
      <title>How to Make Your AI Agent EU AI Act Compliant: A Practical Guide</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Mon, 21 Sep 2026 08:11:09 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/how-to-make-your-ai-agent-eu-ai-act-compliant-a-practical-guide-f9</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/how-to-make-your-ai-agent-eu-ai-act-compliant-a-practical-guide-f9</guid>
      <description>&lt;p&gt;&lt;strong&gt;A step-by-step guide to making AI agents compliant with the EU AI Act&lt;/strong&gt;, covering risk classification, Article 12 logging, human oversight, and post-market monitoring, with a hands-on implementation using the open-source ZizkaDB audit trail database.&lt;/p&gt;

&lt;p&gt;Most AI compliance advice was written for models: an input goes in, a prediction comes out, and you document it. Agents don't work that way. They plan, call tools, read and write memory, hand work to other agents, and act in the real world. When a regulator or a customer's auditor asks "why did your system do that?", the answer is spread across many steps that no single log line explains.&lt;/p&gt;

&lt;p&gt;This guide walks through what the EU AI Act expects from teams building or deploying agents, in the order you'd tackle it. It also shows how to build the runtime evidence layer with ZizkaDB, an open-source audit trail database for AI agents.&lt;/p&gt;

&lt;p&gt;The timeline: deferred, not cancelled``&lt;/p&gt;

&lt;p&gt;Date            What applies&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Feb 2, 2025  Prohibited practices&lt;/li&gt;
&lt;li&gt;Aug 2, 2025  General-purpose AI (GPAI) model obligations&lt;/li&gt;
&lt;li&gt;Aug 2, 2026  Transparency duties under Article 50&lt;/li&gt;
&lt;li&gt;Dec 2, 2027  Standalone high-risk systems (Annex III), after the Digital Omnibus deferral&lt;/li&gt;
&lt;li&gt;Aug 2, 2028  High-risk systems embedded in regulated products (Annex I)&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The Digital Omnibus moved the high-risk dates back but didn't change what the obligations are. Article 12 logging, human oversight, and post-market monitoring are still coming. Logging is the one obligation you can't fix retroactively, because you can't produce records of decisions that were never captured.&lt;/p&gt;

&lt;p&gt;Penalties are real: up to €35 million or 7% of worldwide turnover for prohibited practices, and up to €15 million or 3% for breaches of most other obligations, including those for high-risk systems. Verify these dates against the final published text, since this area is still moving.&lt;/p&gt;

&lt;p&gt;Step 1: Work out what your agent is under the Act&lt;/p&gt;

&lt;p&gt;"AI agent" isn't a legal category. The Act cares about what your system does and who it affects.&lt;/p&gt;

&lt;p&gt;Your role. If you build an agent and place it on the market or put it into service under your name, you're a provider. If you use someone else's agent in your business, you're a deployer. Both carry logging duties: providers under Article 19, deployers under Article 26(6). Many companies are both.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your risk tier.
&lt;/h2&gt;

&lt;p&gt;Prohibited: manipulative techniques, social scoring, and similar practices.&lt;br&gt;
High-risk: systems used in areas like employment, credit and essential services, education, law enforcement, migration, and critical infrastructure.&lt;br&gt;
Limited-risk: systems that interact with people or generate content, which triggers Article 50 transparency duties.&lt;br&gt;
Minimal-risk: everything else, with no specific obligations.&lt;/p&gt;

&lt;p&gt;The agent-specific wrinkle. An agent's capabilities can grow without anyone rewriting its core code. Add a tool, widen a permission, or connect a new data source, and its effective purpose may move from "drafts emails" to "screens candidates." Treat every change to tools, permissions, and model versions as a potential compliance event, and record it.&lt;/p&gt;

&lt;p&gt;Step 2: Map obligations to evidence&lt;/p&gt;

&lt;p&gt;For high-risk systems, the core requirements look like this:&lt;/p&gt;

&lt;p&gt;Article What it asks               What an auditor will want to see&lt;br&gt;
9   Risk management system     Documented risks, mitigations, testing results&lt;br&gt;
10  Data governance            Data provenance, quality checks, bias examination&lt;br&gt;
11  Technical documentation    Architecture, capabilities, limitations, change history&lt;br&gt;
12  Record-keeping             Automatic, lifetime event logs&lt;br&gt;
13  Transparency to deployers  Clear instructions for use&lt;br&gt;
14  Human oversight        Proof humans can understand, intervene, and stop the system&lt;br&gt;
15  Accuracy, robustness, cybersecurity Test results, resilience measures, integrity controls&lt;br&gt;
72  Post-market monitoring    Ongoing performance and behavior data&lt;br&gt;
73  Serious incident reporting Ability to reconstruct and report incidents&lt;/p&gt;

&lt;p&gt;Articles 9, 10, 11, and 13 are mainly documents you write once and maintain. Articles 12, 14, 72, and 73 depend on runtime evidence: records generated while the agent operates. For agents, that second group is where most teams are underprepared.&lt;/p&gt;

&lt;p&gt;Step 3: Lay the paperwork foundation&lt;/p&gt;

&lt;p&gt;Before the runtime layer, get the static documents right and keep them alive:&lt;/p&gt;

&lt;p&gt;A risk register that names agent-specific hazards: tool misuse, prompt injection, runaway loops, silent behavioral drift.&lt;br&gt;
Data governance records for training, fine-tuning, and retrieval sources.&lt;br&gt;
Technical documentation covering the agent's tools, permissions, and model versions, with a change log.&lt;br&gt;
A quality management process that ties releases to testing.&lt;/p&gt;

&lt;p&gt;These are necessary but not sufficient. A perfect risk register won't help when an auditor asks what your agent did for a specific user on a specific Tuesday.&lt;/p&gt;

&lt;p&gt;Step 4: Solve record-keeping, the hard part for agents&lt;/p&gt;

&lt;p&gt;Article 12 requires high-risk systems to technically allow automatic recording of events over the system's lifetime. The logs must help identify situations that may create risk or involve a substantial modification, support post-market monitoring, and support the deployer's monitoring of operation. Deployers must keep the logs under their control for at least six months, and sector rules may require longer.&lt;/p&gt;

&lt;p&gt;Why ordinary application logs fall short for agents:&lt;/p&gt;

&lt;p&gt;Non-determinism. The same input can produce different action sequences, so you can't reproduce behavior by re-running it.&lt;br&gt;
Multi-step chains. A decision emerges from planning, tool calls, and intermediate results. A "final answer" log hides the path.&lt;br&gt;
State. What the agent knew at the moment of a decision often matters more than the decision itself.&lt;br&gt;
Handoffs. Responsibility passes between agents, and the trail must follow it.&lt;/p&gt;

&lt;p&gt;Observability tools tell you how a system performed. An audit trail has to answer a different question: why did this decision happen, and who was responsible? That requires explicit cause-and-effect links between steps, not just timing data.&lt;/p&gt;

&lt;p&gt;What to capture for each session:&lt;/p&gt;

&lt;p&gt;Agent identity, plus the model, prompt, and tool configuration in effect&lt;br&gt;
Inputs received and the context available&lt;br&gt;
Each planning or reasoning step, at a level you can defend&lt;br&gt;
Every tool call with arguments and results&lt;br&gt;
Every output or action, and its recipient&lt;br&gt;
Human interventions: approvals, overrides, stops, and who performed them&lt;br&gt;
Configuration changes such as new tools, changed permissions, or model swaps&lt;/p&gt;

&lt;p&gt;Three design tensions to resolve deliberately:&lt;/p&gt;

&lt;p&gt;Integrity. Logs are only credible if you can show they weren't altered after the fact. The Act doesn't prescribe a mechanism, but tamper-evidence, such as checksummed, ordered events, is what turns a log into evidence.&lt;br&gt;
Retention. Keep logs long enough to meet the six-month floor and your sector rules, but not indefinitely by default. GDPR's storage limitation still applies.&lt;br&gt;
Privacy. Article 12 doesn't mean "log every personal detail." Use pseudonymous references instead of raw identifiers and keep sensitive payloads out of event data where you can. This keeps erasure requests rare and manageable.&lt;/p&gt;

&lt;p&gt;Step 5: Implement the audit layer with ZizkaDB&lt;/p&gt;

&lt;p&gt;ZizkaDB is an open-source, self-hosted audit trail database for AI agents. Its central idea fits Article 12 well. Every agent step is logged with a parent_id pointing to the step that caused it, so you can pick any action and walk back to its root cause with db.why(). That causal chain is what turns a pile of logs into an explanation.&lt;/p&gt;

&lt;p&gt;The open-source repo contains the API, a tenant dashboard, the SDKs, and the MCP server. Data lives in your own Postgres, so you decide where it's hosted. That matters for teams with EU data-residency requirements.&lt;/p&gt;

&lt;p&gt;Get it running&lt;/p&gt;

&lt;p&gt;You need Docker. The first image pull can take 5 to 10 minutes.&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
curl -fsSL &lt;a href="https://raw.githubusercontent.com/Zizka-ai/ZizkaDB/main/scripts/quickstart-remote.sh" rel="noopener noreferrer"&gt;https://raw.githubusercontent.com/Zizka-ai/ZizkaDB/main/scripts/quickstart-remote.sh&lt;/a&gt; | bash&lt;/p&gt;

&lt;p&gt;As with any script piped into a shell, read it before you run it. Or self-host from a clone:&lt;/p&gt;

&lt;p&gt;bash&lt;br&gt;
git clone &lt;a href="https://github.com/Zizka-ai/ZizkaDB.git" rel="noopener noreferrer"&gt;https://github.com/Zizka-ai/ZizkaDB.git&lt;/a&gt; &amp;amp;&amp;amp; cd ZizkaDB&lt;br&gt;
bash scripts/setup-local.sh&lt;/p&gt;

&lt;p&gt;You get the API at localhost:8000, the dashboard at localhost:3001/login, and Swagger docs at localhost:8000/swagger. The local setup uses a built-in dev key, so treat that as a local-testing convenience and configure your own credentials for anything real, following the self-hosting guide in the repo.&lt;/p&gt;

&lt;p&gt;Log decisions with causal links&lt;/p&gt;

&lt;p&gt;Here's the pattern applied to a high-risk scenario: an agent screening credit applications, which falls under the Annex III creditworthiness category. The event names and payload fields are illustrative. The agent, event, data, and parent_id parameters are the ones the repo documents.&lt;/p&gt;

&lt;p&gt;`&lt;code&gt;&lt;/code&gt;python&lt;br&gt;
python&lt;br&gt;
import asyncio&lt;br&gt;
from zizkadb import ZizkaDB&lt;/p&gt;

&lt;p&gt;async def main():&lt;br&gt;
    async with ZizkaDB(host="&lt;a href="http://localhost:8000%22" rel="noopener noreferrer"&gt;http://localhost:8000"&lt;/a&gt;) as db:&lt;br&gt;
        request = await db.log(&lt;br&gt;
            agent="credit-screening-agent",&lt;br&gt;
            event="user_message",&lt;br&gt;
            data={"applicant_ref": "app_7f3a", "task": "assess application"},&lt;br&gt;
        )&lt;br&gt;
        lookup = await db.log(&lt;br&gt;
            agent="credit-screening-agent",&lt;br&gt;
            event="tool_call",&lt;br&gt;
            data={"tool": "credit_lookup", "applicant_ref": "app_7f3a"},&lt;br&gt;
            parent_id=request.event_id,&lt;br&gt;
        )&lt;br&gt;
        decision = await db.log(&lt;br&gt;
            agent="credit-screening-agent",&lt;br&gt;
            event="llm_response",&lt;br&gt;
            data={"model": "gpt-4o", "prompt_version": "v14", "recommendation": "refer_to_human"},&lt;br&gt;
            parent_id=lookup.event_id,&lt;br&gt;
        )&lt;br&gt;
        approval = await db.log(&lt;br&gt;
            agent="credit-screening-agent",&lt;br&gt;
            event="human_review",&lt;br&gt;
            data={"reviewer_ref": "rev_212", "outcome": "approved_with_conditions"},&lt;br&gt;
            parent_id=decision.event_id,&lt;br&gt;
        )&lt;br&gt;
        (await db.why(approval.event_id)).print()&lt;/p&gt;

&lt;p&gt;asyncio.run(main())&lt;br&gt;
&lt;code&gt;&lt;/code&gt;`&lt;/p&gt;

&lt;p&gt;Notice what this gives you. &lt;/p&gt;

&lt;p&gt;The applicant appears only as a pseudonymous reference, the model and prompt version are recorded with the decision, and the human review is a first-class event linked to the recommendation it followed. Running db.why() on the approval walks the whole chain back to the original request. You can do the same from the terminal with zizkadb why , or in the dashboard by opening an event under Activity and choosing the Why? (causal) tab.&lt;/p&gt;

&lt;p&gt;Native integrations&lt;/p&gt;

&lt;p&gt;You don't have to hand-write logging for every framework. ZizkaDB ships packages for the stacks agent teams actually use:&lt;/p&gt;

&lt;p&gt;Stack   Install Notes&lt;br&gt;
Python  pip install zizkadb-sdk Core async SDK&lt;br&gt;
TypeScript  npm install zizkadb-sdk Same package name on npm&lt;br&gt;
LangChain   pip install zizkadb-langchain   Guide&lt;br&gt;
CrewAI  pip install zizkadb-crewai  Guide&lt;br&gt;
LiveKit (voice) pip install zizkadb-livekit One call becomes one session; transcript only, no audio stored (guide)&lt;br&gt;
MCP uvx zizkadb-mcp For MCP clients such as Cursor (README)&lt;br&gt;
Any other agent REST API    Swagger docs on your instance&lt;/p&gt;

&lt;p&gt;The transcript-only design of the LiveKit integration is useful for compliance. You get an auditable record of what a voice agent said and did without retaining audio you'd then have to govern.&lt;/p&gt;

&lt;p&gt;How the features map to the Act&lt;/p&gt;

&lt;p&gt;Requirement ZizkaDB capability&lt;br&gt;
Art. 12: automatic recording    SDKs, framework packages, MCP, and REST let agents write events as they run&lt;br&gt;
Integrity of records    Tamper-evident, checksum-backed decision logs&lt;br&gt;
Art. 73: incident reconstruction    db.why() and session replay walk any action back to its cause&lt;br&gt;
"What did it know at that moment?"  db.at() reconstructs what the agent knew at a given timestamp&lt;br&gt;
Art. 72: post-market monitoring db.baseline() detects when behavior drifts from past sessions&lt;br&gt;
Finding relevant history    db.search() runs semantic search over agent history&lt;br&gt;
GDPR erasure requests   db.forget() erases by metadata filter&lt;br&gt;
Deployment notes&lt;br&gt;
Licensing. The repo is AGPL-3.0, and the MCP server is MIT. Running it unmodified for your own agents is generally straightforward. If you modify it and offer it over a network, AGPL's network clause applies, so involve counsel.&lt;br&gt;
Telemetry. You can turn off ZizkaDB's telemetry with ZIZKADB_TELEMETRY=false, which compliance-minded teams will want to do.&lt;br&gt;
Erasure versus integrity. db.forget() handles GDPR erasure, but deleting from a tamper-evident record deserves a deliberate test. Keep personal data out of payloads by design, then verify in your own environment how erasure interacts with your integrity requirements.&lt;br&gt;
Managed option. If you'd rather not run infrastructure, there's a hosted version at db.zizka.ai.&lt;br&gt;
What ZizkaDB doesn't do&lt;/p&gt;

&lt;p&gt;It doesn't classify your system, write your risk management file, run your data governance, or complete a conformity assessment. It supports record-keeping and monitoring, the parts of compliance that depend on runtime evidence. Treat it as one component of a compliance program, not a substitute for one.&lt;/p&gt;

&lt;p&gt;A practical rollout&lt;br&gt;
Define your event schema first, using the capture list from Step 4. Decide what's logged in full, what's pseudonymized, and what's stored by reference.&lt;br&gt;
Instrument your highest-risk agent end to end, not your easiest one.&lt;br&gt;
Log configuration changes as events so new tools, permissions, and model swaps sit in the same trail as decisions.&lt;br&gt;
Set retention and access rules, and restrict who can read the audit trail itself.&lt;br&gt;
Establish a baseline and use drift alerts as inputs to your risk register.&lt;br&gt;
Run an audit drill (Step 8).&lt;/p&gt;

&lt;p&gt;Step 6: Design human oversight you can prove&lt;/p&gt;

&lt;p&gt;Article 14 requires that people can understand what the system is doing, notice when it misbehaves, avoid over-trusting its output, override or disregard it, and stop it.&lt;/p&gt;

&lt;p&gt;For agents, that means real control points: approval gates before high-impact or irreversible actions, a working stop mechanism, and interfaces that show reviewers why the agent proposes something, not just what.&lt;/p&gt;

&lt;p&gt;"We have a human in the loop" isn't evidence. A record of who reviewed what, when, and what they decided is. In the example above, the human_review event is linked to the recommendation it followed, so an auditor sees the agent's action and the human's response in one causal chain.&lt;/p&gt;

&lt;p&gt;Step 7: Monitor after deployment and be ready to report&lt;/p&gt;

&lt;p&gt;Providers of high-risk systems need post-market monitoring (Article 72) and must report serious incidents (Article 73). Agent behavior can shift when a model provider updates a model, when retrieved data changes, or when users find new ways to use the system.&lt;/p&gt;

&lt;p&gt;Track behavior, not just uptime. Compare action patterns and outcomes against a baseline. In ZizkaDB, that's what db.baseline() is for.&lt;br&gt;
Have an incident playbook. Define what counts as a serious incident, who decides, and how fast you can assemble the facts. Reconstruction speed depends on record quality.&lt;br&gt;
Feed findings back. Monitoring results should update your risk file and documentation.&lt;/p&gt;

&lt;p&gt;Step 8: Run an audit drill&lt;/p&gt;

&lt;p&gt;The best test of your setup is simulating the audit. Pick a random session from three months ago and try to answer:&lt;/p&gt;

&lt;p&gt;What was the agent asked, and what context did it have? (db.at())&lt;br&gt;
What steps did it take, and why? (db.why())&lt;br&gt;
Which model, prompt, and configuration was live at the time?&lt;br&gt;
Did a human review or override anything?&lt;br&gt;
Can you prove the record hasn't changed since?&lt;/p&gt;

&lt;p&gt;If any answer takes more than an hour, or comes from someone's memory instead of a record, you've found a gap while it's still cheap to fix.&lt;/p&gt;

&lt;p&gt;Common mistakes&lt;/p&gt;

&lt;p&gt;Logging outputs only. Auditors care about the path to a decision.&lt;br&gt;
Treating compliance as documentation. Documents describe intent. Runtime records prove behavior.&lt;br&gt;
Ignoring configuration changes. A new tool can change your risk class overnight.&lt;br&gt;
Logging too much personal data. Over-collection creates a GDPR problem while solving an AI Act one.&lt;br&gt;
Waiting for the deadline. You can't backfill records. The trail starts when you start logging.&lt;br&gt;
Overclaiming. Calling your system "compliant" before conformity assessment is a risk in itself. Say "supports" until you can prove "meets."&lt;/p&gt;

&lt;p&gt;Compliance checklist&lt;/p&gt;

&lt;p&gt;Role (provider, deployer, or both) and risk tier documented&lt;br&gt;
 Agent capabilities, tools, and permissions inventoried&lt;br&gt;
 Risk register includes agent-specific hazards&lt;br&gt;
 Data governance and technical documentation maintained&lt;br&gt;
 Automatic event recording with causal links between steps&lt;br&gt;
 Tamper-evidence, retention schedule, and access controls defined&lt;br&gt;
 Personal data minimized or pseudonymized in logs&lt;br&gt;
 Human oversight controls built and their use recorded&lt;br&gt;
 Post-market monitoring with drift detection running&lt;br&gt;
 Incident playbook written and tested&lt;br&gt;
 Audit drill completed and gaps closed&lt;/p&gt;

&lt;p&gt;Conclusion&lt;/p&gt;

&lt;p&gt;Making an AI agent EU AI Act compliant comes down to building a system that can explain itself after the fact: what it did, why, under whose oversight, and with what integrity guarantees. The paperwork matters, but the parts that depend on runtime evidence are where agents differ from traditional software, and where retrofitting is impossible.&lt;/p&gt;

&lt;p&gt;Start with classification and documentation, then put an audit layer under your highest-risk agent now. You can clone ZizkaDB from GitHub, connect it through the Python or TypeScript SDK, LangChain, CrewAI, LiveKit, or MCP, and start recording decisions today, so that when the obligations arrive, the records already exist.&lt;/p&gt;

&lt;p&gt;This article is general information, not legal advice.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>discuss</category>
      <category>debugging</category>
    </item>
    <item>
      <title>The Science of Machine Learning vs. the Push for AI Deployment</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Sat, 19 Sep 2026 13:26:25 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/the-science-of-machine-learning-vs-the-push-for-ai-deployment-3p04</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/the-science-of-machine-learning-vs-the-push-for-ai-deployment-3p04</guid>
      <description>&lt;p&gt;The premise was that AI would change the world. It would supercharge business processes, alter industrial mechanics, create entirely new industries, and lead us toward a world of abundant production. Elon Musk has even suggested that money and work might one day become unnecessary, leaving only abundance. It was a bold vision, and it was presented with great confidence.&lt;/p&gt;

&lt;p&gt;To a large extent, the promise held up in the GenAI era. We created millions of near-perfect photos and experimented with video creation. ChatGPT answered almost everything we asked, and Claude handled deep workflows that used to take teams days. Cursor became a favorite among coders who had spent their careers writing syntax by hand. This naturally led to rapid commercialization, and the valuations of frontier models rose to extraordinary levels, helped by the American talent for marketing ideas on a grand scale. Given Silicon Valley’s track record of changing the world we live in, many of us were convinced that the world was changing once again, and the fear of missing out spread across every industry, from banks to retailers to manufacturers. Nobody wanted to be left behind.&lt;/p&gt;

&lt;p&gt;Then the conversation began to shift toward reliability and determinism. The arrival of open-weight and open-source models challenged the notion, strongly promoted by the frontier LLM companies, that hundreds of billions of dollars were needed to make LLMs near perfect. If a capable model could be downloaded, fine-tuned, and run on your own infrastructure, the competitive moat looked smaller than advertised. So a new race began: the deployment of AI, again led by Silicon Valley through influential gatekeepers like Y Combinator, which shape what gets funded, what gets attention, and which story the market hears next.&lt;/p&gt;

&lt;p&gt;Today, two kinds of companies are emerging globally. The first are those deploying AI across almost every industry vertical, working to keep the promise of AI alive. The second are those taking a more scientific approach, trying to make AI more deterministic through guardrails, observability layers, and similar tools. Both are responding to the same challenge: the technology was presented faster than it could be adopted.&lt;/p&gt;

&lt;p&gt;We are now at a crossroads. We are trying to work against the very science of probability that gave rise to LLMs and, with them, the whole AI phenomenon. These systems are powerful precisely because they work with probabilities rather than fixed rules. Trying to replace that probability with deterministic outcomes is extremely difficult, if not impossible. It is a bit like asking the ocean to behave like a swimming pool.&lt;/p&gt;

&lt;p&gt;Some may say that I am being pessimistic or contrarian, or that I am going against the natural evolution of AI. I would respectfully disagree. First, I am a founder, not a researcher whose role is to offer opposing views. I build, I sell, and I sit across the table from people who have to make these systems work in real businesses. Second, I am pragmatic and fully aligned with the vision that AI can change the world we live in, though perhaps not in the way it is often promised. I believe AI is as significant as the Internet, if not more so. Like the Internet, it will empower us, bring more human creativity, push us to rethink how we work, and help us create more powerful processes. The Internet did not remove humans from the loop; it gave them more leverage. The idea that we can simply sit back and let everything happen autonomously, at scale, is not a realistic expectation.&lt;/p&gt;

&lt;p&gt;Before looking at both kinds of efforts, it is worth pausing on a recent development that I read as a shift in messaging from Silicon Valley.&lt;/p&gt;

&lt;p&gt;Dario Amodei is preparing for an IPO of a company that is still loss-making, at a valuation of around 2 trillion. There is nothing wrong with that, and I do not question it morally, but it comes with its own dynamics. When a company goes public, its story has to hold up under the scrutiny of public markets, and that naturally changes how the story is told.&lt;/p&gt;

&lt;p&gt;He recently wrote an article calling for a slowdown in frontier AI progress and describing it as a challenge for the human race. Musk and Altman have voiced agreement. Not long ago, these same leaders spoke to us about AGI, about a world where money may not be needed, about abundant production, and about the point of singularity. Now they are calling for a slowdown. Why? My view is that this is not about the capabilities of frontier models. It reflects the fact that adoption at scale has not yet followed. Capability is racing ahead, but enterprises are not absorbing it at the same speed, and the gap between what these models can do and what businesses can trust them to do is where the real challenge lies. That gap has led to the efforts of vertical AI companies and of those working to enforce determinism, so let us look at each of them.&lt;/p&gt;

&lt;p&gt;Vertical AI Solutions / Application Layers&lt;/p&gt;

&lt;p&gt;This is one way, and perhaps the most discussed, most promoted, and most funded way, to deploy AI into processes, industries, and businesses for commercial use. Every sector now has its own group of startups: AI for legal, AI for healthcare, AI for customer support, AI for finance. I personally believe this is the way forward, but there is a fundamental challenge: it sits in tension with the science of machine learning. The issue is not the process itself but the way it is being positioned.&lt;/p&gt;

&lt;p&gt;Vertical AI companies in any given sector are promising deterministic outcomes, which runs directly against how machine learning works. I understand that outputs have to be quantified in order to sell them, since no enterprise buyer signs a contract based on “it usually works well.” But that quantification cannot be produced deterministically, because we are no longer in the API era, where 2+2 was always 4, no matter where it was calculated or how many times you asked. We are in the AI era, where AI relies on the most probable next token. The power it brings, and the scale of output it delivers, reach far beyond the API era, but the certainty being promised remains difficult to reconcile with the science it relies on.&lt;/p&gt;

&lt;p&gt;Let us take an example. Suppose a vertical AI company promises 40% higher customer satisfaction if you deploy its solution. The question is: how can they guarantee that it will always remain 40%? Being cloud-native (as most of them are) and reliant on a model, how can they promise a deterministic output? Models can change over time through the data they accumulate, the tool calls they make, the errors that occur, the API calls, and much more. The underlying model can be updated, the context can change, the data seen in production can differ from the data used in testing, and a small change in any of these can shift behavior in ways nobody planned for. Predicting a deterministic output therefore runs against the logic of agent behavior drift.&lt;/p&gt;

&lt;p&gt;I am not suggesting that 40% is an unfounded claim. It may well be accurate on the day it is measured. My point is only that it cannot be guaranteed, and this is where the challenge for enterprise adoption begins. A number that holds in a pilot may not hold six months later in production, and it may go unnoticed until the impact is felt.&lt;/p&gt;

&lt;p&gt;Consider a bank. A credit score is deterministic and derived from multiple factors. The same profile cannot receive different scores each time, because a customer, a regulator, or a court can ask why one person received one score on Monday and another on Friday, and “the model responded differently” is not an acceptable answer. Furthermore, the EU AI Act, which is a much-needed step for the industry, requires compliance with clear rules on transparency, traceability, and human oversight. Systems that cannot explain their own behavior will find it difficult to meet those requirements. This is where the second type of company comes in.&lt;/p&gt;

&lt;p&gt;Determinism-Enforcing Companies&lt;/p&gt;

&lt;p&gt;These companies understand the science of probability well and are working to manage it with every technique available. They deserve credit for taking the problem seriously, although the results so far have been limited.&lt;/p&gt;

&lt;p&gt;For example, some introduced RAG (retrieval-augmented generation). The idea was that if a context layer was provided, the AI’s output would stay within the boundaries of the data provided. In practice, this did not fully work. Even within a given set of data, the science of machine learning remains in play, and the LLM still selects the most probable next answer. The model can still misread a document, give more weight to one source over another, or blend retrieved content with what it already “knows.” The output remains a probability, just a better-informed one.&lt;/p&gt;

&lt;p&gt;After that, companies introduced context layers, memory, and hardcoded guardrails. Each of these adds some safety, but each also adds complexity and new points of failure, and none has fully solved the problem so far. At least, I have not yet come across a company that claims full compliance with the EU AI Act. This may be one reason for slower adoption, and possibly for the delay by policymakers in fully implementing the Act. If the industry cannot yet demonstrate that it meets the rules, regulators face a difficult choice between enforcing standards that very few can meet and extending the deadlines.&lt;/p&gt;

&lt;p&gt;So what is the solution?&lt;/p&gt;

&lt;p&gt;As I mentioned earlier, I see myself as a pragmatic person, and I believe in solutions more than opinions. My conclusion remains the same: AI adoption is a necessity, but not autonomous adoption. I believe in empowering humans with AI, and my answer is to make processes auditable.&lt;/p&gt;

&lt;p&gt;If we cannot make probability deterministic, we should be honest about it. Instead of trying to remove uncertainty from the model, we should make that uncertainty visible, measurable, and controllable. The premise on which we are building AI infrastructure at Zizka AI (ZizkaDB) is simple: making AI processes auditable, with human oversight, so that we keep the power AI offers while retaining a human check on it.&lt;/p&gt;

&lt;p&gt;In practice, this means three things. First, humans decide what level of agent drift is tolerable, because the acceptable level of risk for a marketing assistant is very different from that for a credit decision. Second, we can understand why an agent does what it does, so its decisions are not a black box but something a person can inspect, question, and explain. Third, we can replay everything in production, so that when something goes wrong, or even when something goes right, we gain deep insight into how it happened and can learn from it.&lt;/p&gt;

&lt;p&gt;Essentially, we are trying to reverse the trend of handing everything over to AI, and instead place the power of AI in human hands so that humans can decide and act. AI provides the capability, while humans provide judgment and accountability. I believe this is the version of the future that businesses, regulators, and ordinary people can genuinely trust.&lt;/p&gt;

&lt;p&gt;If you are interested in knowing more about ZizkaDB, check our open-source repository here.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Zizka-ai/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/Zizka-ai/ZizkaDB&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://zizka.ai" rel="noopener noreferrer"&gt;https://zizka.ai&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>claude</category>
      <category>database</category>
    </item>
    <item>
      <title>Help needed to hit 100 Star Milestone</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Fri, 18 Sep 2026 12:33:33 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/help-needed-to-hit-100-star-milestone-4lh4</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/help-needed-to-hit-100-star-milestone-4lh4</guid>
      <description>&lt;p&gt;Hello Community!&lt;/p&gt;

&lt;p&gt;4122 Downloads, 294 Active Installations and &lt;br&gt;
80 Stars , Need 20 More  to hit first 100!&lt;/p&gt;

&lt;p&gt;If you are using ZizkaDB, let me know your feedback, if you arent, Use it , if you think there is some improvement needed, give us a PR.&lt;/p&gt;

&lt;p&gt;Most Importantly help us to reach 100 Stars &lt;/p&gt;

&lt;p&gt;Support Please .&lt;/p&gt;

&lt;p&gt;Opensource milestone - ZizkaDB&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/Zizka-ai/ZizkaDB" rel="noopener noreferrer"&gt;https://github.com/Zizka-ai/ZizkaDB&lt;/a&gt;&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>development</category>
      <category>zizkadb</category>
      <category>operationaldatabase</category>
    </item>
    <item>
      <title>Cloud Hosting Is the Data Hostage Model, and It Has No Place in the AI Era</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Fri, 11 Sep 2026 13:41:08 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/cloud-hosting-is-the-data-hostage-model-and-it-has-no-place-in-the-ai-era-356i</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/cloud-hosting-is-the-data-hostage-model-and-it-has-no-place-in-the-ai-era-356i</guid>
      <description>&lt;p&gt;The thesis, stated plainly: cloud hosting did not just happen to make customer data hard to leave with. The architecture was built around continuous delivery and operational convenience, and one of its natural side effects was that data became hard to move. That side effect quietly became part of the retention story, whether or not any individual vendor set out to build it that way. I call this the Data Hostage Model, not as an accusation against any one company, but as a description of what the architecture structurally produces. It worked well for two decades. I don’t think it is the right foundation for the AI era, and this piece explains why, in three stages. First, why the model won the Application Era and where its logic starts to strain. Second, why AI agent operational data is a category of information that shouldn’t be built on that logic in the first place. Third, why VPC deployment and licensing, not data leverage, are the more durable way forward, and why that is the direction we are building ZizkaDB at Zizka AI.&lt;/p&gt;

&lt;p&gt;Part One: Understanding the Data Hostage Model, Not Just Naming It&lt;br&gt;
A disagreement worth having in public&lt;br&gt;
A while back I asked the CEO of Aiven a direct question: what’s the actual path forward for open source companies to monetize? His answer, in essence, was that cloud hosting eventually wins. However you start, the managed-cloud model is where sustainable open source businesses end up.&lt;/p&gt;

&lt;p&gt;I think he’s largely right about what happens in practice, and I respect the case. Aiven has built a genuinely large business on a real premise: take strong open-source data infrastructure, run it so customers don’t have to, and get paid for the operational burden you take off their plate. By their own account, a meaningful part of why customers stay isn’t the software itself, it’s the trust built up over years of handling security, patching, and uptime reliably. That’s a legitimate insight, and it has earned real revenue and real customer goodwill.&lt;/p&gt;

&lt;p&gt;I don’t disagree with the diagnosis. I disagree that this is the model to build the next decade on. And I think the AI era is exactly when that disagreement starts to matter in practice, not just in principle.&lt;/p&gt;

&lt;p&gt;How the model actually works&lt;br&gt;
Here is the mechanic worth naming plainly, without dressing it up and without overstating it either. The code is free, so nobody pays for the software directly. What gets sold instead is operational convenience: someone else runs it, patches it, keeps it online. To deliver that convenience, the vendor has to hold the data. And once the data is sitting inside someone else’s infrastructure, wired into someone else’s dashboards and backup systems, leaving stops being a quick decision. It becomes a migration project.&lt;/p&gt;

&lt;p&gt;That gap between deciding to leave and actually being able to leave is not necessarily a deliberate trap. Most vendors building this way are optimizing for reliability and product velocity, not for lock-in as an end in itself. But the gap is real, and it functions as retention whether or not anyone intended it to. A customer who stays because switching would take months of engineering time and a customer who stays because the product earns it every quarter look identical on a retention chart. The Data Hostage Model, as an architecture, doesn’t distinguish between those two customers. That’s the part worth being honest about.&lt;/p&gt;

&lt;p&gt;This worked well, and reasonably, for most application data, because the tradeoff was worth it. A production Postgres cluster or a Kafka pipeline isn’t usually sensitive in a way that makes centralized hosting feel risky. The value was real operational convenience, and the cost was a manageable amount of trust extended to a vendor with every incentive to keep that trust intact.&lt;/p&gt;

&lt;p&gt;What the arrangement depends on, though, is trust standing in for verification. You trust the vendor is patching on schedule. You trust they aren’t looking at your data. You trust their security posture. You don’t independently verify most of it, you take their word, backed by their reputation and their SLA. That was a reasonable trade for the Application Era. I think it gets harder to justify once the data in question is no longer just sitting there.&lt;/p&gt;

&lt;p&gt;Why AI agents change the calculus&lt;br&gt;
Here’s the shift: AI agents don’t just store your data, they act on it, autonomously, in production, making decisions that have real consequences. Approving a refund, calling a tool, escalating a support ticket, taking an action on someone’s behalf. When something goes wrong, the question isn’t just whether the data was secure. It’s whether you can see, step by step, why the agent did what it did, and explain that to a regulator, an auditor, or a customer who was affected.&lt;/p&gt;

&lt;p&gt;That’s a materially different requirement than uptime and patching. It’s evidentiary, not just operational. And evidentiary needs are hard to satisfy through a trust relationship alone. They’re better satisfied by being able to inspect the chain yourself, on your own infrastructure, without waiting on a vendor’s support queue at the exact moment you need an answer.&lt;/p&gt;

&lt;p&gt;This isn’t a hypothetical concern. Regulation is already moving in this direction. The EU AI Act is pushing toward exactly this kind of demand: agent behavior needs to be explainable and auditable, not just monitored by whoever happens to be hosting it. An answer like the AI just made a mistake is becoming less acceptable, in a legal sense as much as a product sense.&lt;/p&gt;

&lt;p&gt;It’s not only regulation. It’s the practical reality that the operational logs of an AI agent, the decisions, the tool calls, the causal chain of why something happened, are often more sensitive than the application data sitting next to them, because they reveal how your system reasons, what it has access to, and where it can go wrong. Centralizing that inside a third party’s cloud is a bigger ask than centralizing a database backup ever was, and it’s worth treating it as a genuinely different category rather than assuming the old model still applies.&lt;/p&gt;

&lt;p&gt;Even Aiven has visibly moved in this direction itself, offering deployment directly into a customer’s own cloud account for data residency and private networking. I read that as a useful signal: when a company whose business is built on centralized hosting starts building a path for customers to keep infrastructure in their own account, it suggests the market is already asking for more control than the original model assumed it would want.&lt;/p&gt;

&lt;p&gt;The reframe: from control to verification&lt;br&gt;
Here’s the reframe at the center of this piece. In the Application Era, the thing that mattered most was control: control over infrastructure, control over reliability, control over how convenient it was to stay. Trust was largely a story built on top of that control, earned by vendors who exercised it responsibly over time.&lt;/p&gt;

&lt;p&gt;In the AI era, I think the thing that matters most shifts toward verification: not just trusting that data is handled responsibly, but being able to check, directly, on infrastructure you own, without needing to take anyone’s word for it. That’s a more durable form of trust, because it doesn’t depend on a vendor’s continued good behavior. It’s trust that holds up on its own.&lt;/p&gt;

&lt;p&gt;This is the direction we’re building in at Zizka AI. ZizkaDB exists to answer the question an AI agent’s operators need answered under pressure: why did the agent do that, what did it know when it acted, has its behavior drifted. We believe that answer needs to be verifiable on infrastructure the customer controls, not something requested from a vendor’s dashboard after the fact. That’s why the product is open source under AGPL from the start, with self-hosting as a first-class option rather than an enterprise add-on.&lt;/p&gt;

&lt;p&gt;To be clear, this isn’t a claim that cloud-hosted vendors are acting in bad faith, or that Aiven’s approach is wrong for the problems it was built to solve. It’s an argument that the category of data involved in AI agent operations, more sensitive, more regulated, more consequential when something breaks, deserves a different default than the one the Application Era settled on.&lt;/p&gt;

&lt;p&gt;Cloud hosting won the last era on a genuine insight about trust. The rest of this piece is about what trust needs to look like when the systems being trusted are the ones making decisions.&lt;/p&gt;

&lt;p&gt;Part Two: Why Agent Operational Data Deserves a Different Default&lt;br&gt;
The assumption worth questioning&lt;br&gt;
Here’s an assumption that sounds reasonable and, I think, doesn’t hold up well under scrutiny: an AI agent’s logs are just another kind of data, so they can be hosted the same way any other database has always been hosted.&lt;/p&gt;

&lt;p&gt;That assumption applies the Application Era’s default to a new category of information without checking whether the category actually fits. It fits reasonably well for rows in a customer table. It fits for a Kafka topic full of order events. It fits less well for the record of why an autonomous system decided to do what it did, and the difference is worth taking seriously.&lt;/p&gt;

&lt;p&gt;What agent operational data actually is&lt;br&gt;
Application data is mostly inert until something queries it. A row in a Postgres table sits there. It doesn’t reveal how your business reasons, only what it recorded. If a vendor holds it and something goes wrong, you’ve lost some convenience, not much else.&lt;/p&gt;

&lt;p&gt;Agent operational data is different in kind. It’s the trace of a decision process: what the agent knew at the moment it acted, which tool it called and why, what context shaped that choice, how the outcome might have differed if one earlier step had gone differently. It’s the reasoning of a system standing in for a human making judgment calls. That distinction matters, because it means agent operational data carries three properties application data usually doesn’t.&lt;/p&gt;

&lt;p&gt;First, it’s diagnostic of failure. When an agent makes a bad call, the operational record is often the only way to understand what actually happened. Losing easy access to that record costs more than convenience, it costs your ability to explain yourself when it matters most.&lt;/p&gt;

&lt;p&gt;Second, it reveals business logic. A causal chain showing how an agent decided to approve a refund, or escalate a ticket, or skip a policy check, is close to a readable trace of your internal decision rules, more like source code than a customer record.&lt;/p&gt;

&lt;p&gt;Third, it’s often legally significant in a way most application data isn’t. Under emerging regulation like the EU AI Act, being able to explain and audit an autonomous decision is close to a requirement, not a nice-to-have. That turns agent operational data from a debugging convenience into something you may need to produce, under time pressure, to a regulator or an affected customer.&lt;/p&gt;

&lt;p&gt;None of these three properties apply cleanly to a typical managed database workload. All three apply directly to what an agent observability and audit layer holds. Treating them the same by default is where the model starts to strain.&lt;/p&gt;

&lt;p&gt;A conflict of interest worth naming&lt;br&gt;
There’s a second issue underneath the first, and it’s less about sensitivity and more about incentives. Many observability tools for AI agents, cloud-hosted by design, aggregate data across customers to improve their own product, benchmark performance, and inform future features. That’s a normal part of how SaaS businesses grow, not evidence of bad intent. But it does create a structural tension worth naming: the party holding the record of your agent’s failures is often the same party whose product improves by learning from failures across its entire customer base.&lt;/p&gt;

&lt;p&gt;Even with good contracts and good intentions, that’s an arrangement customers can’t fully verify from the outside. Nobody worries much that a hosted Postgres backup is teaching a vendor about their business. It’s reasonable to think more carefully about that when the data in question is the causal record of how your AI agents make decisions, since that record is more revealing, and more valuable to more parties, than a typical backup.&lt;/p&gt;

&lt;p&gt;The verification problem&lt;br&gt;
There’s a third issue that I think gets the least attention. If the reason you need agent operational visibility in the first place is that the agent’s own reasoning isn’t fully transparent, then the system you use to audit that agent should ideally not be another layer you have to take on faith. If your audit layer is itself a closed, cloud-hosted service, you’ve moved the transparency problem up one level instead of resolving it. It becomes harder to independently confirm that logging is complete, that nothing was dropped, that the causal chain you’re shown is the full and accurate one.&lt;/p&gt;

&lt;p&gt;Many tools in this space, including well-regarded ones built around span-based tracing, are genuinely useful for performance debugging but weren’t designed to solve this specific problem: proving, independently, on infrastructure you control, that a causal chain is complete and unaltered. Explicit causal lineage, the ability to trace backward from an outcome to the exact decision that produced it, is a stricter requirement than watching a trace of spans in a dashboard, and it holds up best when nobody else needs to be trusted to have captured it faithfully.&lt;/p&gt;

&lt;p&gt;Where this points&lt;br&gt;
Put these three factors together, evidentiary weight, a real conflict of interest, and the verification problem, and the conclusion isn’t that cloud hosting is unusable for agent operational data. It’s that defaulting to the same architecture used for application data, without adjusting for what this data actually is, is a mismatch worth correcting. What matters most here isn’t a vendor’s assurances. It’s the customer’s own ability to verify, on their own terms, at the exact moment something has already gone wrong.&lt;/p&gt;

&lt;p&gt;This is why, at Zizka AI, ZizkaDB wasn’t designed as a cloud-first product with self-hosting added later as an afterthought. It was built the other way around: self-hosting under an open license as the default, with managed cloud offered as a convenience for teams who want it, not as the direction the architecture quietly steers everyone toward. The causal lineage, the explicit decision chains, the drift baselines, all of it is designed to be inspected on infrastructure the customer actually owns, because that’s the version of this that holds up once the stakes are taken seriously.&lt;/p&gt;

&lt;p&gt;The Application Era’s model was built for data that mostly just sits there. Agent operational data doesn’t just sit there, it explains and, when needed, testifies. That difference is enough, on its own, to justify a different default for where it lives.&lt;/p&gt;

&lt;p&gt;Part Three: The Case for VPC Deployment and Licensing&lt;br&gt;
Naming the actual choice&lt;br&gt;
Every company building AI agent infrastructure eventually has to answer one question honestly, whether or not it shows up in the marketing: does the business get stronger when it’s harder for a customer to leave with their data, or does it get stronger some other way. Most architecture decisions end up answering this question implicitly, whether the founders intended to or not.&lt;/p&gt;

&lt;p&gt;Much of the software industry’s last two decades leaned toward the first answer, often without anyone deliberately choosing it. Retention got built, in part, by making data portability inconvenient, not solely by making the product good enough that leaving would feel like a loss. That’s the pattern I’ve been calling the Data Hostage Model throughout this piece. It’s not a claim that any individual vendor set out to trap customers. It’s an observation that a business whose retention depends partly on data being hard to move has a quiet, structural incentive to keep it that way, regardless of intent.&lt;/p&gt;

&lt;p&gt;This section makes the case for the second answer, as an actual architecture and an actual pricing model, not just a values statement.&lt;/p&gt;

&lt;p&gt;Why VPC deployment is more than a compliance feature&lt;br&gt;
VPC deployment, running the software inside infrastructure the customer owns and controls, tends to get framed as a compliance checkbox. Data residency, sure. Something for the security review, sure. That framing understates what it actually does.&lt;/p&gt;

&lt;p&gt;Deploying inside a customer’s own VPC changes who’s structurally capable of what. It removes the vendor’s role as the single point of access to the most sensitive data the system produces, because there’s no centralized copy to protect or misuse. It reduces how much a customer has to trust a vendor’s internal access controls, since there’s no vendor-side copy for those controls to govern. And it directly addresses the conflict of interest and verification problems described earlier, because there’s nothing aggregated across customers to learn from, and nothing standing between the customer and their own evidence.&lt;/p&gt;

&lt;p&gt;For AI agents specifically, this matters even more. If the causal record of your agent’s decisions lives inside your own infrastructure, inspectable directly, you don’t need anyone to vouch for its completeness. You can verify it yourself, exactly when a regulator, a customer, or your own team needs an answer. That’s a different kind of confidence than an uptime guarantee, and it’s closer to what the AI era actually requires.&lt;/p&gt;

&lt;p&gt;This is why VPC deployment sits at the center of how we think about ZizkaDB, not as an enterprise tier layered on top of a cloud-first product, but as the default posture the system is designed around from the start. Self-hosting under an open license, with managed cloud available for teams who prefer the convenience.&lt;/p&gt;

&lt;p&gt;Why the current playbook is hard to walk away from&lt;br&gt;
It’s worth being honest about the tension here rather than glossing over it. The dominant open-source monetization model, hosting the code in the cloud and charging for the operational burden taken off a customer’s hands, works partly because it reproduces some of the dynamics described above, usually without deliberate intent. Once a customer’s data and workloads are running in a vendor’s cloud, switching away carries real friction, and that friction is part of what makes retention numbers look strong to investors.&lt;/p&gt;

&lt;p&gt;It’s understandable why founders build this way. It works, it’s well understood, and there’s a clear, well-validated playbook, including from people who’ve executed it very well and who genuinely believe, with good reason, that trust is a real part of why their customers stay. But it’s worth being clear-eyed that some of that retention reflects switching cost as much as product quality, and that the two are easy to conflate if you’re not looking closely.&lt;/p&gt;

&lt;p&gt;What licensing actually charges for&lt;br&gt;
If VPC deployment removes the vendor’s structural leverage over customer data, the business model has to be built around something else: the software and the expertise behind it, sold directly, rather than sold indirectly through the cost of leaving.&lt;/p&gt;

&lt;p&gt;A license-based model charges for what a customer is genuinely getting: production-grade software, the specific engineering behind causal lineage and drift detection and reliable operation at scale, ongoing improvements, and support when something breaks. None of that requires holding the customer’s operational data as leverage. It requires the software being good enough, and the support being responsive enough, that renewing feels like an easy decision rather than a forced one.&lt;/p&gt;

&lt;p&gt;This is a harder business to build, at least early on. Data leverage is an effective retention mechanic, and giving it up means the product and the relationship have to earn renewal on their own. I think that trade is worth making, for two reasons. First, because in a category where the underlying data is increasingly sensitive and regulated, customer tolerance for centralized control is likely to keep declining, regardless of what any individual vendor prefers. Second, because a business built on genuine product quality tends to be more durable over time than one built primarily on switching cost, even if it grows more slowly at first.&lt;/p&gt;

&lt;p&gt;The valuation question this raises&lt;br&gt;
There’s a broader implication worth mentioning, briefly, because it extends beyond any single company’s product decisions. A meaningful share of current valuations in software, and increasingly in AI infrastructure, are priced partly on the assumption that data gravity will keep customers locked in, the way it did through the social media and SaaS eras. If the AI era genuinely shifts what customers value from control toward verification, some of that valuation logic may be resting on an assumption that doesn’t hold as firmly going forward. A moat built on switching cost is only as strong as customers’ willingness to accept it, and that willingness looks like it’s declining faster for AI agent operational data than it ever did for a CRM record or a social graph.&lt;/p&gt;

&lt;p&gt;I’m not raising this to predict a market correction or to single out any particular company. I’m raising it because it should change what founders optimize for now, while this category is still being defined, rather than later, once the market has already reassessed the difference between durable product advantages and data leverage that looked like one for a while.&lt;/p&gt;

&lt;p&gt;Where this leaves the argument&lt;br&gt;
The first part of this piece described the Data Hostage Model for what it is: an architecture that produces retention partly through switching cost, often without anyone deliberately designing it that way, and asked whether that’s the right foundation to keep building on. The second part argued that AI agent operational data is a category of information poorly suited to that model, given its evidentiary weight, the conflicts of interest it creates, and the verification problem it raises. This final part has made the case for what follows from both: VPC deployment as the architecture, because it removes structural leverage rather than just asking customers to trust that it won’t be misused, and licensing as the business model, because it charges for what’s actually valuable, the software and the expertise behind it, rather than for the inconvenience of leaving.&lt;/p&gt;

&lt;p&gt;This is the direction we’re building ZizkaDB in at Zizka AI. It’s likely a slower path to revenue than the alternative. I think it’s the more honest one, and the one more likely to hold up as the systems being trusted are increasingly the ones making the decisions.&lt;/p&gt;

&lt;p&gt;The Article is originally published in medium and can be accessed via this link : &lt;a href="https://medium.com/@MirArshadTalpur/cloud-hosting-is-the-data-hostage-model-and-it-has-no-place-in-the-ai-era-6d3fc8f26dd3?postPublishedType=initial" rel="noopener noreferrer"&gt;https://medium.com/@MirArshadTalpur/cloud-hosting-is-the-data-hostage-model-and-it-has-no-place-in-the-ai-era-6d3fc8f26dd3?postPublishedType=initial&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Causal Lineage, Session Replay, and the Missing Layer in Enterprise AI</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Tue, 08 Sep 2026 11:21:31 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/causal-lineage-session-replay-and-the-missing-layer-in-enterprise-ai-1ac8</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/causal-lineage-session-replay-and-the-missing-layer-in-enterprise-ai-1ac8</guid>
      <description>&lt;p&gt;Why the next wave of AI adoption in Europe depends on infrastructure that can explain itself&lt;/p&gt;

&lt;p&gt;Zizka DB — EU AI Adoption&lt;br&gt;
For the last two years, the conversation around enterprise AI has been dominated by capability. Bigger models, longer context windows, more agentic workflows. What has lagged behind is a much less glamorous but far more consequential category: the infrastructure that lets a business actually understand what its AI systems did and why. In Europe in particular, that gap is not a nice to have. It is the reason pilots stall before they reach production.&lt;/p&gt;

&lt;p&gt;This is the problem ZizkaDB was built to solve, and it is worth walking through why each of its core capabilities, causal lineage, session replay, and drift detection, maps so directly onto what European enterprises need before they will trust an AI system with anything that matters.&lt;/p&gt;

&lt;p&gt;The limits of a trace&lt;br&gt;
Most observability tools for AI agents give you a trace. A trace is a record of events: a call went out, a response came back, latency was this many milliseconds, tokens cost this much. That is useful for debugging performance problems. It is close to useless for answering the question that actually matters after something goes wrong: why did the agent decide to do this particular thing, at this particular moment, given what it knew.&lt;/p&gt;

&lt;p&gt;A trace shows you the sequence. It does not show you the reasoning chain, the state the agent was actually operating on, or how one decision fed into the next. When a customer gets the wrong answer, or a compliance officer asks how a specific output was produced, a pile of spans does not answer that. Someone has to reconstruct the story by hand, usually under time pressure, usually without confidence that the reconstruction is even accurate.&lt;/p&gt;

&lt;p&gt;Causal lineage: from what happened to why it happened&lt;br&gt;
This is the core idea behind causal lineage, and it is the feature that separates ZizkaDB from a conventional tracing tool. Instead of storing flat, disconnected events, ZizkaDB links each decision an agent makes back to the specific inputs, tool calls, and prior state that produced it. Every event carries a parent, so you can walk backward through an agent’s behavior the same way you would walk backward through a chain of custody.&lt;/p&gt;

&lt;p&gt;For a European enterprise, this is not an engineering convenience. It is closer to a legal requirement in spirit, even where it is not yet one in letter. Under frameworks like the EU AI Act and existing sector rules in finance and healthcare, organizations are expected to be able to explain automated decisions, not just log that they happened. Causal lineage turns explainability from a manual, after the fact investigation into something the system produces as a byproduct of normal operation. When a regulator or an internal auditor asks how a decision was reached, the answer already exists. Nobody has to reconstruct it from memory or guesswork.&lt;/p&gt;

&lt;p&gt;Session replay: turning an incident into evidence&lt;br&gt;
The second piece is session replay, and its value becomes obvious the moment something goes wrong in production. Without it, an incident review usually looks the same way everywhere: a user reports a bad answer, someone pulls up logs that show fragments of activity, and the team spends hours trying to piece together what the agent actually knew at the time it acted. Often the honest conclusion is that nobody can say for certain.&lt;/p&gt;

&lt;p&gt;Session replay removes that uncertainty. It lets a team reconstruct an agent’s session end to end and see exactly what the agent had access to, what it retrieved, and what path it took to its final output. That is the difference between saying we believe this is what happened and saying here is exactly what happened, and here is the evidence.&lt;/p&gt;

&lt;p&gt;For enterprises operating under audit obligations, this is the raw material that compliance is actually built from. An audit trail is not a dashboard. It is a reconstructable record that can be handed to a third party and stand on its own. Session replay is what makes that possible for AI systems in the same way transaction logs make it possible for financial systems.&lt;/p&gt;

&lt;p&gt;Drift detection: an honest response to a probabilistic technology&lt;br&gt;
The third capability addresses a problem that is more cultural than technical. A lot of AI vendors, particularly in vertical markets like legal and insurance, have tried to win over cautious European buyers by promising consistency: the system will always behave this way, the accuracy is guaranteed. That promise does not hold up, because the underlying models are probabilistic. Behavior shifts when a prompt changes, when a model is upgraded, when the data the agent is retrieving changes shape.&lt;/p&gt;

&lt;p&gt;ZizkaDB does not try to paper over that reality. Its drift detection continuously compares current agent behavior against an established baseline and flags meaningful shifts, so a change in behavior is caught and documented rather than discovered by a customer weeks later. This is a more honest position than promising determinism, and it is also a more useful one. It gives a business the ability to say, with evidence, that it monitors for behavioral change and responds to it, which is a far stronger compliance posture than a marketing claim of guaranteed accuracy that nobody can actually back up.&lt;/p&gt;

&lt;p&gt;Why self-hosting matters as much as the features&lt;br&gt;
There is one more piece that is easy to overlook if you are only thinking about features: deployment model. ZizkaDB is available as a fully open source, self-hosted stack, not only as a managed cloud product. For a European enterprise, this matters for reasons that have nothing to do with cost. It means audit trails, decision lineage, and session data can live entirely inside the organization’s own infrastructure, under its own data residency and access controls, rather than inside a third party’s cloud in a jurisdiction the compliance team has to separately evaluate.&lt;/p&gt;

&lt;p&gt;Combined with causal lineage and session replay, this turns ZizkaDB into something closer to a compliance foundation than a debugging tool. The enterprise keeps full control of the evidence its AI systems generate, which is precisely the condition European regulators and internal governance teams have been asking for.&lt;/p&gt;

&lt;p&gt;The real barrier was never AI itself&lt;br&gt;
None of this is really about making AI smarter. It is about making AI systems accountable in a way that matches how European institutions already operate. Enterprises here are not rejecting AI because they doubt its usefulness. They are withholding trust from systems that cannot explain themselves, cannot be reconstructed after the fact, and cannot demonstrate that someone is watching for when their behavior changes.&lt;/p&gt;

&lt;p&gt;Causal lineage answers why a decision was made. Session replay proves what actually happened. Drift detection shows that behavior is being watched honestly rather than promised away. Together, these are not add on features. They are the missing layer that turns an AI pilot into something a European enterprise can actually put its name behind.&lt;/p&gt;

&lt;p&gt;That is the layer ZizkaDB is building. And it is likely to be the layer that determines which AI vendors actually succeed in this market, and which ones stay stuck in procurement.&lt;/p&gt;

&lt;p&gt;If your team is running production agents and cannot yet answer the question why did it do that, that gap is worth closing before the next incident forces the question.&lt;/p&gt;

&lt;p&gt;The article is originally published in Medium and can be viewed here:&lt;a href="https://medium.com/@MirArshadTalpur/causal-lineage-session-replay-and-the-missing-layer-in-enterprise-ai-c1ba194f4f09?postPublishedType=initial" rel="noopener noreferrer"&gt;https://medium.com/@MirArshadTalpur/causal-lineage-session-replay-and-the-missing-layer-in-enterprise-ai-c1ba194f4f09?postPublishedType=initial&lt;/a&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>euaiact</category>
      <category>europe</category>
    </item>
    <item>
      <title>Europe Doesn’t Have an AI Problem. It Has a Trust Problem.</title>
      <dc:creator>Mir Arshad Ali Talpur</dc:creator>
      <pubDate>Wed, 02 Sep 2026 15:01:36 +0000</pubDate>
      <link>https://dev.to/mir_arshadalitalpur_1b3/europe-doesnt-have-an-ai-problem-it-has-a-trust-problem-3pc9</link>
      <guid>https://dev.to/mir_arshadalitalpur_1b3/europe-doesnt-have-an-ai-problem-it-has-a-trust-problem-3pc9</guid>
      <description>&lt;p&gt;Europe has no shortage of AI talent, research output, or ambition. What it has struggled with is turning pilots into production. While American and Asian enterprises push agentic AI into core workflows, European boards keep AI projects in a permanent proof-of-concept holding pattern. The irony is that the regulation many blame for this, the EU AI Act, may end up being the thing that finally gives European enterprises the confidence to industrialize AI at scale. This is also the bet a young infrastructure company, Zizka AI, is making with its product ZizkaDB.&lt;/p&gt;

&lt;p&gt;Why Enterprise AI Adoption Has Been Slow in the EU&lt;br&gt;
The gap isn’t a lack of appetite. Surveys show the opposite: the majority of CEOs list accelerating AI among their top three priorities, and most large organizations already use AI in at least one business function. The blockage sits between experimentation and deployment, and it comes down to a handful of compounding factors.&lt;/p&gt;

&lt;p&gt;Regulatory uncertainty, not regulation itself, is the real drag. The AI Act has been amended repeatedly over the past two years, most recently through the Digital Omnibus on AI, which pushed back high-risk compliance deadlines and reshaped documentation and conformity-assessment requirements. Formal political agreement was reached in May 2026, endorsed by Parliament in June, and the Omnibus entered into force in July 2026, but businesses were told throughout that the August 2026 deadline remained legally live until publication. For a legal and compliance team, that kind of moving target is worse than a strict rule: it’s very hard to build a governance program against a deadline that keeps shifting underneath you.&lt;/p&gt;

&lt;p&gt;Compliance is perceived as a cost center, not infrastructure. Legal teams see documentation, conformity assessments, and registration obligations; engineering teams see a checklist bolted onto a system that was never built to be inspected. Most agent stacks generate logs, not evidence. When something goes wrong, a chatbot gives a wrong refund policy, an agent skips a required approval step, teams can see that it happened but not why, which is precisely the kind of causal explanation regulators, auditors, and increasingly customers expect.&lt;/p&gt;

&lt;p&gt;Fragmentation across 27 member states adds friction on top of the EU-level rules. Even as Brussels tries to harmonize standards, national implementation, sectoral overlaps (like the Machinery Regulation), and inconsistent guidance from standardization bodies have left many enterprises unsure which rulebook applies to their specific use case, especially for anything touching finance, health, energy, or HR.&lt;/p&gt;

&lt;p&gt;Risk aversion is rational, not just cultural. EY has found that most C-suite leaders now rank AI regulatory non-compliance as their top AI-related risk, ahead of cost or performance concerns. In a market where a poorly governed AI system can trigger real penalties, waiting and watching is often the economically sensible move, even if it’s strategically costly.&lt;/p&gt;

&lt;p&gt;Put together: European enterprises aren’t AI-skeptical. They’re evidence-poor. They lack the operational tooling to prove, to a regulator or to themselves, what their AI systems actually did and why.&lt;/p&gt;

&lt;p&gt;How the AI Act Could Actually Industrialize AI, Not Just Constrain It&lt;br&gt;
It’s easy to read the AI Act purely as a brake. But regulation has industrialized entire sectors before by doing something enterprises can’t do for themselves: creating a common, enforceable definition of trustworthiness. Pharmaceutical manufacturing didn’t scale despite GMP standards, it scaled because of them, since a shared bar of evidence let hospitals, insurers, and regulators trust products from companies they’d never audited directly.&lt;/p&gt;

&lt;p&gt;The AI Act is starting to play a similar role, and the recent Omnibus amendments push it further in that direction rather than away from it:&lt;/p&gt;

&lt;p&gt;Longer, staged runways for high-risk systems (with separate fixed deadlines for different high-risk categories) give enterprises time to build governance as real infrastructure instead of a rushed compliance sprint.&lt;br&gt;
Expanded regulatory sandboxes, including an EU-level sandbox, let companies test AI systems, including agentic ones, in supervised real-world conditions rather than guessing at what compliant looks like.&lt;br&gt;
Simplified obligations extended from SMEs to small mid-caps widen the on-ramp beyond just the largest players, which matters because Europe’s AI economy is disproportionately mid-market.&lt;br&gt;
A single horizontal framework, however imperfect in its rollout, is still a more industrializable target than 27 divergent national approaches. It’s the difference between building one audit pipeline and building twenty-seven.&lt;br&gt;
The common thread across almost every one of these provisions, transparency, human oversight, technical documentation, post-market monitoring, is that they all reduce to the same underlying requirement: an enterprise has to be able to reconstruct and explain what its AI system did. That’s not a legal nicety. It’s an engineering problem. And it’s the same problem that separates AI demos from AI that a bank, hospital, or logistics operator can actually run in production.&lt;/p&gt;

&lt;p&gt;Where ZizkaDB Comes In&lt;br&gt;
This is the layer Zizka AI is building for with ZizkaDB, an open-source (AGPL-3.0) and managed-cloud operational database purpose-built for production AI agents. Rather than treating logging as an afterthought bolted onto an agent framework, ZizkaDB is designed around the questions the AI Act effectively forces every deployer to answer:&lt;/p&gt;

&lt;p&gt;Causal lineage, not just logs. ZizkaDB records each agent step with an explicit parent-child chain, so a team can call a why() query and walk backward from a bad outcome, a wrong answer, a skipped policy check, to the exact user message, tool call, or context state that caused it. That’s the difference between having logs somewhere and being able to produce, on demand, the decision trail an auditor or regulator would ask for.&lt;br&gt;
Time-travel state snapshots. Instead of only knowing an agent misbehaved, teams can isolate the exact context window the agent was working from at that moment, which is essential for diagnosing drift after a prompt or model change, one of the more common failure modes in production agent fleets.&lt;br&gt;
Drift baselines and behavioral alerts. ZizkaDB compares live agent behavior against a baseline and flags regressions, giving teams a way to catch a silently-changed system before customers do, directly supporting the kind of post-market monitoring the AI Act expects from providers and deployers of higher-risk systems.&lt;br&gt;
Human-in-the-loop interception points, so governance can trigger before a risky action executes rather than being reconstructed afterward from a postmortem.&lt;br&gt;
Deployment flexibility that matches EU procurement realities. Self-hosted open source for organizations that need data sovereignty, or managed cloud for teams that want to move faster, matters in a region where data residency and sovereignty concerns are their own adoption blocker, separate from the AI Act itself.&lt;br&gt;
None of this is marketed as AI Act compliance software, and it shouldn’t be reduced to that. It’s infrastructure for making agents debuggable and reliable, full stop. But the overlap with what the Act increasingly demands is not a coincidence. As Zizka AI’s own team has put it, the industry has spent enormous engineering effort trying to force fundamentally probabilistic models to behave like deterministic APIs. That effort is largely wasted; models will keep drifting as weights update and context accumulates. The more tractable goal is what they call constrained reliability through auditability: accepting that AI systems are probabilistic, and building the operational tooling to observe, explain, and govern them anyway.&lt;/p&gt;

&lt;p&gt;The Bigger Picture&lt;br&gt;
Regulation and industrialization aren’t opposites, they’ve historically been sequential. The AI Act’s early years looked chaotic because Europe was trying to write the rulebook and build the technology at the same time, with each amendment cycle adding uncertainty rather than removing it. But as the framework stabilizes, longer timelines, clearer sandboxes, harmonized obligations, the enterprises that win won’t be the ones that lobbied hardest against it. They’ll be the ones that already built the operational plumbing to answer why their AI did what it did before anyone asked.&lt;/p&gt;

&lt;p&gt;That’s the bet behind tools like ZizkaDB: that in Europe, the fastest path to enterprise-scale AI isn’t around the AI Act, but through it, with an audit trail as the foundation, not an afterthought.&lt;/p&gt;

&lt;p&gt;This article reflects the regulatory landscape as of early September 2026. The AI Act’s implementation timeline continues to evolve; readers should confirm current deadlines via the European Commission’s AI Act Service Desk before making compliance decisions.&lt;/p&gt;

&lt;p&gt;The article is originally published in medium and can be accessed via this link :&lt;a href="https://medium.com/@MirArshadTalpur/europe-doesnt-have-an-ai-problem-it-has-a-trust-problem-c0ee81ad7bef" rel="noopener noreferrer"&gt;https://medium.com/@MirArshadTalpur/europe-doesnt-have-an-ai-problem-it-has-a-trust-problem-c0ee81ad7bef&lt;/a&gt;&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
