DEV Community

Cover image for The at() Function: Reconstruct What Your Agent Looked Like on Any Given Day
Mir Arshad Ali Talpur
Mir Arshad Ali Talpur

Posted on

The at() Function: Reconstruct What Your Agent Looked Like on Any Given Day

This is the fifth article in the series where I go through each function ZizkaDB offers and explain what it means in real business terms. So far we have covered Causal Lineage, why an agent did something, Session Replay, what happened in one session, Behavioral Drift, whether your agent has changed, and forget(), how you erase a user’s data cleanly. This one covers at(), how you reconstruct what your entire agent system looked like at any specific moment in the past, not just one session, the whole configuration.

It is easy to confuse this with Session Replay, so it is worth being precise about the difference. Replay reconstructs one session, the conversation and actions that happened within it. at() reconstructs the system itself, which prompt version was live, which tools were available, which model you were calling, what your routing logic looked like, as of a specific timestamp. Replay answers what happened in this conversation. at() answers what was true about the agent on that day. You often need both together, because a session only makes sense in the context of the system that produced it.

What at() is
Agent systems change constantly. Prompts get rewritten, tools get added or deprecated, models get upgraded, routing logic gets adjusted, thresholds get tuned. None of this is usually version controlled the way application code is, and even when it is, the live production configuration at a given moment is not always obvious six weeks later just by looking at a git log. An at() query lets you ask the system directly, what was the state of this agent at this timestamp, and get back the actual configuration that was running then, not your best guess reconstructed from commit history and deployment logs.

This matters most in two situations. The first is incident response, when a session from three weeks ago behaved strangely and you need to know whether that was caused by the prompt in use at the time, a tool that has since been replaced, or a model version that has since been upgraded. Without at(), you are cross referencing deploy logs, Slack messages, and memory to guess what was live on that date. The second is regulatory evidence, where the question an auditor asks is never what does your agent do now, it is what did your agent do on the date in question, and being able to answer that precisely is the difference between a satisfying answer and an educated guess.

Consider a concrete case. A customer disputes a decision your agent made two months ago. Your prompt has been updated four times since then, one of your tools has been swapped for a faster provider, and the model has been upgraded once. Without at(), reconstructing what the agent actually looked like on the day of that decision means digging through deployment history across several systems and hoping nothing was missed. With at(), you query that timestamp and get back the exact configuration, the prompt version, the tools, the model, as it existed that day. You are not reasoning about what probably changed since then, you are looking at what was actually true.

Where you see benefits in the first 30 days
Old incidents stop being unsolvable. When a session from weeks or months ago needs investigating, you are not forced to guess whether something that has since changed was the cause. You pull the exact system state from that date and check directly.

Regulatory evidence requests get a precise answer instead of an approximate one. Articles 12 and 26 of the EU AI Act, and most audit frameworks generally, ask what your system did at a point in time, not what it does now. at() turns that into a direct query rather than a reconstruction project.

Rollback decisions get easier to evaluate. When you are deciding whether to revert a recent change, you can compare the current configuration against the one from before the change directly, rather than relying on memory of what the diff was supposed to do.

Postmortems get more accurate. A common failure in incident reviews is debugging against the current system instead of the system as it was when the incident happened, because by the time anyone looks, several things have already changed. at() removes that blind spot.

Due diligence in acquisitions or partnerships moves faster. When a partner or acquirer asks how your agent behaved during a specific period, you can show them directly instead of reconstructing it from scattered records.

Modeled numbers
These are estimates built from assumptions, not measured customer results. Swap in your own figures and the math still holds.

Take investigating older incidents first. Assume a team handles 8 incidents a month involving sessions more than a week old, at 100 euro per engineering hour. Without at(), tracing the historical configuration for each one takes about 5 hours, costing 4,000 euro a month in total. With at(), the same lookup takes about 30 minutes per incident, bringing the monthly cost down to 400 euro.

Regulatory or audit evidence requests follow a similar pattern, though they happen far less often, around 4 times a year, usually involving 2 engineers at 100 euro per hour. Reconstructing historical system state manually takes about 4 days per request, costing 6,400 euro each time. With at(), the same request takes about 4 hours, costing 800 euro.

Postmortem accuracy rework is a smaller but telling case. Assume a team runs 20 postmortems a month, and 1 in 5 is later found to have been based on the current configuration rather than the historical one, requiring about 3 hours of rework at 100 euro per hour. That works out to 4 reworked postmortems a month, costing 1,200 euro. With at(), the correct historical state is pulled the first time, so this rework cost drops to close to zero.

Put together, that is roughly 4,800 euro a month in saved investigation and rework time for a team with a reasonable volume of historical debugging, plus a meaningful reduction in time spent on audit and regulatory evidence requests when they come up. The audit line is intermittent but expensive when it hits, a single poorly prepared evidence request during due diligence or a regulatory inquiry tends to cost far more than a year of simply having the capability on hand.

Why it matters beyond savings
Every agent system accumulates a history of small, mostly undocumented changes, a prompt tweak here, a tool swap there, a threshold adjusted during an incident and never written down. Over months, the gap between what your records say was running and what was actually running tends to grow. at() closes that gap by making the actual historical state queryable rather than reconstructed from memory and documentation that is usually incomplete. For a vertical AI company, that turns “we believe this is what the system looked like then” into “here is what the system looked like then,” which matters every time you debug an old incident, answer an auditor, or need to explain a decision you made months ago with confidence instead of a best guess.

If you tell me your vertical and how often your agent configuration changes, I can tailor these numbers to your case.

Want to test ZizkaDB on your agent? Try our open source version here: https://github.com/ZIZKA-AI-SL/ZizkaDB

Interested in a design partnership? Fill in the form on our site or reach me directly at founder@zizka.ai

The article was originally published in Medium
can be viewed here : https://medium.com/@MirArshadTalpur/the-at-function-reconstruct-what-your-agent-looked-like-on-any-given-day-04ce74acc2f3?postPublishedType=repub

Top comments (1)

Collapse
 
arhancanli profile image
Arhan Canli •

The distinction from Session Replay is useful. "What was true about the agent on that day" is the question every incident review ends up asking.

One subtlety that matters once you have corrections: there are two different "as of" questions. What configuration was actually live at time T (what the system knew then), versus what we now believe was live at T after fixing a mislogged deploy. Databases call this bitemporal: valid time and recorded time. We hit the same thing with SEC financial data, where a company's revenue "as of" a date and the later restated figure are both correct answers to different questions, and mixing them up quietly breaks backtests.

Does at() keep both? If someone corrects a record about what was deployed last month, can you still query what the system believed last month?