At 02:01 one night, Telegram got a message from my personal portal:
🟢 You can go for 95-octane
It pointed to the right station and said there was no line. About an hour later I decided to use it. Home automation logged the car leaving at 03:07 and coming back at 03:30: a 23-minute trip.
The real notification and the two home-automation messages that followed (in Russian).
The useful part isn't that the recommendation happened to be right. It's that until that moment I didn't have to watch anything. Getting there took one wrong turn: the coding agent quickly built a working MVP, and I then realized I had automated the wrong part of the problem.
The problem
A fuel shortage turned filling up into a small operational process. There are a few stations near me worth driving to, I only need 95-octane, and "fuel showed up" means little if there's a line of dozens of cars. Public services already reported availability and lines, but every check still meant: open the service, find my stations, check 95 specifically, look at lines, judge how fresh the data is, compare, decide. Then repeat a while later.
The data existed. What was missing was a personal layer that knew my criteria and did the routine part.
MVP: a few curl calls, a problem description, a template
I gave a coding agent (I don't remember which one, and it doesn't matter here):
- a description of the problem;
- about 5–6
curlexamples that made the available API clear; - a requirement to use my existing FastAPI project template.
The agent worked out the data structure, wrote the project documentation, and implemented the MVP: station search, a watchlist, periodic data collection, history, and a screen with the current state of the selected stations.
In git, the first commit with the plan landed on September 12 at 09:24, and the MVP commit at 11:44. That is not "2 hours 20 minutes of human work": git doesn't show how much time a person spent or what happened inside the agent session. It does show how fast a raw input became a working product.
It worked, and answered the wrong question
Technically the MVP did exactly what I had asked: it collected data, showed my stations, gave a current snapshot. Looking at the whole screen, I still had to read several cards and interpret the situation myself. I had automated getting the data, not the part that annoyed me.
The first version answered "What's happening at my gas stations right now?" I needed "Should I go for 95-octane now, or wait?"
I took the working screen to ChatGPT, not as a coding task but as a product review: what decision should the user get on the main screen? One AI had acted as the implementer of the original spec; the second one looked at the result from outside. That review produced the idea of turning the monitor into a decision assistant.
Why one status field isn't enough
The naive version:
95-octane available -> go
95-octane not available -> don't go
It fails on real data. The source API has no formal public spec, and different parts of a response can describe the situation with different freshness and semantics. The overall status can say fuel is available while more recent data shows a line of 20–50 cars at the same station.
A decision has to account for:
- whether 95-octane specifically is confirmed;
- how fresh that confirmation is;
- what the line looks like, and how fresh that is;
- whether the current state can be trusted at all.
And one rule had to be stated explicitly:
No fresh data does not mean no fuel.
If the system doesn't know, it says UNKNOWN instead of pretending there's no fuel.
From snapshots to a deterministic decision
A normalization layer now sits between the external API and the decision:
API snapshot
↓
StationState
↓
StateTransition
↓
DecisionEngine
↓
GO / WAIT / NO_OPTIONS / UNKNOWN
StationState describes the situation in the app's own terms: is my fuel available, what's the line, when was it observed, how fresh is the data. DecisionEngine looks at all watched stations at once:
-
GO: at least one suitable station has confirmed 95-octane and an acceptable line; -
WAIT: fuel is available, but the options don't work for me, e.g. because of the line; -
NO_OPTIONS: no watched station has a confirmed suitable option; -
UNKNOWN: fresh data isn't enough for an honest decision.
This is a plain deterministic algorithm. There's no LLM at runtime. The answer had to be repeatable, depend on measurable conditions, be easy to diagnose, be explainable through the underlying facts, and behave the same after every restart. AI was used to build the system, review the interface, and design and implement changes; the runtime decision is plain rules.
AI helps build the decision system, but the decision system itself doesn't have to be an AI system.
Decision first, evidence after
The station cards stayed, but became secondary. The main element is the aggregate verdict with an explanation, for example GO followed by "95-octane confirmed: Gazpromneft, Primorskoe Highway 251 — no line."
Top: the decision across all stations. Below: the per-station data it's based on (UI in Russian).
The first version said: here are three cards, figure it out. The second says: you can go now, here's why, and the underlying data is below.
Each station also has a history by day, week or month, with an hour-by-hour timeline: 95 with no line, short line, long line, no 95, no data. There is no ML model predicting when a fuel truck arrives. The history lets a human spot patterns (what part of the day a station tends to get 95, when lines form), so the operational decision goes to the algorithm and the deeper interpretation stays with me.
The hardest part: teaching it to stay quiet
Once DecisionEngine existed, the system started understanding transitions: fuel appeared, the line changed, data recovered, the verdict changed. Every one of them is interesting technically. To the user, most aren't:
fuel appeared
↓
the line is still too long
↓
final verdict = WAIT
A Telegram message about that demands attention and changes nothing about what I do. The rule became:
I don't care about every transition and intermediate status. I only want to know when I can actually go.
Telegram stopped being a log of internal events and became a channel for actionable events only. The pipeline got a third step:
data
↓
decision
↓
is it worth interrupting the person right now?
The detailed log didn't go away; it moved inside the portal as an audit trail. For every Telegram message you can see the previous state, what changed, what event occurred, whether the overall verdict changed, why the portal decided to notify, and the exact text sent. If it tells me "go" at three in the morning, I can reconstruct afterwards exactly which change made it decide to bother me. With a rule-based system, that full cause-and-effect chain is available.
Who did what
The coding agent got a few curl calls, a problem description and a template, and built the first working service. After that, AI implemented the new state models, the decision engine, transitions, Telegram, history, audit and tests.
The most important changes didn't come from the code. The MVP worked, and only seeing it in a real interface showed it answered the wrong question. Then the working Telegram integration exposed that a correct event system can still be too chatty.
real problem → task for the agent → working MVP
→ human looks at the result → wrong thing was automated
→ external AI review → new product model
→ real-world use → one more simplification
I spent much less time implementing classes and more on: what problem the system actually solves, how to tell the first version isn't enough, which decision can be trusted to an algorithm, what data is sufficient, when the system may interrupt a person, and whether the result is useful in real life. Sometimes a working product is exactly what you need to see that the task should have been framed differently.
A thin personal layer over mass-market data
Existing services aren't bad; they show information to thousands of people with different routes, cars and tolerance for lines. They don't need to know that I care only about 95-octane, only a few stations, won't accept a certain line, and want to be called only when there's a practical reason to go.
someone else's data
+ my context
+ my rules
+ my history
+ my notifications
= a personal decision assistant
A service like this for one person used to look disproportionately expensive in time. With coding agents, that experiment costs something else entirely.
The best quality criterion I found for it isn't the number of screens or endpoints, or whether "AI" appears in the architecture. It's that most of the time the system asks nothing of me, and shows up exactly when there's something for me to do.
The full case study, including the first MVP screen, the station history and the Telegram audit log, is on my site: https://hram.github.io/en/articles/gdebenz/


Top comments (0)