<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Sarit Chauhan</title>
    <description>The latest articles on DEV Community by Sarit Chauhan (@sarit_chauhan).</description>
    <link>https://dev.to/sarit_chauhan</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118589%2Fd3a34088-7c49-471a-8479-2308eb8cdf5b.jpg</url>
      <title>DEV Community: Sarit Chauhan</title>
      <link>https://dev.to/sarit_chauhan</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/sarit_chauhan"/>
    <language>en</language>
    <item>
      <title>Letting an LLM write alerting rules, but never letting it flip the switch</title>
      <dc:creator>Sarit Chauhan</dc:creator>
      <pubDate>Fri, 18 Sep 2026 16:05:14 +0000</pubDate>
      <link>https://dev.to/sarit_chauhan/letting-an-llm-write-alerting-rules-but-never-letting-it-flip-the-switch-pjk</link>
      <guid>https://dev.to/sarit_chauhan/letting-an-llm-write-alerting-rules-but-never-letting-it-flip-the-switch-pjk</guid>
      <description>&lt;p&gt;This is the platform behind &lt;a href="https://lnkd.in/p/dWp56Xfy" rel="noopener noreferrer"&gt;taabi Nexus&lt;/a&gt;, which Taabi Mobility launched this week. The launch video has three real calls it made into trucks. This is the engineering side.&lt;/p&gt;

&lt;p&gt;Last Wednesday a truck driver heard his dashcam intercom tell him, in Hindi, that his seat belt was off. He said yes, he could hear. It asked if he was wearing it now. He pulled the belt across while answering. Nobody had dialled that call. A fleet manager had typed a sentence into a text box a few days earlier, and that sentence did.&lt;/p&gt;

&lt;p&gt;I've spent the last few weeks building the thing in between. Dashcams raise alerts: drowsiness, phone use, seat belt, hard braking, about forty kinds. A manager writes what should happen, in English. The platform messages, waits, re-checks, calls the truck, and writes down the follow-ups. Here is what I'd tell someone starting the same build.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model writes a document. Code runs it.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My first instinct was to let the model read the sentence and act. That's a chatbot with side effects, and nobody is going to let one ring their drivers. So the LLM's only job became translating the sentence into a rule document with six keys, and a plain engine executes that:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;{&lt;br&gt;
  "trigger":   { "event_types": ["fatigueWarn"] },&lt;br&gt;
  "where":     { "field": "speed_kmph", "op": "gt", "value": 60 },&lt;br&gt;
  "aggregate": { "window": { "type": "sliding", "size": "PT10M" }, "group_by": ["vehicle_no"],&lt;br&gt;
                 "function": { "name": "count" }, "having": { "op": "gte", "value": 2 } },&lt;br&gt;
  "actions":   [ { "action": { "type": "notify", "channel": "telegram", "to": { "role": "manager" } } },&lt;br&gt;
                 { "action": { "type": "call", "to": { "selector": "driver" } }, "wait": "PT10M",&lt;br&gt;
                   "recheck": { "event_types": ["fatigueWarn"], "same": ["vehicle_no"] } } ],&lt;br&gt;
  "throttle":  { "cooldown": "PT1H", "dedup_key": "{vehicle_no}" }&lt;br&gt;
}&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;That is "drowsy above 60 km/h twice in ten minutes, tell the manager; still drowsy ten minutes later, call the truck". The schema is a Pydantic v2 model, exported as JSON Schema and as TypeScript types, so the engine, the simulator, the UI's plain-words summary and the evals all point at the same thing. Freezing it early was the best decision in the project, and the one I almost skipped because it felt like premature ceremony.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where the flow stops&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The authoring agent is a LangGraph graph with two interrupts. If the sentence is missing a threshold or a channel, the graph stops and asks. It doesn't guess. Once a draft validates (schema plus about thirty semantic checks, written in code, with the errors fed back to the model for up to three repairs), it is replayed over the tenant's last seven days and comes back with a number: "would have fired 27 times, 58 held back by the throttle". That number is what turns "sounds right" into "yes". It has also caught a rule that would have fired four thousand times.&lt;/p&gt;

&lt;p&gt;Then the second interrupt: a person approves, and approval only saves a draft. Activating is a separate button in a separate service behind a role check. The agent has no tool that can press it. The same shape carries into the calls: a call step is placed by the notifier, never by a model, uncertain calls wait in a queue for a manager, and every call ends with follow-ups a person closes. I stopped calling this "human in the loop" in meetings and started saying "this is where the state machine stops", which is what it actually is.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Two bugs I'd rather not have found in production&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;"At most one call per vehicle per hour" came out of the composer as a cap per rule, which in my schema means one call per hour for the entire fleet. The schema was fine. The problem was that the summary the manager approves said "at most one call per hour" and let both readings through. Now it says exactly what the schema means, because the summary is the thing people actually read.&lt;/p&gt;

&lt;p&gt;The other one was latency. A yes-or-no decision inside a live phone call went through the full agent subprocess and took 18 to 30 seconds, which is most of the call. One direct Messages API call with a tight schema took 1.2 seconds. Heavy machinery where you need tools and memory; not inside a call.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What real trucks taught me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Tests and a simulated intercom get you to the door. On the other side: engine noise under every word the driver says, Hindi speech recognition that is fine on the agent's own lines and shaky on the driver's, and no way for the agent to see that the belt actually went on. So it asks yes-or-no questions, treats an unclear answer as "not confirmed", and re-checks the alert stream after the call instead of trusting the transcript. Budget more time for the microphone than for the model. I did not, and I would now.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The boring parts that let me sleep&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Every table carries tenant_id and Postgres row-level security is forced, so a forgotten WHERE returns nothing rather than everything. The tenant comes from the JWT, never from a request body or the model, and the MCP server the agent reads through takes a per-tenant token where the token is the tenant; there is no argument for the model to fill with someone else's id.&lt;/p&gt;

&lt;p&gt;Every alert gets a trace_id at ingest that rides in a Kafka header and is bound into every log line downstream; with OpenTelemetry on, it's the trace id too. Someone asks what happened to an alert, I paste one id into Grafana and see it accepted, persisted, matched and delivered across four services. Under load, 96,000 alerts went through with ingest p99 at 94 ms and nothing in the dead-letter queue, and when I killed the persist consumer on purpose an 86,704-message backlog drained in under five minutes with nothing lost. I wrote those numbers into the service contracts so they would be argued with, not remembered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Stack&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Python 3.12 and FastAPI in a uv workspace, Pydantic v2 for the contracts, Kafka through aiokafka (a compacted topic for rule updates, so an engine restart rebuilds its rules from Kafka alone), Postgres 16 with RLS and Alembic, Redis for dedup and cooldowns. LangGraph for the agents, Claude via the Agent SDK, a custom MCP server of curated read-only tools, Langfuse for traces, prompts and the eval datasets, with pytest golden suites failing the build. OpenTelemetry to Tempo, Prometheus rules with promtool tests, Grafana provisioned from code. React 18, Vite and TanStack on the front, Playwright end to end. Docker Compose for now; the rules compile to an IR so the engine can move to Flink without the rules changing.&lt;/p&gt;

&lt;p&gt;If you've shipped natural-language rules or agents with a person in the path, I'd genuinely like to know where you put the switch. Mine is two buttons and a queue. &lt;/p&gt;

&lt;p&gt;Much more to come next. &lt;/p&gt;

</description>
      <category>llm</category>
      <category>architecture</category>
      <category>agents</category>
      <category>automation</category>
    </item>
    <item>
      <title>but.. but they did this in one week</title>
      <dc:creator>Sarit Chauhan</dc:creator>
      <pubDate>Thu, 10 Sep 2026 06:11:07 +0000</pubDate>
      <link>https://dev.to/sarit_chauhan/but-but-they-did-this-in-one-week-3cfm</link>
      <guid>https://dev.to/sarit_chauhan/but-but-they-did-this-in-one-week-3cfm</guid>
      <description>&lt;p&gt;i wrote this while claude wrote the code.&lt;/p&gt;

&lt;p&gt;its 3am. fix agent is running for the third time, i have been waiting on it more than an hour now. no progress bar no ETA nothing, just the spinner. and i am sitting here with nothing to do because there is literally nothing for me to do till it finishes.&lt;/p&gt;

&lt;p&gt;i have been reading this community for long time, silent reader, never posted. this is the first one. i think people here will understand this better than anyone around me right now because everyone else only sees the output. "you shipped this in one week??" yes. we did. but i am not sure anymore who is "we".&lt;/p&gt;

&lt;p&gt;remember the whole thing. robots will do the hard work, humans will do the art.&lt;/p&gt;

&lt;p&gt;ok. robot is doing the hard work. and i guess this is my art, a blog post at 3am while the machine writes the thing i used to write.&lt;/p&gt;

&lt;p&gt;the project - its an AI control tower for a logistics client. it keeps drivers and vehicle owners notified proactively, tracks the trips, and one of the main features is it detects when driver is getting drowsy and alerts him so he doesnt sleep on the wheel. accidents prevention basically. good product, i actually like it, thats not the issue. and the code is good also. no doubt about it, honestly better than what i would write at this hour.&lt;/p&gt;

&lt;p&gt;but its not my code. thats the issue.&lt;/p&gt;

&lt;p&gt;llm was supposed to help devs. and it does, i am not going to lie about that. but what happened is the bar got raised. now the expectation is one week for what used to be 2 months. and whatever work is left for me is the most boring part of the whole thing and i am doing that for even longer hours than before.&lt;/p&gt;

&lt;p&gt;boring meaning, no thrill. no problem to solve. only thrill left is solving the business problem, and ya that was always the point i know that. but earlier we had small wins on the way. big bugs that keep you awake because you want to be awake. fixing the unfixable. that feeling when you finally see it and go ohhhh and everyone in the room knows. thats mostly gone now and i think we will never have it again like before.&lt;/p&gt;

&lt;p&gt;when was the last time you even read the error properly. like actually read the whole trace and went and fixed it yourself. for me its been a while. there is no time for that, we are operating on top of code where we have not read or understood 90% of it. we are not devs on these projects, we are like supervisors of something we dont fully understand.&lt;/p&gt;

&lt;p&gt;and then the 5 hr limit.&lt;/p&gt;

&lt;p&gt;for complex stuff you basically have to plan your sleep around when the limit resets. run a phase, hit the wall, wait, run next phase. i ended up purchasing another premium with my other id just to keep working. then also i took a friend's account for a day. i think a lot of people are handling 2-3 subscriptions right now quietly and its exactly like managing multiple resources on a team, except the resources go to sleep on a fixed schedule and you dont.&lt;/p&gt;

&lt;p&gt;and the agent stalls. one hour into the run. no error no timeout, just stopped. and you are sitting there saying just get it done bro. to a text box. at 3am.&lt;/p&gt;

&lt;p&gt;never felt this useless in the process of coding before. not as a fresher, not on my worst day. useless is the word. the thing got built, mostly by claude, got deployed, launched to the customer, real drivers are getting alerts from it. and i am living in fear waiting for it to break because i know when it breaks my only option is to go and beg claude again.&lt;/p&gt;

&lt;p&gt;ok i dont want this to be only a rant (may be it is). some things that actually helped in this week, if you are in the same place -&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;p&gt;make phase wise plans before you write a single prompt. not "build the app". phase 1 auth, phase 2 data model, phase 3 apis like that. each phase small enough to finish in one session. this alone saved a lot of the 3am pain.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;write handover notes after every phase. what is done, what is half done, what is broken, what next phase needs to know. i keep it in a plain md file in the repo. when context resets or limit hits or you switch account you just paste the notes and continue from where you left instead of explaining everything again.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;read the error yourself first. even if you are going to give it to the agent anyway. 2 minutes, actually read it. half the time you will see it. other half at least you know what to ask instead of "fix it" for the third time.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;debugging is the real skill now. not typing code. if you cant debug then the first time something breaks in prod you are fully at the mercy of the model. protect this skill, use it on purpose even when you dont need to.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;know your reset time. sounds stupid but plan the heavy phases right after reset and do the reading and planning while you are locked out.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;keep at least one part of the codebase that you understand end to end. for me its the deployment and the db layer. if everything is on fire atleast i have somewhere to stand.&lt;/p&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;so is this it? i honestly dont know. client is happy. thing works. by every outside measure this was a big win and i should be feeling great.&lt;/p&gt;

&lt;p&gt;but i keep coming back to - what was my contribution. better prompts? staying awake? knowing which account still has hours left?&lt;/p&gt;

&lt;p&gt;may be this is just what the job is now and i have to make peace with it. may be the thrill comes back in some other form which i cant see right now. may be in 2 years this post will read like someone crying about compilers.&lt;/p&gt;

&lt;p&gt;or may be lot of us are feeling this and not saying it because the output looks so good that complaining feels ungrateful.&lt;/p&gt;

&lt;p&gt;anyone here also did the same "one week" thing recently? what's your view?&lt;/p&gt;

&lt;p&gt;agent finished btw. 3:40. going to sleep.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>career</category>
      <category>discuss</category>
      <category>productivity</category>
    </item>
  </channel>
</rss>
