<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Vaibhav Shukla</title>
    <description>The latest articles on DEV Community by Vaibhav Shukla (@vaibhav_shukla_20).</description>
    <link>https://dev.to/vaibhav_shukla_20</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3800008%2F34e17d59-5601-4b5d-a5e1-013b7f8b4cbd.png</url>
      <title>DEV Community: Vaibhav Shukla</title>
      <link>https://dev.to/vaibhav_shukla_20</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vaibhav_shukla_20"/>
    <language>en</language>
    <item>
      <title>Building an Adaptive Authorization Layer with SigNoz and OpenTelemetry</title>
      <dc:creator>Vaibhav Shukla</dc:creator>
      <pubDate>Sun, 26 Jul 2026 05:38:53 +0000</pubDate>
      <link>https://dev.to/vaibhav_shukla_20/building-an-autonomy-error-budget-gateway-with-signoz-and-opentelemetry-4ia3</link>
      <guid>https://dev.to/vaibhav_shukla_20/building-an-autonomy-error-budget-gateway-with-signoz-and-opentelemetry-4ia3</guid>
      <description>&lt;p&gt;A lot of agent implementations right now have static permissions. If an agent's allowed to run migrations, it stays allowed to run migrations, until a human happens to notice something's wrong.&lt;/p&gt;

&lt;p&gt;SRE teams don't run production like that. Burn through your error budget and deployments freeze automatically, nobody has to catch it manually. So for the Agents of SigNoz hackathon, I built LEASH to see what happens if an agent's permissions worked on a similar principle: retained through sustained healthy execution, and lost automatically the moment that breaks down.&lt;/p&gt;

&lt;p&gt;This runs in a simulated staging environment with failure injection I control. The goal was proving permissions can change mid-run, without a human in the loop, driven entirely by the alert conditions SigNoz evaluates from real telemetry.&lt;/p&gt;

&lt;p&gt;The idea&lt;/p&gt;

&lt;p&gt;LEASH is a broker that sits between the agent and every tool it's allowed to call, enforcing decisions triggered by alert conditions evaluated from SigNoz telemetry. Static permissions assume authorization doesn't need to adapt at runtime. LEASH treats an agent's permission tier as runtime state that changes based on observed behavior.&lt;/p&gt;

&lt;p&gt;The authorization policy itself is intentionally simple, a threshold and a window. The novelty isn't the policy, it's closing the feedback loop between observability and authorization.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3el0gn7ope0sf9d4zfk.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fa3el0gn7ope0sf9d4zfk.png" alt=" " width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;When failures cross a threshold, SigNoz fires an alert, hits a webhook on the broker, and the broker downgrades the agent's tier before its next privileged calls are authorized. LEASH effectively turns observability into a feedback controller for agent permissions. Observability is no longer just something you check after the damage is done, it decides what the agent's allowed to do while it's still running.&lt;/p&gt;

&lt;p&gt;I split this into four separate services on purpose. An agent shouldn't be able to touch the system deciding its own permissions, keeping them genuinely separate processes was the point, not an implementation detail. LEASH itself runs as a standalone HTTP service the agent calls through for every tool invocation, not a sidecar or proxy sitting transparently in the network path.&lt;/p&gt;

&lt;p&gt;Tiers&lt;br&gt;
T3 — read, write, destructive cleanup&lt;br&gt;
T2 — read and write, destructive calls blocked&lt;br&gt;
T1 — read-only&lt;/p&gt;

&lt;p&gt;Agent runs at T3 normally. The broker tracks the current consecutive-failure streak per tool and exports it as a metric, SigNoz's alert rule watches that. After 3 consecutive migration failures inside a 5 minute window, SigNoz fires, and in testing, the whole pipeline, failure to alert to webhook to tier downgrade, lands in under 10 seconds. Try a destructive call after that and you get a hard 403, with the exact trace ID of the failures that caused it attached as evidence.&lt;/p&gt;

&lt;p&gt;That trace ID isn't decoration. Paste it into SigNoz and you see the actual failed spans behind the decision. The agent isn't refused because a prompt told it to behave, it's refused because there's a real record sitting in SigNoz.&lt;/p&gt;

&lt;p&gt;Recovery isn't built yet, downgrade is one-directional within a run right now. Automatic recovery after a clean streak is the obvious next piece, and it'll need real hysteresis, not just a raw success count, or a fast recovering agent could flap between tiers on noisy data.&lt;/p&gt;

&lt;p&gt;Why OpenTelemetry specifically&lt;/p&gt;

&lt;p&gt;The broker doesn't have its own monitoring protocol. It just emits standard OTel telemetry, and SigNoz already understands it, no custom integration on either end. That's what let SigNoz's alert engine drive the authorization decision, instead of remaining something you'd only inspect after an incident.&lt;/p&gt;

&lt;p&gt;Getting there wasn't the policy logic, that part's genuinely simple. It was the webhook. Docs tell you SigNoz alerts can call one, they don't tell you the exact JSON shape you get for your specific alert type. Had to trigger a real failure, watch the alert actually fire, and read what landed on my endpoint before the handler worked right. Also learned a 60 second window just reacts to noise, 5 minutes with 3 consecutive failures is what turned it into a real signal.&lt;/p&gt;

&lt;p&gt;Scope&lt;/p&gt;

&lt;p&gt;LEASH isn't trying to figure out if an agent is aligned. It's only deciding whether the agent retains permission to keep doing higher-risk things.&lt;/p&gt;

&lt;p&gt;That's a real limit, worth being honest about. It reasons through execution health, tool failure rate, nothing more. An agent that repeatedly fails migrations loses privileges. An agent that succeeds at every call while doing the wrong thing doesn't, deleting the right table for the wrong reason still looks like a healthy 200 to this system. Closing that needs richer signals than pass or fail, future versions could incorporate policy violations or anomalous tool sequences.&lt;/p&gt;

&lt;p&gt;What I'd change next&lt;/p&gt;

&lt;p&gt;Tool risk tiers are hardcoded right now, a real deployment needs that configurable without redeploying the broker. There's one alert rule for one failure mode. The natural next step, generalize this so any policy-relevant metric crossing a threshold can demote any agent, plus an actual path back up with proper hysteresis.&lt;/p&gt;

&lt;p&gt;Takeaway&lt;/p&gt;

&lt;p&gt;A prompt telling an agent not to do destructive things after errors is a suggestion, not enforcement, the model can ignore it or misread it. What actually stops something is a system outside the model, reading real telemetry, blocking the call before it executes. That's what SigNoz's alerts and OpenTelemetry's traces gave me here, and a prompt never could.&lt;/p&gt;

&lt;p&gt;Code: github.com/Vaibhav13Shukla/LEASH&lt;/p&gt;

&lt;p&gt;Built for the Agents of SigNoz hackathon, Track 1.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>automation</category>
      <category>devops</category>
    </item>
    <item>
      <title>I Found 60 Accessibility Bugs on Google Support That No Tool Catches</title>
      <dc:creator>Vaibhav Shukla</dc:creator>
      <pubDate>Mon, 16 Mar 2026 17:12:59 +0000</pubDate>
      <link>https://dev.to/vaibhav_shukla_20/i-found-60-accessibility-bugs-on-google-support-that-no-tool-catches-bkg</link>
      <guid>https://dev.to/vaibhav_shukla_20/i-found-60-accessibility-bugs-on-google-support-that-no-tool-catches-bkg</guid>
      <description>&lt;p&gt;I broke Google Support last week. Or rather, I proved it was already broken.&lt;/p&gt;

&lt;p&gt;Not in English. In English it works fine. Lighthouse gives it a decent score, axe-core agrees, everyone moves on.&lt;/p&gt;

&lt;p&gt;I checked the Japanese version.&lt;/p&gt;

&lt;p&gt;The HTML tag said lang="en". On a page written entirely in Japanese. Every ARIA label — Search, Close, Main menu — was still in English. The search placeholder said "Describe your issue" in English. The page title said "Google Help" in English.&lt;/p&gt;

&lt;p&gt;Sixty accessibility attributes on the Japanese page were never translated.&lt;/p&gt;

&lt;p&gt;A blind person using a Japanese screen reader on that page would hear English words randomly mixed into Japanese audio. The screen reader sees lang="en", switches to an English voice, and tries to pronounce Japanese characters with English rules. Complete mess.&lt;/p&gt;

&lt;p&gt;And the kicker: Lighthouse says the page is fine.&lt;/p&gt;

&lt;p&gt;Why Lighthouse misses it&lt;/p&gt;

&lt;p&gt;Lighthouse checks whether an aria-label exists. It does not check what language it is in. So aria-label="Search" on a Japanese page gets a green checkmark because the attribute is present.&lt;/p&gt;

&lt;p&gt;Same with alt text. Same with placeholders. Same with the lang attribute — Lighthouse checks that lang exists on the html tag, but not whether it matches the actual content language.&lt;/p&gt;

&lt;p&gt;To find these issues, you have to do something no tool currently does: open the same page in two languages and compare the accessibility attributes side by side.&lt;/p&gt;

&lt;p&gt;That is what I spent the last few days building.&lt;/p&gt;

&lt;p&gt;What I built&lt;/p&gt;

&lt;p&gt;Locali11y fetches the English and Japanese (or any other language) versions of the same page. It parses both with Cheerio — no browser needed, just raw HTML parsing. Then it runs twenty checks that are specifically designed around what breaks during translation.&lt;/p&gt;

&lt;p&gt;Not generic WCAG checks. Targeted questions like:&lt;/p&gt;

&lt;p&gt;Is the aria-label on this button identical to the English version? If yes, it was probably never translated.&lt;/p&gt;

&lt;p&gt;Does the html lang attribute match the page locale? If the page is in Japanese but lang says "en", something went wrong in the deployment.&lt;/p&gt;

&lt;p&gt;Is the placeholder text the same as the English page? Is the page title? The meta description?&lt;/p&gt;

&lt;p&gt;These are simple string comparisons. The code that does the core detection is about twenty lines. It collects all ARIA labels from the English page, then checks each label on the translated page to see if it matches. If a Japanese page has aria-label="Search" and the English page also has aria-label="Search", that label was never localized.&lt;/p&gt;

&lt;p&gt;No machine learning. No complex parsing. Just a comparison nobody built into a tool before.&lt;/p&gt;

&lt;p&gt;What I found on real sites&lt;/p&gt;

&lt;p&gt;Google Support: 60 locale-specific issues on the Japanese page. Score dropped from 84 to 71. The html lang bug is the worst one — it affects how the entire page is announced by screen readers.&lt;/p&gt;

&lt;p&gt;IKEA Japan: carousel buttons saying "See previous items" and "See next items" in English on a Japanese page. Those are navigation controls that blind users rely on.&lt;/p&gt;

&lt;p&gt;The pattern was the same everywhere I tested. Teams translate the visible content — headings, paragraphs, button text. They forget the invisible stuff — ARIA labels, alt attributes, placeholder values, metadata. And no tool in the standard workflow catches the gap because every tool checks one page at a time.&lt;/p&gt;

&lt;p&gt;The technical choices&lt;/p&gt;

&lt;p&gt;I used Cheerio instead of Playwright because the attributes I care about exist in the static HTML. Running a full browser for each page would add thirty seconds per locale and create deployment headaches. Cheerio parses HTML in milliseconds.&lt;/p&gt;

&lt;p&gt;The app is built with Next.js and uses Supabase for storing audit results. The scoring groups issues by check type so that one category (like missing alt on fifty product images) cannot tank the entire score alone.&lt;/p&gt;

&lt;p&gt;I used Lingo.dev for the app's own translations. The tool itself works in English, Spanish, Japanese, and Chinese. Building a localization auditor with poor localization would have been embarrassing.&lt;/p&gt;

&lt;p&gt;What surprised me&lt;/p&gt;

&lt;p&gt;Browsers do not fix this problem. Chrome's translate feature changes visible text on a page. It does not touch HTML attributes. So even if you translate a page with Chrome, the aria-labels stay English. There was a Chrome bug filed about this in 2019. In 2023 it regressed. As of mid-2025, no major browser reliably translates hidden accessibility attributes.&lt;/p&gt;

&lt;p&gt;The EU Accessibility Act requires accessible digital products across all supported languages. Having lang="en" on a Japanese page is a compliance issue. Most companies have no idea this is happening on their sites because no tool checks for it.&lt;/p&gt;

&lt;p&gt;What is missing&lt;/p&gt;

&lt;p&gt;Locali11y currently checks one page per locale. A full site crawler would be more useful but was too complex for the hackathon timeline. The scoring system is simple and approximate. The fix suggestions use AI and would need human review for production use.&lt;/p&gt;

&lt;p&gt;If someone forks this, the most useful next step would be CI integration — running these checks automatically on every deploy and failing the build when locale accessibility drops below a threshold.&lt;/p&gt;

&lt;p&gt;Links&lt;/p&gt;

&lt;p&gt;Live: &lt;a href="https://locali11y.vercel.app/en" rel="noopener noreferrer"&gt;https://locali11y.vercel.app/en&lt;/a&gt;&lt;br&gt;
Code: &lt;a href="https://github.com/Vaibhav13Shukla/locali11y" rel="noopener noreferrer"&gt;https://github.com/Vaibhav13Shukla/locali11y&lt;/a&gt;&lt;br&gt;
Video: &lt;a href="https://youtu.be/dWck2xBytMs" rel="noopener noreferrer"&gt;https://youtu.be/dWck2xBytMs&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Built for the Lingo.dev Multilingual Hackathon #3. App localization powered by Lingo.dev.&lt;/p&gt;

</description>
      <category>webdev</category>
      <category>lingdotdev</category>
      <category>hackathon</category>
      <category>javascript</category>
    </item>
    <item>
      <title>The AI That Never Forgets — My Vision Agents Hackathon Journey</title>
      <dc:creator>Vaibhav Shukla</dc:creator>
      <pubDate>Sun, 01 Mar 2026 18:44:00 +0000</pubDate>
      <link>https://dev.to/vaibhav_shukla_20/the-ai-that-never-forgets-my-vision-agents-hackathon-journey-34g</link>
      <guid>https://dev.to/vaibhav_shukla_20/the-ai-that-never-forgets-my-vision-agents-hackathon-journey-34g</guid>
      <description>&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fyrrlxnmlj8nb2mw5zjs0.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fyrrlxnmlj8nb2mw5zjs0.jpeg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fui86gn5k03nouade73hs.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fui86gn5k03nouade73hs.jpeg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0qn9a5rfdma6816fqsb.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.amazonaws.com%2Fuploads%2Farticles%2Fn0qn9a5rfdma6816fqsb.jpeg" alt=" "&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  How I solved video AI's biggest blind spot — amnesia — by building a real-time temporal memory engine on top of the Vision Agents SDK by Stream.
&lt;/h2&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://img.shields.io/badge/ARGUS-The_AI_That_Never_Forgets-00d4aa?style=for-the-badge&amp;amp;logo=openai&amp;amp;logoColor=white&amp;amp;labelColor=101010&amp;amp;logoWidth=30&amp;amp;pd=20" rel="noopener noreferrer"&gt;https://img.shields.io/badge/ARGUS-The_AI_That_Never_Forgets-00d4aa?style=for-the-badge&amp;amp;logo=openai&amp;amp;logoColor=white&amp;amp;labelColor=101010&amp;amp;logoWidth=30&amp;amp;pd=20&lt;/a&gt;
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;🏆 &lt;strong&gt;Built for the &lt;a href="https://www.wemakedevs.org/hackathons/vision" rel="noopener noreferrer"&gt;Vision Possible: Agent Protocol Hackathon&lt;/a&gt;&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
⚡ &lt;strong&gt;Powered by &lt;a href="https://github.com/GetStream/Vision-Agents" rel="noopener noreferrer"&gt;Vision Agents SDK by Stream&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  🎥 Watch ARGUS in Action
&lt;/h2&gt;

&lt;p&gt;Before I explain how I built it, you have to see it to believe it. Here is ARGUS detecting objects, tracking them over time, and answering questions about the past in real-time.&lt;/p&gt;

&lt;p&gt;

&lt;/p&gt;
&lt;div class="crayons-card c-embed text-styles text-styles--secondary"&gt;
    &lt;div class="c-embed__content"&gt;
      &lt;div class="c-embed__body flex items-center justify-between"&gt;
        &lt;a href="PASTE_YOUR_VIDEO_LINK_HERE" rel="noopener noreferrer" class="c-link fw-bold flex items-center"&gt;
          &lt;span class="mr-2"&gt;PASTE_YOUR_VIDEO_LINK_HERE&lt;/span&gt;
          

        &lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
&lt;/div&gt;




&lt;p&gt;&lt;em&gt;(If the video doesn't load, &lt;a href="https://youtu.be/DXfTK2XOzMk" rel="noopener noreferrer"&gt;click here to watch the demo&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  🚨 The Problem: AI Has Amnesia
&lt;/h2&gt;

&lt;p&gt;I realized something frustrating while testing modern Video AI demos. They are brilliant at telling you what is happening &lt;em&gt;right now&lt;/em&gt;, but they are terrible at telling you what happened &lt;em&gt;5 minutes ago&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;If I drop my keys and ask a standard AI agent, "Where are my keys?", it looks at the &lt;em&gt;current&lt;/em&gt; frame, sees nothing, and says: "I don't see any keys."&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Vision Agents SDK documentation actually highlighted this limitation:&lt;/strong&gt;&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"Longer videos can cause the AI to lose context. For instance, if it's watching a soccer match, it will get confused after 30 seconds."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That was my lightbulb moment. 💡&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The Mission:&lt;/strong&gt; Build &lt;strong&gt;ARGUS&lt;/strong&gt;, a real-time agent that doesn't just "see" video—it &lt;strong&gt;remembers&lt;/strong&gt; it.&lt;/p&gt;




&lt;h2&gt;
  
  
  🧠 What is ARGUS?
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ARGUS&lt;/strong&gt; is a multimodal AI agent that watches live video, tracks objects using computer vision, and maintains a &lt;strong&gt;Temporal Memory Engine&lt;/strong&gt;. &lt;/p&gt;

&lt;p&gt;Unlike standard agents that process &lt;code&gt;Frame → Detect → Forget&lt;/code&gt;, ARGUS uses a stateful pipeline:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;&lt;br&gt;
&lt;code&gt;mermaid&lt;br&gt;
graph LR&lt;br&gt;
A[Camera Feed] --&amp;gt; B(YOLO26 Detection)&lt;br&gt;
B --&amp;gt; C{Temporal Memory Engine}&lt;br&gt;
C --&amp;gt; D[Update Object History]&lt;br&gt;
C --&amp;gt; E[Log Events]&lt;br&gt;
D &amp;amp; E --&amp;gt; F[LLM Context]&lt;br&gt;
F --&amp;gt; G((Voice Response))&lt;/code&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Key Capabilities
👁️ Real-time Tracking: Uses YOLO26 Nano + ByteTrack to assign persistent IDs to objects.
🕰️ Time Travel: Can answer "What did I hold up 2 minutes ago?"
📍 Spatial Awareness: Converts raw coordinates into human terms like "top-left" or "center."
🗣️ Voice Interaction: Full duplex voice conversation with &amp;lt;1s latency.
💬 Real Conversations With ARGUS
These are actual interactions from my testing sessions:

Terminal Logs showing events
Real-time event logging showing objects appearing and moving.

I Said  ARGUS Responded
"What do you see?"  "Person ID:2 at middle-center, Cup ID:3 at bottom-right"
"What am I holding?"    "You appear to be holding a bottle, ID:7"
"What just moved?"  "Cup moved from bottom-left to bottom-right at 2:05 PM"
"Summarize everything"  "Person appeared at center 30s ago. Cup moved left to right at 2:05"
⚡ Response time: ~1 second
🧠 All answers came from temporal memory — not from re-analyzing the video frame.

🏗️ The Architecture &amp;amp; Tech Stack
I needed a stack that was fast, cheap, and capable of handling real-time video streams without melting my laptop.

Component   Technology  Why I Chose It
Framework   Vision Agents SDK   It handled all the WebRTC/Audio/Video piping for me.
Vision Model    YOLO26 Nano Benchmarked at 130ms/frame on CPU. Fast &amp;amp; Accurate.
Reasoning   Llama 3.3 via OpenRouter    Fast inference with tool-calling capabilities.
Speech  Deepgram (STT) + ElevenLabs (TTS)   The lowest latency combo available.
Transport   Stream Edge Network Kept video latency under 30ms.
🛠️ The Build Journey
1. The "Secret Weapon": Temporal Memory Engine
This is the heart of the project. I wrote a custom Python class that sits between the vision processor and the LLM.

Instead of feeding raw video frames to the LLM (which is slow and expensive), I feed it structured event logs.

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Python&lt;/p&gt;

&lt;h1&gt;
  
  
  Core logic: If an object moves zones, log it.
&lt;/h1&gt;

&lt;p&gt;if old_zone != zone:&lt;br&gt;
    self._log("moved", f"{class_name} (ID:{track_id}) moved from {old_zone} to {zone}")&lt;br&gt;
When I ask, "Where is the cup?", the LLM receives this context injection:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;
[ARGUS MEMORY]

Cup (ID:2): Last seen at bottom-right at 12:05 PM.
Person (ID:1): Currently visible at center.
Event: Cup moved from left to right 30 seconds ago.
2. Building the Custom Processor
Using the SDK's VideoProcessorPublisher pattern was intuitive. I could access the raw av.VideoFrame, run my YOLO inference, draw bounding boxes, and push the frame back to the browser.

ARGUS Detection View
ARGUS tracking objects with persistent IDs and Spatial Zones.

3. Solving the Latency Problem
My first prototype had 5-second delays. To fix this, I optimized ruthlessly:

Switched from Gemini (Rate limits) to OpenRouter/Llama.
Switched YOLO11 to YOLO26 Nano (7.7 FPS on CPU).
Used human-readable zones ("top-left") instead of raw coordinates, reducing token usage for the LLM.
🧪 Benchmark Results
I ran a diagnostic script to prove efficiency on a standard laptop (No GPU):

&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Benchmark Results&lt;/p&gt;

&lt;p&gt;Model   Speed   Max FPS Verdict&lt;br&gt;
YOLO26 Nano 130ms   7.7 ✅ Winner&lt;br&gt;
YOLOv8 Nano 138ms   7.2 Solid&lt;br&gt;
YOLO11 Small    310ms   3.2 Too slow&lt;/p&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;

The Vision Agents SDK was crucial here. Because it handles the video transport efficiently, I could use all my CPU cycles for the actual detection logic.

💡 The "Aha!" Moment
The magic happened during a test run. I held up a water bottle, put it down, and waited. Then I asked:

Me: "What did I just show you?"

ARGUS: "You were holding a bottle (ID:7) at the center of the screen about 15 seconds ago."

The Aha Moment

It wasn't looking at the bottle now. It remembered. That feeling of interacting with an AI that has object permanence is wild.

🌍 Why This Matters
Hackathons often produce cool demos that don't solve real problems. ARGUS solves the context window problem for video.

By abstracting video into structured temporal data, we can build agents that:

Monitor security feeds for hours and summarize activity.
Help find lost items in a room.
Analyze workflow efficiency in factories.
The Vision Agents SDK made this possible by removing the complexity of WebRTC and audio handling, allowing me to focus entirely on the memory innovation.

🔗 Links &amp;amp; Resources
Code Repository: GitHub - [ARGUS](https://github.com/Vaibhav13Shukla/argus)
Vision Agents SDK: Star the Repo!
Hackathon: Vision Possible
Thanks to Stream and WeMakeDevs for this challenge. It pushed me to build something I didn't think was possible in a weekend!
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>hackathon</category>
      <category>computervision</category>
    </item>
  </channel>
</rss>
