<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: chinmay garg</title>
    <description>The latest articles on DEV Community by chinmay garg (@chinmay_garg).</description>
    <link>https://dev.to/chinmay_garg</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F217922%2Fea7c2e55-e473-4f83-9a95-e162c758259c.jpg</url>
      <title>DEV Community: chinmay garg</title>
      <link>https://dev.to/chinmay_garg</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/chinmay_garg"/>
    <language>en</language>
    <item>
      <title>On-device agents need their own identity: a delegation design for the edge</title>
      <dc:creator>chinmay garg</dc:creator>
      <pubDate>Sat, 03 Oct 2026 04:30:00 +0000</pubDate>
      <link>https://dev.to/chinmay_garg/on-device-agents-need-their-own-identity-a-delegation-design-for-the-edge-4cj8</link>
      <guid>https://dev.to/chinmay_garg/on-device-agents-need-their-own-identity-a-delegation-design-for-the-edge-4cj8</guid>
      <description>&lt;p&gt;Local inference solves where the data lives. It says nothing about who the agent is when it calls out.&lt;/p&gt;

&lt;p&gt;Phone silicon can now run large mixture-of-experts models from flash and build a personal knowledge graph without a network call. Qualcomm's Snapdragon Summit pitch (Sep 22-24) is a personal agent that keeps your data on the device. Its own list of what such an agent needs includes "secure permissions," and I found no published spec for how an on-device agent proves its identity or its delegated authority to the service it calls. The one place delegation limits were named was the Qualcomm-Mastercard agentic commerce announcement on Sep 22.&lt;/p&gt;

&lt;p&gt;I haven't built on this silicon. This post is about the design problem, which applies whatever the chip.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv03tedzc6nvbn23jjnn9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fv03tedzc6nvbn23jjnn9.png" alt="AI Agents working within governance" width="799" height="436"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The failure mode
&lt;/h2&gt;

&lt;p&gt;Local processing is not local authority. Sooner or later the agent calls an API, and that API has to answer three questions: which agent is this, who delegated this action, and when does the delegation expire. A long-lived token in app storage answers none of them. It also sits next to a model that reads untrusted text all day.&lt;/p&gt;

&lt;p&gt;The cloud version of this failure is public. An OpenAI agent routed around blocks on Services Australia's Medicare portal. The access happened Jun 18, the agency was notified Sep 10, and it was disclosed Sep 24. That is 84 days from access to disclosure, in part because the activity was hard to attribute.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;strong&gt;A design sketch&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Per-instance agent key.&lt;/strong&gt; Generate a non-exportable keypair in the hardware-backed keystore when the agent is installed. The agent's identity is that key, not the user's session.&lt;br&gt;
&lt;strong&gt;Delegation as a signed, short-lived token.&lt;/strong&gt; The user authorizes a task through an OS-level user-presence prompt. The result is a token naming the agent key, the task, the allowed tools and resources, and an expiry in minutes.&lt;br&gt;
&lt;strong&gt;Proof of possession on every call.&lt;/strong&gt; The agent signs each outbound request with its key. A lifted token is useless without the key that never leaves the hardware.&lt;br&gt;
&lt;strong&gt;Approval out of band.&lt;/strong&gt; "The user approved" is a signed assertion from the OS prompt, checked by the receiving service. A field the agent wrote into the request is untrusted data.&lt;br&gt;
&lt;strong&gt;Separate logs.&lt;/strong&gt; The device logs agent actions apart from user actions, and the receiving service logs the agent key and delegation ID next to the user account.&lt;/p&gt;

&lt;p&gt;The receiving side verifies signatures. It never has to trust the agent's description of itself.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What it costs&lt;/strong&gt;&lt;br&gt;
Extra latency per call: one signature, plus verification on the server.&lt;br&gt;
Operational: token issuance and expiry handling for every task.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where I am unsure&lt;/strong&gt;&lt;br&gt;
What happens when the device is offline and the delegation expires mid-task. Fail closed is safe and annoying.&lt;br&gt;
Whether attestation should tell the receiving service which model the agent runs. Useful for risk decisions, and easy to spoof without a trusted root.&lt;/p&gt;

&lt;p&gt;How this maps to India's DPDP Act, which expects you to show who handled a personal-data record and on what basis. An agent on a phone that forwards a record to a cloud service breaks that trail unless the delegation reference travels with it.&lt;/p&gt;

&lt;p&gt;If you have shipped any of this on-device, I would like to hear what broke.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>agents</category>
    </item>
    <item>
      <title>AI on Edge: One forward pass, one letter: a decision model that runs on a laptop</title>
      <dc:creator>chinmay garg</dc:creator>
      <pubDate>Fri, 02 Oct 2026 04:58:31 +0000</pubDate>
      <link>https://dev.to/chinmay_garg/ai-on-edge-one-forward-pass-one-letter-a-decision-model-that-runs-on-a-laptop-1c82</link>
      <guid>https://dev.to/chinmay_garg/ai-on-edge-one-forward-pass-one-letter-a-decision-model-that-runs-on-a-laptop-1c82</guid>
      <description>&lt;p&gt;&lt;strong&gt;How I get a confidence score for every option without generating text, and why that score should only ever make a gate stricter.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A voice bot does not need a paragraph about what the caller wants. It needs a label, and a number that says how sure the model is. So I trained a small model to give exactly that, and I wanted to see how far one 24 GB MacBook could take it.&lt;/p&gt;

&lt;p&gt;TypeSafe AI has described this style publicly as "decision models". This is my own version, built in the open-source forge fine-tuning tool. It is not their model, and I make no comparison to it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The setup&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The model is Gemma 4 E2B (4-bit) with a LoRA adapter, rank 16. The adapter is 52 MB. Training peaked at about 5.7 GB of memory and took about an hour. The data was seven public datasets, plus synthetic rows for four behaviour tasks made by a local model on the same Mac. No paid API and no cloud GPU.&lt;/p&gt;

&lt;p&gt;Every example has the same shape: the conversation, a question, and lettered options.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="err"&gt;Customer:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;I&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;need&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;to&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;move&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;my&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;appointment,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;it's&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;urgent.&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="err"&gt;Question:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;Which&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;request&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;is&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;customer&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;making?&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;Options:&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;A)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;reschedule&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;an&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;appointment&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;B)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;cancel&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;an&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;appointment&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="err"&gt;C)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;ask&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;about&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="err"&gt;pricing&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model is trained to answer with one letter. At test time I do not let it write anything. I run the prompt once, read the raw scores for the letters A, B and C, and turn those into probabilities. A temperature fitted on a validation set makes them well calibrated. There is nothing to parse, and the answer can never be an invalid label.&lt;/p&gt;

&lt;p&gt;Three details mattered:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Options are shuffled and the question is reworded on every training row, because models favour certain letters. Evaluation keeps a fixed order.&lt;/li&gt;
&lt;li&gt;Multi-select questions become many yes/no questions, so one format covers everything.&lt;/li&gt;
&lt;li&gt;I read the scores in-process. The server I use only returns the top 11 log-probabilities, so a 13-option list would silently lose probability mass.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;What good calibration gives you&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On about 68,000 test rows, the calibration error was between 0.006 and 0.04 on most tasks. Median latency was about 85 ms per decision (p95 117 ms), measured on the laptop GPU. That allows a simple rule: act when the model is at least 80% sure, and otherwise hand off.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Test&lt;/th&gt;
&lt;th&gt;Handled at 0.8 or higher&lt;/th&gt;
&lt;th&gt;Accuracy on those&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;CLINC150 (held out)&lt;/td&gt;
&lt;td&gt;90%&lt;/td&gt;
&lt;td&gt;98.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MASSIVE intent&lt;/td&gt;
&lt;td&gt;79%&lt;/td&gt;
&lt;td&gt;95.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;BANKING77 (held out)&lt;/td&gt;
&lt;td&gt;74%&lt;/td&gt;
&lt;td&gt;92.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;The part that is not about the model&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A confident 0.92 says the label is probably right. It does not say the agent is allowed to act. The message being scored is text the caller wrote, so any model that reads it can be pushed around by it.&lt;/p&gt;

&lt;p&gt;This is how I would combine the two in an agent gateway:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;check identity and scoped permission first
if denied: stop

score = decision_model(message)
if score is below the threshold: send to a human
else: go ahead
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The score can turn a "go ahead" into a "send to a human". It can never turn a "denied" into a "go ahead". Identity and scoped permission decide. The model only narrows. And because the model runs locally, the customer's message never goes to a third-party API, which helps with DPDPA.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What I am not claiming&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The "held-out" intent tests pick from about ten options with random wrong answers. That is easier than the published 77-way and 151-way benchmarks.&lt;br&gt;
The behaviour-task tests are small, and the test rows were made by the same kind of model that made the training rows.&lt;br&gt;
I have not tested on real recorded calls. Everything above is on public data.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Where I am stuck&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Training loss stays flat near ln(26), which is a uniform guess over the answer letters, for roughly the first 2,000 steps. Then it drops sharply. A lower learning rate moved the plateau but did not remove it. I do not know why.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open question:&lt;/strong&gt; &lt;/p&gt;

&lt;p&gt;where should the hand-off threshold live? Per task, per data type, or per permission? I lean towards per data type. If you run something like this in production, I would like to know what you chose.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mcp</category>
      <category>slm</category>
    </item>
  </channel>
</rss>
