<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Gaurav Talesara</title>
    <description>The latest articles on DEV Community by Gaurav Talesara (@gaurav_talesara).</description>
    <link>https://dev.to/gaurav_talesara</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3724256%2F3ab1fc88-aa7b-426c-9783-c019bd8e5915.png</url>
      <title>DEV Community: Gaurav Talesara</title>
      <link>https://dev.to/gaurav_talesara</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/gaurav_talesara"/>
    <language>en</language>
    <item>
      <title>3 Redis Design Failures You Should Avoid Before They Become Production Incidents</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Wed, 30 Sep 2026 17:01:48 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/3-redis-design-failures-you-should-avoid-before-they-become-production-incidents-2pe2</link>
      <guid>https://dev.to/gaurav_talesara/3-redis-design-failures-you-should-avoid-before-they-become-production-incidents-2pe2</guid>
      <description>&lt;p&gt;Caching is usually introduced for one simple reason:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Take pressure off the database and make reads faster.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The architecture often looks straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
  v
Application
  |
  +----&amp;gt; Redis
  |
  +----&amp;gt; Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A request checks Redis first.&lt;/p&gt;

&lt;p&gt;If the value exists, return it.&lt;/p&gt;

&lt;p&gt;If it doesn't, query the database, put the result into Redis, and return it.&lt;/p&gt;

&lt;p&gt;Simple.&lt;/p&gt;

&lt;p&gt;Until the system meets real production traffic.&lt;/p&gt;

&lt;p&gt;That's where caching stops being just a performance optimization and becomes a &lt;strong&gt;system-design problem&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Let's walk through how a seemingly simple cache evolves through three stages—and the failure modes that appear at each stage.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Stale Data
&lt;/h2&gt;

&lt;p&gt;The first version of a cache is usually &lt;strong&gt;cache-aside&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The application owns the caching logic.&lt;/p&gt;

&lt;p&gt;The flow is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   |
   v
Check Redis
   |
   +---- HIT ----&amp;gt; Return cached value
   |
   +---- MISS ---&amp;gt; Query database
                     |
                     v
                 Store in Redis
                     |
                     v
                  Return
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For read-heavy workloads, this can dramatically reduce database traffic.&lt;/p&gt;

&lt;p&gt;Suppose we have a user profile:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;42&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Gaurav"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"email"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gaurav@example.com"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application stores it under:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user:42
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything works.&lt;/p&gt;

&lt;p&gt;Until the user changes their name.&lt;/p&gt;

&lt;p&gt;The database is updated:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gaurav
    ↓
Gaurav Talesara
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But Redis still contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gaurav
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have two versions of reality.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database → Gaurav Talesara
Redis    → Gaurav
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database is correct.&lt;/p&gt;

&lt;p&gt;The application can still return the wrong answer.&lt;/p&gt;

&lt;p&gt;This is &lt;strong&gt;stale data&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why this matters
&lt;/h3&gt;

&lt;p&gt;Caching introduces another copy of your data.&lt;/p&gt;

&lt;p&gt;The moment you do that, you have to answer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;When does the cached copy stop being valid?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's the beginning of cache consistency design.&lt;/p&gt;




&lt;h1&gt;
  
  
  2. TTL Helps—but Doesn't Solve Everything
&lt;/h1&gt;

&lt;p&gt;A common next step is adding a &lt;strong&gt;TTL (Time To Live)&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;user:42
TTL = 5 minutes
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After five minutes, Redis considers the entry expired.&lt;/p&gt;

&lt;p&gt;The next request becomes a cache miss:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Redis MISS
    |
    v
Database
    |
    v
Fresh value
    |
    v
Redis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you eventual freshness without requiring every application write to explicitly update the cache.&lt;/p&gt;

&lt;p&gt;But TTL introduces a trade-off.&lt;/p&gt;

&lt;h3&gt;
  
  
  Longer TTL
&lt;/h3&gt;

&lt;p&gt;You get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fewer database reads&lt;/li&gt;
&lt;li&gt;better cache hit rate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But potentially:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;older data&lt;/li&gt;
&lt;li&gt;longer periods of inconsistency&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Shorter TTL
&lt;/h3&gt;

&lt;p&gt;You get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;fresher data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But potentially:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;more cache misses&lt;/li&gt;
&lt;li&gt;more database traffic&lt;/li&gt;
&lt;li&gt;lower cache efficiency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So TTL isn't really a "freshness setting."&lt;/p&gt;

&lt;p&gt;It's a &lt;strong&gt;consistency vs. performance trade-off&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And some data simply can't wait five minutes.&lt;/p&gt;

&lt;p&gt;If a user changes their email address, waiting for TTL expiration may be unacceptable.&lt;/p&gt;

&lt;p&gt;That's where &lt;strong&gt;explicit invalidation&lt;/strong&gt; comes in.&lt;/p&gt;




&lt;h1&gt;
  
  
  3. Invalidation on Write
&lt;/h1&gt;

&lt;p&gt;Instead of waiting for the cache to expire, the application can invalidate the relevant key when the database changes.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;UPDATE users
SET name = 'Gaurav Talesara'
WHERE id = 42;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;DELETE user:42 FROM Redis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next read becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Redis MISS
    |
    v
Database
    |
    v
Fresh data
    |
    v
Redis
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives you a much fresher cache.&lt;/p&gt;

&lt;p&gt;But now the application has to maintain the relationship between:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;database writes ↔ cache invalidation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And that creates its own failure scenarios.&lt;/p&gt;

&lt;p&gt;What if the database update succeeds but cache invalidation fails?&lt;/p&gt;

&lt;p&gt;You can end up with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database → NEW value
Redis    → OLD value
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What if the invalidation happens first but the database transaction fails?&lt;/p&gt;

&lt;p&gt;Now the cache is empty even though the database still contains the old value.&lt;/p&gt;

&lt;p&gt;There is no universal answer here.&lt;/p&gt;

&lt;p&gt;The right approach depends on the consistency requirements of the application.&lt;/p&gt;

&lt;p&gt;But there is another problem that has nothing to do with stale data.&lt;/p&gt;

&lt;p&gt;And this one can take down the database.&lt;/p&gt;




&lt;h1&gt;
  
  
  4. Cache Stampede
&lt;/h1&gt;

&lt;p&gt;Imagine you have a popular product page:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;product:123
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Normally it receives thousands of requests.&lt;/p&gt;

&lt;p&gt;Because the data is cached, most of those requests never reach the database.&lt;/p&gt;

&lt;p&gt;That's exactly what we wanted.&lt;/p&gt;

&lt;p&gt;Now the cache entry expires.&lt;/p&gt;

&lt;p&gt;At almost the same moment, thousands of requests arrive.&lt;/p&gt;

&lt;p&gt;They all see:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CACHE MISS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And they all do this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request 1 → DB
Request 2 → DB
Request 3 → DB
Request 4 → DB
...
Request 5000 → DB
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database was protected by Redis a moment ago.&lt;/p&gt;

&lt;p&gt;Now thousands of requests are hitting it simultaneously.&lt;/p&gt;

&lt;p&gt;This is a &lt;strong&gt;cache stampede&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Another name you'll often see for this pattern is &lt;strong&gt;thundering herd&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The important point is that the problem isn't simply that the cache missed.&lt;/p&gt;

&lt;p&gt;The problem is that &lt;strong&gt;many requests independently react to the same cache miss&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  5. Defense #1: Single Flight
&lt;/h1&gt;

&lt;p&gt;One way to solve this is to allow only one request to refresh a missing key.&lt;/p&gt;

&lt;p&gt;Suppose 5,000 requests miss:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;             Redis MISS
                  |
        +---------+---------+
        |                   |
    Request 1            Requests 2-5000
        |                   |
      DB query              WAIT
        |                   |
        +--------+----------+
                 |
          Fresh result
                 |
             Redis SET
                 |
        Everyone gets result
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Request #1 becomes responsible for fetching the data.&lt;/p&gt;

&lt;p&gt;The other requests wait for the result instead of independently hitting the database.&lt;/p&gt;

&lt;p&gt;This pattern is often called &lt;strong&gt;single-flight&lt;/strong&gt; or &lt;strong&gt;request coalescing&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The key idea is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;One cache miss should not become 5,000 database queries.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;This is particularly useful when a small number of keys receive a very large amount of traffic.&lt;/p&gt;




&lt;h1&gt;
  
  
  6. Defense #2: TTL Jitter
&lt;/h1&gt;

&lt;p&gt;There is another failure pattern that can happen even when you aren't dealing with one single hot key.&lt;/p&gt;

&lt;p&gt;Imagine millions of cache entries are created at roughly the same time.&lt;/p&gt;

&lt;p&gt;If every entry gets exactly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TTL = 300 seconds
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then many of them can expire around the same time.&lt;/p&gt;

&lt;p&gt;That can create a synchronized wave of cache misses.&lt;/p&gt;

&lt;p&gt;Instead, you can introduce &lt;strong&gt;jitter&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Rather than:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TTL = 300 seconds
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you might use something conceptually like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TTL = 300 + random(0, 60)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now expiration is spread across time.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                    ↓
              EVERYTHING EXPIRES
                    ↓
             DATABASE SPIKE
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you get something more like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;      ↓      ↓          ↓    ↓       ↓
   expire  expire     expire expire  expire
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database sees a more distributed load pattern.&lt;/p&gt;

&lt;p&gt;TTL jitter doesn't eliminate cache misses.&lt;/p&gt;

&lt;p&gt;It reduces the probability of &lt;strong&gt;synchronized expiration&lt;/strong&gt; becoming a traffic spike.&lt;/p&gt;




&lt;h1&gt;
  
  
  7. Defense #3: Hot-Key Replication
&lt;/h1&gt;

&lt;p&gt;Now consider a different situation.&lt;/p&gt;

&lt;p&gt;One key becomes extremely popular.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;product:123
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;might represent a product that suddenly goes viral.&lt;/p&gt;

&lt;p&gt;You can have thousands or millions of requests for essentially the same piece of data.&lt;/p&gt;

&lt;p&gt;Even if Redis is distributed, concentrating a huge amount of traffic around one key can create a bottleneck.&lt;/p&gt;

&lt;p&gt;This is a &lt;strong&gt;hot key&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;One approach is to replicate frequently accessed data across multiple cache nodes so requests don't all depend on the same location.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                  product:123
                       |
            +----------+----------+
            |          |          |
            v          v          v
         Redis-1    Redis-2    Redis-3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now requests can be distributed rather than concentrating entirely on one cache location.&lt;/p&gt;

&lt;p&gt;But this introduces another design question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How do you keep replicated hot data consistent?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Again, the optimization introduces another system-design trade-off.&lt;/p&gt;




&lt;h1&gt;
  
  
  8. The Pattern Behind All Three Failures
&lt;/h1&gt;

&lt;p&gt;This is the part that's easy to miss.&lt;/p&gt;

&lt;p&gt;You start with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You add Redis:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application
     |
   Redis
     |
 Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have better read performance.&lt;/p&gt;

&lt;p&gt;Then production reveals stale data.&lt;/p&gt;

&lt;p&gt;So you add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;TTL + Invalidation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then production traffic reveals cache stampedes.&lt;/p&gt;

&lt;p&gt;So you add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Single Flight
TTL Jitter
Request Coalescing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then traffic concentration reveals hot keys.&lt;/p&gt;

&lt;p&gt;So you consider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Replication
Load Distribution
Failure Isolation
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The architecture keeps evolving.&lt;/p&gt;

&lt;p&gt;That's because caching isn't simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Put Redis in front of the database."&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You're introducing another distributed component with its own:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;consistency behavior&lt;/li&gt;
&lt;li&gt;expiration behavior&lt;/li&gt;
&lt;li&gt;concurrency&lt;/li&gt;
&lt;li&gt;failure modes&lt;/li&gt;
&lt;li&gt;traffic patterns&lt;/li&gt;
&lt;li&gt;recovery requirements&lt;/li&gt;
&lt;/ul&gt;




&lt;h1&gt;
  
  
  9. What to Design Before Adding Redis
&lt;/h1&gt;

&lt;p&gt;Before introducing a cache, I'd want the team to answer at least these questions.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. What can be stale?
&lt;/h3&gt;

&lt;p&gt;Not every piece of data has the same consistency requirement.&lt;/p&gt;

&lt;p&gt;Ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Can this value be 1 second old?
10 seconds?
5 minutes?
Never?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That answer should influence the caching strategy.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. What invalidates the data?
&lt;/h3&gt;

&lt;p&gt;Define the relationship between writes and cache entries.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User updated
    ↓
Database committed
    ↓
Invalidate user:42
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Don't leave invalidation as an afterthought.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. What happens when the key expires?
&lt;/h3&gt;

&lt;p&gt;Ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 request misses?
100?
10,000?
1,000,000?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the answer is "they all query the database," you may have a stampede waiting to happen.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. What happens when one key becomes extremely popular?
&lt;/h3&gt;

&lt;p&gt;Measure key access patterns.&lt;/p&gt;

&lt;p&gt;A distributed cache doesn't automatically mean distributed traffic.&lt;/p&gt;

&lt;p&gt;You can still have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;99% of requests → 1 key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  5. What happens when Redis is unavailable?
&lt;/h3&gt;

&lt;p&gt;This is one of the most important questions.&lt;/p&gt;

&lt;p&gt;If Redis disappears, does the application:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fail closed?
      or
Fall back to database?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If everything falls back to the database simultaneously, Redis failure can become a database incident.&lt;/p&gt;

&lt;p&gt;Your fallback path needs capacity planning too.&lt;/p&gt;




&lt;h1&gt;
  
  
  10. The Real Lesson
&lt;/h1&gt;

&lt;p&gt;Caching is one of the easiest performance optimizations to add to an architecture.&lt;/p&gt;

&lt;p&gt;It's also easy to underestimate what you're introducing.&lt;/p&gt;

&lt;p&gt;The progression often looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cache
  ↓
Stale Data
  ↓
TTL / Invalidation
  ↓
Cache Stampede
  ↓
Single Flight / Jitter
  ↓
Hot Keys
  ↓
Replication / Distribution
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each step solves a real problem.&lt;/p&gt;

&lt;p&gt;And each step introduces another design decision.&lt;/p&gt;

&lt;p&gt;That's the real lesson:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A cache doesn't remove complexity. It moves complexity somewhere else in the system.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So before adding Redis, don't only ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"How much faster will this make our application?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Ask:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What happens when the cache is stale?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What happens when the key expires?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What happens when 10,000 requests miss simultaneously?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What happens when one key receives 100× normal traffic?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"What happens when Redis itself fails?"&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Those questions are where caching stops being a feature and becomes &lt;strong&gt;system design&lt;/strong&gt;.&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>database</category>
      <category>performance</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>I Tried to Break a $12 Server. It Took 5,000+ Concurrent Users.</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Sat, 26 Sep 2026 05:15:45 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/i-tried-to-break-a-12-server-it-took-5000-concurrent-users-3hbe</link>
      <guid>https://dev.to/gaurav_talesara/i-tried-to-break-a-12-server-it-took-5000-concurrent-users-3hbe</guid>
      <description>&lt;p&gt;There is a common assumption in backend engineering:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If you expect traffic to grow, you need to start thinking about scaling infrastructure early.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;More servers.&lt;br&gt;
More CPU.&lt;br&gt;
A dedicated database.&lt;br&gt;
Redis.&lt;br&gt;
Load balancers.&lt;br&gt;
Containers.&lt;br&gt;
Kubernetes.&lt;/p&gt;

&lt;p&gt;Sometimes you do.&lt;/p&gt;

&lt;p&gt;But sometimes the application isn't actually asking for more infrastructure yet.&lt;/p&gt;

&lt;p&gt;I wanted to see how far a very small server could go before reaching a real bottleneck.&lt;/p&gt;

&lt;p&gt;So I rented a server for &lt;strong&gt;$12/month&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The machine had:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1 CPU&lt;/li&gt;
&lt;li&gt;2 GB RAM&lt;/li&gt;
&lt;li&gt;Nginx&lt;/li&gt;
&lt;li&gt;Node.js&lt;/li&gt;
&lt;li&gt;PostgreSQL&lt;/li&gt;
&lt;li&gt;No dedicated database server&lt;/li&gt;
&lt;li&gt;No Redis&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then I started increasing the load until something broke.&lt;/p&gt;

&lt;p&gt;The interesting part wasn't how cheap the server was.&lt;/p&gt;

&lt;p&gt;It was &lt;strong&gt;what became the bottleneck first&lt;/strong&gt;.&lt;/p&gt;


&lt;h2&gt;
  
  
  The experiment
&lt;/h2&gt;

&lt;p&gt;The goal was simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Find the point where the application starts showing meaningful performance degradation, identify the bottleneck, and improve it without upgrading the server.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I used a small API application running on a single machine.&lt;/p&gt;

&lt;p&gt;The application had a database-backed feed/workflow-style endpoint that required reading data from PostgreSQL and generating a response for the client.&lt;/p&gt;

&lt;p&gt;The database was running on the same machine as the application.&lt;/p&gt;

&lt;p&gt;There was no separate cache layer.&lt;/p&gt;

&lt;p&gt;No Redis.&lt;/p&gt;

&lt;p&gt;No horizontally scaled application servers.&lt;/p&gt;

&lt;p&gt;Just:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Internet
   |
 Nginx
   |
 Node.js
   |
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The server cost about &lt;strong&gt;$12/month&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That was the entire point.&lt;/p&gt;

&lt;p&gt;I wanted to see what the application could actually handle before adding infrastructure.&lt;/p&gt;




&lt;h2&gt;
  
  
  Making the test more realistic
&lt;/h2&gt;

&lt;p&gt;A simple benchmark that repeatedly calls one endpoint isn't particularly interesting.&lt;/p&gt;

&lt;p&gt;Real users don't behave like that.&lt;/p&gt;

&lt;p&gt;They open something.&lt;/p&gt;

&lt;p&gt;They read.&lt;/p&gt;

&lt;p&gt;They perform an action.&lt;/p&gt;

&lt;p&gt;They come back.&lt;/p&gt;

&lt;p&gt;They refresh.&lt;/p&gt;

&lt;p&gt;They request more data.&lt;/p&gt;

&lt;p&gt;So I used &lt;strong&gt;k6&lt;/strong&gt; to create virtual users that followed a more realistic pattern.&lt;/p&gt;

&lt;p&gt;The virtual users would:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Request the feed/workflow data.&lt;/li&gt;
&lt;li&gt;Perform an occasional write/action.&lt;/li&gt;
&lt;li&gt;Request the data again.&lt;/li&gt;
&lt;li&gt;Repeat the cycle.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I also populated the database with a meaningful amount of data rather than testing against an empty database.&lt;/p&gt;

&lt;p&gt;The load generator itself ran from a second VM so that the machine being tested wasn't also responsible for generating all the traffic.&lt;/p&gt;

&lt;p&gt;The goal was to put pressure on the application server while keeping the test environment simple.&lt;/p&gt;




&lt;h1&gt;
  
  
  Starting with a small load
&lt;/h1&gt;

&lt;p&gt;I didn't jump directly to thousands of users.&lt;/p&gt;

&lt;p&gt;The load increased progressively:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10
50
100
200
1,000
2,000
3,000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This matters because performance problems aren't always obvious when you only look at average latency.&lt;/p&gt;

&lt;p&gt;At low traffic, everything looked healthy.&lt;/p&gt;

&lt;p&gt;As the number of virtual users increased, the system continued to respond normally.&lt;/p&gt;

&lt;p&gt;Then something started changing around the higher loads.&lt;/p&gt;




&lt;h1&gt;
  
  
  2,000 users: the first warning sign
&lt;/h1&gt;

&lt;p&gt;At around &lt;strong&gt;2,000 virtual users&lt;/strong&gt;, the average response time still looked reasonable.&lt;/p&gt;

&lt;p&gt;But the tail latency started moving.&lt;/p&gt;

&lt;p&gt;This is where metrics like &lt;strong&gt;p95 and p99&lt;/strong&gt; become much more useful than simply looking at average latency.&lt;/p&gt;

&lt;p&gt;Imagine these two systems:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;System A
Average: 100 ms
p95:     150 ms
p99:     200 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;versus:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;System B
Average: 100 ms
p95:     700 ms
p99:     2,000 ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both have the same average.&lt;/p&gt;

&lt;p&gt;But the user experience is very different.&lt;/p&gt;

&lt;p&gt;The second system has a tail-latency problem.&lt;/p&gt;

&lt;p&gt;That was the first signal that the server was approaching a limit.&lt;/p&gt;




&lt;h1&gt;
  
  
  3,000 users: the test failed
&lt;/h1&gt;

&lt;p&gt;I increased the load again.&lt;/p&gt;

&lt;p&gt;At approximately &lt;strong&gt;3,000 virtual users&lt;/strong&gt;, the experiment crossed the failure criteria.&lt;/p&gt;

&lt;p&gt;That gave me an important data point:&lt;/p&gt;

&lt;p&gt;The server wasn't simply getting "a little slower."&lt;/p&gt;

&lt;p&gt;There was a point where additional concurrency started pushing the system beyond the acceptable performance envelope.&lt;/p&gt;

&lt;p&gt;So instead of continuing to increase the load blindly, I backed down.&lt;/p&gt;

&lt;p&gt;I wanted to understand the bottleneck first.&lt;/p&gt;




&lt;h1&gt;
  
  
  2,500 concurrent users
&lt;/h1&gt;

&lt;p&gt;At around &lt;strong&gt;2,500 concurrent users&lt;/strong&gt;, the system handled approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~232 requests/sec
~288 ms p95 latency
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those numbers were interesting.&lt;/p&gt;

&lt;p&gt;But the most interesting information came from the resource utilization.&lt;/p&gt;

&lt;p&gt;The server wasn't running out of memory.&lt;/p&gt;

&lt;p&gt;RAM usage was still &lt;strong&gt;under 1 GB&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;PostgreSQL was also behaving normally.&lt;/p&gt;

&lt;p&gt;The CPU, however, was sitting around &lt;strong&gt;90%&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That changed the direction of the investigation.&lt;/p&gt;

&lt;p&gt;The first instinct could have been:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"The server is too small. Buy a bigger one."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the measurements were telling a different story.&lt;/p&gt;




&lt;h1&gt;
  
  
  Finding the actual bottleneck
&lt;/h1&gt;

&lt;p&gt;The resource picture looked roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RAM       &amp;lt; 1 GB
Postgres  Healthy
CPU       ~90%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The CPU was the constraint.&lt;/p&gt;

&lt;p&gt;This is an important distinction.&lt;/p&gt;

&lt;p&gt;If RAM had been nearly exhausted, adding memory would have been a reasonable direction.&lt;/p&gt;

&lt;p&gt;If PostgreSQL had been saturated, database optimization would have been the next area to investigate.&lt;/p&gt;

&lt;p&gt;But neither was the immediate problem.&lt;/p&gt;

&lt;p&gt;The application was spending too much CPU doing work for requests.&lt;/p&gt;

&lt;p&gt;So I asked a different question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How much of this work actually needs to happen for every request?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That led directly to caching.&lt;/p&gt;




&lt;h1&gt;
  
  
  Optimization #1: Cache the response in Node.js
&lt;/h1&gt;

&lt;p&gt;The feed/workflow response wasn't changing every millisecond.&lt;/p&gt;

&lt;p&gt;Yet the application was repeatedly doing the same expensive work for requests that could safely receive a recently generated response.&lt;/p&gt;

&lt;p&gt;So I introduced a very simple cache.&lt;/p&gt;

&lt;p&gt;The idea was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   |
   v
Is cached response available?
   |
  Yes ----&amp;gt; Return cached response
   |
  No
   |
   v
Query database
   |
   v
Build response
   |
   v
Cache for 1 second
   |
   v
Return response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail was the TTL:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1 second.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This wasn't an attempt to build a complicated distributed caching architecture.&lt;/p&gt;

&lt;p&gt;It was simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;If another request arrives immediately after this one, don't repeat the same work unnecessarily.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That small change had a significant effect.&lt;/p&gt;

&lt;p&gt;The system moved from roughly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2,500 concurrent users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;4,000 concurrent users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;before reaching the same kind of limitation.&lt;/p&gt;

&lt;p&gt;And I hadn't changed the server.&lt;/p&gt;

&lt;p&gt;No additional CPU.&lt;/p&gt;

&lt;p&gt;No additional RAM.&lt;/p&gt;

&lt;p&gt;No Redis cluster.&lt;/p&gt;

&lt;p&gt;No second application server.&lt;/p&gt;

&lt;p&gt;The workload had simply changed.&lt;/p&gt;




&lt;h1&gt;
  
  
  Why did a 1-second cache help so much?
&lt;/h1&gt;

&lt;p&gt;Caching works because computation has a cost.&lt;/p&gt;

&lt;p&gt;Consider a simplified request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Node.js
   ↓
PostgreSQL query
   ↓
Process data
   ↓
Serialize response
   ↓
Send response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If 100 requests arrive close together and they all need essentially the same data, doing the complete operation 100 times may be unnecessary.&lt;/p&gt;

&lt;p&gt;With a short-lived cache:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request 1
   ↓
Generate response
   ↓
Cache

Request 2 ─┐
Request 3 ─┤
Request 4 ─┤──&amp;gt; Cached response
Request 5 ─┘
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The expensive work can be shared across multiple requests.&lt;/p&gt;

&lt;p&gt;That reduces CPU work.&lt;/p&gt;

&lt;p&gt;It can also reduce database work.&lt;/p&gt;

&lt;p&gt;And because the response is already available, the application has less work to perform per request.&lt;/p&gt;

&lt;p&gt;The important point is that caching didn't make the CPU faster.&lt;/p&gt;

&lt;p&gt;It &lt;strong&gt;reduced how much CPU work was required&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  Optimization #2: Move the cache closer to the edge
&lt;/h1&gt;

&lt;p&gt;The next question was:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Why should Node.js be responsible for serving a response that Nginx can serve directly?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The architecture originally looked like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
 Nginx
  |
 Node.js
  |
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After introducing application-level caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
 Nginx
  |
 Node.js
  |
Cache
  |
PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That was already better.&lt;/p&gt;

&lt;p&gt;But there was another opportunity.&lt;/p&gt;

&lt;p&gt;If Nginx could cache the response, a cache hit wouldn't need to travel through the Node.js application at all.&lt;/p&gt;

&lt;p&gt;So the architecture became conceptually:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
 Nginx
  |
  +---- Cache hit ----&amp;gt; Response
  |
  +---- Cache miss ---&amp;gt; Node.js
                           |
                       PostgreSQL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the request could potentially be handled before reaching the application layer.&lt;/p&gt;

&lt;p&gt;This is an important optimization because every request that Nginx can satisfy is a request that doesn't consume Node.js CPU.&lt;/p&gt;




&lt;h1&gt;
  
  
  The result
&lt;/h1&gt;

&lt;p&gt;After moving the caching responsibility to Nginx, the system pushed &lt;strong&gt;past 5,000 concurrent users&lt;/strong&gt; on the same $12/month server.&lt;/p&gt;

&lt;p&gt;The progression was approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;~2,500 users
     |
     | 1-second application cache
     v
~4,000 users
     |
     | Move cache to Nginx
     v
5,000+ users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The infrastructure didn't become more powerful.&lt;/p&gt;

&lt;p&gt;The application simply stopped doing unnecessary work for every request.&lt;/p&gt;

&lt;p&gt;That distinction is the main lesson from the experiment.&lt;/p&gt;




&lt;h1&gt;
  
  
  Concurrent users are not total users
&lt;/h1&gt;

&lt;p&gt;One important clarification:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5,000 concurrent users does not mean the server can only support 5,000 users total.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Concurrency and total user count are different measurements.&lt;/p&gt;

&lt;p&gt;If an application has 100,000 registered users but only 500 are actively making requests at a given moment, the server is dealing with roughly 500 concurrent users, not 100,000.&lt;/p&gt;

&lt;p&gt;That's why saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This server supports 5,000 users"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;would be misleading.&lt;/p&gt;

&lt;p&gt;The accurate statement from this experiment is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The setup was able to push past 5,000 concurrent virtual users under this particular load-test workload.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That doesn't automatically translate into a production capacity number.&lt;/p&gt;

&lt;p&gt;Real applications have different request patterns, payload sizes, database queries, background jobs, connection behavior, and traffic distributions.&lt;/p&gt;

&lt;p&gt;Benchmarks are useful when we understand exactly what was benchmarked.&lt;/p&gt;




&lt;h1&gt;
  
  
  What I would not conclude from this experiment
&lt;/h1&gt;

&lt;p&gt;This experiment doesn't prove that everyone should run production systems on a $12 server.&lt;/p&gt;

&lt;p&gt;It doesn't.&lt;/p&gt;

&lt;p&gt;A production system may need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;High availability&lt;/li&gt;
&lt;li&gt;Backups&lt;/li&gt;
&lt;li&gt;Replication&lt;/li&gt;
&lt;li&gt;Monitoring&lt;/li&gt;
&lt;li&gt;Failover&lt;/li&gt;
&lt;li&gt;Security controls&lt;/li&gt;
&lt;li&gt;Multiple application instances&lt;/li&gt;
&lt;li&gt;Database replicas&lt;/li&gt;
&lt;li&gt;Disaster recovery&lt;/li&gt;
&lt;li&gt;Horizontal scaling&lt;/li&gt;
&lt;li&gt;Background workers&lt;/li&gt;
&lt;li&gt;Rate limiting&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those requirements are independent of whether a small server can handle a particular workload.&lt;/p&gt;

&lt;p&gt;The experiment answers a narrower question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How much performance can I get from simple infrastructure before infrastructure itself becomes the problem?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;In this case, quite a lot.&lt;/p&gt;




&lt;h1&gt;
  
  
  The bigger lesson: measure before scaling
&lt;/h1&gt;

&lt;p&gt;One of the easiest mistakes in backend engineering is solving a capacity problem before identifying the actual constraint.&lt;/p&gt;

&lt;p&gt;You see latency increasing.&lt;/p&gt;

&lt;p&gt;You add a bigger server.&lt;/p&gt;

&lt;p&gt;But what if the application is wasting CPU?&lt;/p&gt;

&lt;p&gt;You see database queries getting slower.&lt;/p&gt;

&lt;p&gt;You add a larger database instance.&lt;/p&gt;

&lt;p&gt;But what if the same expensive query is being executed thousands of times unnecessarily?&lt;/p&gt;

&lt;p&gt;You see more traffic.&lt;/p&gt;

&lt;p&gt;You add Redis.&lt;/p&gt;

&lt;p&gt;But what if a simple Nginx cache would have handled the workload?&lt;/p&gt;

&lt;p&gt;Infrastructure can absolutely solve performance problems.&lt;/p&gt;

&lt;p&gt;But infrastructure should be guided by measurements.&lt;/p&gt;

&lt;p&gt;In this experiment, the sequence was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Load test
   ↓
Observe latency
   ↓
Measure resources
   ↓
Find CPU bottleneck
   ↓
Reduce repeated work
   ↓
Test again
   ↓
Find the next limit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That loop is more valuable than any particular server size.&lt;/p&gt;




&lt;h1&gt;
  
  
  A simple scaling mindset
&lt;/h1&gt;

&lt;p&gt;When an application starts struggling, I like thinking about the problem in this order:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. What exactly is getting slower?
&lt;/h3&gt;

&lt;p&gt;Don't stop at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"The API is slow."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Average latency&lt;/li&gt;
&lt;li&gt;p95&lt;/li&gt;
&lt;li&gt;p99&lt;/li&gt;
&lt;li&gt;Requests/sec&lt;/li&gt;
&lt;li&gt;Error rate&lt;/li&gt;
&lt;li&gt;CPU&lt;/li&gt;
&lt;li&gt;Memory&lt;/li&gt;
&lt;li&gt;Database utilization&lt;/li&gt;
&lt;li&gt;Connection counts&lt;/li&gt;
&lt;li&gt;Network&lt;/li&gt;
&lt;li&gt;Disk I/O&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. What resource is actually saturated?
&lt;/h3&gt;

&lt;p&gt;Find the constraint.&lt;/p&gt;

&lt;p&gt;In this experiment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RAM     → fine
Postgres → fine
CPU     → ~90%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That pointed the investigation toward application work.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Can the work be eliminated?
&lt;/h3&gt;

&lt;p&gt;Before making hardware bigger, ask:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Are we calculating the same thing repeatedly?&lt;/li&gt;
&lt;li&gt;Can the result be cached?&lt;/li&gt;
&lt;li&gt;Can a query be avoided?&lt;/li&gt;
&lt;li&gt;Can data be precomputed?&lt;/li&gt;
&lt;li&gt;Can work move to a cheaper layer?&lt;/li&gt;
&lt;li&gt;Can Nginx/CDN handle something before it reaches the application?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  4. Test the change
&lt;/h3&gt;

&lt;p&gt;Don't assume the optimization worked.&lt;/p&gt;

&lt;p&gt;Run the same workload again.&lt;/p&gt;

&lt;p&gt;Compare the metrics.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Scale only when necessary
&lt;/h3&gt;

&lt;p&gt;Eventually, the $12 server will hit another limit.&lt;/p&gt;

&lt;p&gt;That's completely fine.&lt;/p&gt;

&lt;p&gt;Scaling isn't the failure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Scaling without understanding why you're scaling is the expensive part.&lt;/strong&gt;&lt;/p&gt;




&lt;h1&gt;
  
  
  What surprised me most
&lt;/h1&gt;

&lt;p&gt;The surprising part wasn't that a cheap server handled thousands of concurrent virtual users.&lt;/p&gt;

&lt;p&gt;It was how quickly the result changed after removing unnecessary work.&lt;/p&gt;

&lt;p&gt;At roughly 2,500 concurrent users, CPU was close to its limit.&lt;/p&gt;

&lt;p&gt;A one-second cache increased the tested capacity to roughly 4,000.&lt;/p&gt;

&lt;p&gt;Moving that cache to Nginx pushed the test past 5,000.&lt;/p&gt;

&lt;p&gt;Same machine.&lt;/p&gt;

&lt;p&gt;Same CPU.&lt;/p&gt;

&lt;p&gt;Same RAM.&lt;/p&gt;

&lt;p&gt;The biggest change was the amount of work the application had to perform.&lt;/p&gt;

&lt;p&gt;That's a useful reminder for any backend system:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Before adding capacity, find out whether you can remove the work creating the demand.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Sometimes the cheapest performance optimization isn't a bigger server.&lt;/p&gt;

&lt;p&gt;It's one less thing your server has to do.&lt;/p&gt;

</description>
      <category>cloud</category>
      <category>webdev</category>
      <category>devops</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Jev: AI for Decisions, Not Just Generation</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Thu, 24 Sep 2026 18:05:32 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/jev-ai-for-decisions-not-just-generation-3j88</link>
      <guid>https://dev.to/gaurav_talesara/jev-ai-for-decisions-not-just-generation-3j88</guid>
      <description>&lt;p&gt;Most AI applications today are built around generation.&lt;/p&gt;

&lt;p&gt;Give a model some context → ask a question → get text back.&lt;/p&gt;

&lt;p&gt;But many application workflows don't actually need generated text.&lt;/p&gt;

&lt;p&gt;They need a &lt;strong&gt;decision&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Should this support ticket go to billing or technical support?&lt;/p&gt;

&lt;p&gt;Is this lead worth prioritizing?&lt;/p&gt;

&lt;p&gt;Should an AI agent retry a failed action?&lt;/p&gt;

&lt;p&gt;Does this content need human review?&lt;/p&gt;

&lt;p&gt;Which tool should an agent call next?&lt;/p&gt;

&lt;p&gt;This is where Jev gets interesting.&lt;/p&gt;

&lt;p&gt;Instead of treating every AI problem as a text-generation problem, Jev is designed around &lt;strong&gt;decision-making&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You provide the current state and typed questions, and the model can return structured decisions along with probabilities and confidence.&lt;/p&gt;

&lt;p&gt;A simplified flow looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Application State
       ↓
      Jev
       ↓
Decision + Probability + Confidence
       ↓
Application Action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"route"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"billing"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"probability"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.92&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"high"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Your application can then use that output directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where this could be useful
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Support automation&lt;/strong&gt;&lt;br&gt;
→ classify and route incoming tickets&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Sales systems&lt;/strong&gt;&lt;br&gt;
→ score and prioritize leads&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI agents&lt;/strong&gt;&lt;br&gt;
→ decide which tool or action should happen next&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Content pipelines&lt;/strong&gt;&lt;br&gt;
→ identify what requires human review&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Risk workflows&lt;/strong&gt;&lt;br&gt;
→ evaluate a situation before allowing an automated action&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Data processing&lt;/strong&gt;&lt;br&gt;
→ classify and triage unstructured inputs&lt;/p&gt;

&lt;p&gt;The interesting architectural idea isn't simply adding another AI model.&lt;/p&gt;

&lt;p&gt;It's separating &lt;strong&gt;generation&lt;/strong&gt; from &lt;strong&gt;decision-making&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A general-purpose LLM is great when the output needs to be language.&lt;/p&gt;

&lt;p&gt;But when the application needs something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;route = billing
retry = false
risk = medium
next_action = create_ticket
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;a decision-oriented model can fit the workflow much more naturally.&lt;/p&gt;

&lt;p&gt;I'm interested to see where this pattern goes as AI applications move from chat interfaces toward systems that continuously make decisions and take actions.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>jev</category>
      <category>llm</category>
    </item>
    <item>
      <title>Voice AI can sound human and still fail at sales.

The real problem is often conversation logic: understanding customer state, adapting questions, and knowing when to hand off to a human.

I wrote about the architecture behind that.</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Sun, 13 Sep 2026 12:01:28 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/voice-ai-can-sound-human-and-still-fail-at-sales-the-real-problem-is-often-conversation-logic-4k7l</link>
      <guid>https://dev.to/gaurav_talesara/voice-ai-can-sound-human-and-still-fail-at-sales-the-real-problem-is-often-conversation-logic-4k7l</guid>
      <description>&lt;div class="ltag__link--embedded"&gt;
  &lt;div class="crayons-story "&gt;
  &lt;a href="https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9" class="crayons-story__hidden-navigation-link"&gt;Voice AI Doesn't Have a Voice Problem. It Has a Conversation Problem.&lt;/a&gt;


  &lt;div class="crayons-story__body crayons-story__body-full_post"&gt;
    &lt;div class="crayons-story__top"&gt;
      &lt;div class="crayons-story__meta"&gt;
        &lt;div class="crayons-story__author-pic"&gt;

          &lt;a href="/gaurav_talesara" class="crayons-avatar  crayons-avatar--l  "&gt;
            &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3724256%2F3ab1fc88-aa7b-426c-9783-c019bd8e5915.png" alt="gaurav_talesara profile" class="crayons-avatar__image"&gt;
          &lt;/a&gt;
        &lt;/div&gt;
        &lt;div&gt;
          &lt;div&gt;
            &lt;a href="/gaurav_talesara" class="crayons-story__secondary fw-medium m:hidden"&gt;
              Gaurav Talesara
            &lt;/a&gt;
            &lt;div class="profile-preview-card relative mb-4 s:mb-0 fw-medium hidden m:inline-block"&gt;
              
                Gaurav Talesara
                
                
              
              &lt;div id="story-author-preview-content-4643467" class="profile-preview-card__content crayons-dropdown branded-7 p-4 pt-0"&gt;
                &lt;div class="gap-4 grid"&gt;
                  &lt;div class="-mt-4"&gt;
                    &lt;a href="/gaurav_talesara" class="flex"&gt;
                      &lt;span class="crayons-avatar crayons-avatar--xl mr-2 shrink-0"&gt;
                        &lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3724256%2F3ab1fc88-aa7b-426c-9783-c019bd8e5915.png" class="crayons-avatar__image" alt=""&gt;
                      &lt;/span&gt;
                      &lt;span class="crayons-link crayons-subtitle-2 mt-5"&gt;Gaurav Talesara&lt;/span&gt;
                    &lt;/a&gt;
                  &lt;/div&gt;
                  &lt;div class="print-hidden"&gt;
                    
                      Follow
                    
                  &lt;/div&gt;
                  &lt;div class="author-preview-metadata-container"&gt;&lt;/div&gt;
                &lt;/div&gt;
              &lt;/div&gt;
            &lt;/div&gt;

          &lt;/div&gt;
          &lt;a href="https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9" class="crayons-story__tertiary fs-xs"&gt;&lt;time&gt;Sep 13&lt;/time&gt;&lt;span class="time-ago-indicator-initial-placeholder"&gt;&lt;/span&gt;&lt;/a&gt;
        &lt;/div&gt;
      &lt;/div&gt;

    &lt;/div&gt;

    &lt;div class="crayons-story__indention"&gt;
      &lt;h2 class="crayons-story__title crayons-story__title-full_post"&gt;
        &lt;a href="https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9" id="article-link-4643467"&gt;
          Voice AI Doesn't Have a Voice Problem. It Has a Conversation Problem.
        &lt;/a&gt;
      &lt;/h2&gt;
        &lt;div class="crayons-story__tags"&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/ai"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;ai&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/automation"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;automation&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/architecture"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;architecture&lt;/a&gt;
            &lt;a class="crayons-tag  crayons-tag--monochrome " href="/t/performance"&gt;&lt;span class="crayons-tag__prefix"&gt;#&lt;/span&gt;performance&lt;/a&gt;
        &lt;/div&gt;
      &lt;div class="crayons-story__bottom"&gt;
        &lt;div class="crayons-story__details"&gt;
            &lt;a href="https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9#comments" class="crayons-btn crayons-btn--s crayons-btn--ghost crayons-btn--icon-left flex items-center"&gt;
              

              2&lt;span class="hidden s:inline"&gt;&amp;nbsp;comments&lt;/span&gt;
            &lt;/a&gt;
        &lt;/div&gt;
        &lt;div class="crayons-story__save"&gt;
          &lt;small class="crayons-story__tertiary fs-xs mr-2"&gt;
            5 min read
          &lt;/small&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;


</description>
    </item>
    <item>
      <title>Voice AI Doesn't Have a Voice Problem. It Has a Conversation Problem.</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Sun, 13 Sep 2026 11:55:18 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9</link>
      <guid>https://dev.to/gaurav_talesara/voice-ai-doesnt-have-a-voice-problem-it-has-a-conversation-problem-pm9</guid>
      <description>&lt;h2&gt;
  
  
  Voice AI can sound completely human and still be terrible at sales.
&lt;/h2&gt;



&lt;p&gt;The voice can be natural, the response time can be fast, and the agent can handle thousands of calls. It can qualify leads, answer questions, schedule follow-ups, and update the CRM.&lt;/p&gt;

&lt;p&gt;And the sales results can still be disappointing.&lt;/p&gt;

&lt;p&gt;The problem is that sounding human and having a good sales conversation are two different engineering problems.&lt;/p&gt;

&lt;p&gt;A customer rarely follows a predefined script. They interrupt, change direction, raise objections, ask their own questions, or reveal the most important information halfway through the conversation.&lt;/p&gt;

&lt;p&gt;That is where many Voice AI sales systems start to struggle.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem with linear sales conversations
&lt;/h2&gt;

&lt;p&gt;A typical sales agent follows a simple flow: introduce the company, ask qualification questions, explain the product, handle objections, and try to book a meeting.&lt;/p&gt;

&lt;p&gt;It looks reasonable on a whiteboard.&lt;/p&gt;

&lt;p&gt;Real customers don't behave that way.&lt;/p&gt;

&lt;p&gt;Two customers can give the same answer but need completely different follow-ups. One may already be comparing vendors and ready to buy. Another may only be researching because a problem recently appeared.&lt;/p&gt;

&lt;p&gt;A rigid workflow sees the same answer and moves to the next predefined question.&lt;/p&gt;

&lt;p&gt;A good salesperson thinks about what the answer actually means.&lt;/p&gt;

&lt;p&gt;That difference is the core problem.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Model the customer, not just the lead
&lt;/h2&gt;

&lt;p&gt;Most sales systems already have plenty of customer data: lead score, company information, previous interactions, campaign source, product interest, and CRM history.&lt;/p&gt;

&lt;p&gt;That information is useful, but the agent still needs to answer one question during the conversation:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What should happen next?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For that, the system needs conversational state.&lt;/p&gt;

&lt;p&gt;The agent should understand what the customer is trying to solve, how urgent the problem is, whether they are actively evaluating solutions, what they already know, and which objections have appeared.&lt;/p&gt;

&lt;p&gt;That state should change as the conversation develops.&lt;/p&gt;

&lt;p&gt;A customer can move from unfamiliar with the product to curious, then to problem-aware and eventually ready for a sales conversation. The opposite can happen too. A customer may discover that the timing is wrong or that the product isn't a fit.&lt;/p&gt;

&lt;p&gt;The system should respond to those changes instead of forcing everyone through the same flow.&lt;/p&gt;

&lt;p&gt;This is why I think a Voice AI sales agent needs a state machine, not a better prompt.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Make every answer useful
&lt;/h2&gt;

&lt;p&gt;The next question should depend on what the customer just said.&lt;/p&gt;

&lt;p&gt;Suppose a customer says price is the main concern.&lt;/p&gt;

&lt;p&gt;The agent shouldn't simply record "price objection" and continue with the script. It should understand what is behind that answer.&lt;/p&gt;

&lt;p&gt;Maybe the customer is comparing vendors. Maybe the budget isn't approved. Maybe the value isn't clear. Maybe they are interested but not ready to commit.&lt;/p&gt;

&lt;p&gt;Each situation requires a different response.&lt;/p&gt;

&lt;p&gt;The same applies to positive signals.&lt;/p&gt;

&lt;p&gt;If someone says they are interested, that doesn't automatically mean they want a demo. They may be researching the market or trying to understand whether the product solves a specific problem.&lt;/p&gt;

&lt;p&gt;The job of the next question is to reduce that uncertainty.&lt;/p&gt;

&lt;p&gt;This is where conversation logic becomes more important than prompt length.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Optimize for progression, not meetings
&lt;/h2&gt;

&lt;p&gt;A common mistake is making "book a meeting" the main objective.&lt;/p&gt;

&lt;p&gt;That can push the agent to ask for a meeting even when the customer isn't ready.&lt;/p&gt;

&lt;p&gt;A better objective is progression.&lt;/p&gt;

&lt;p&gt;The next useful action might be a meeting. It could also be sending pricing information, arranging a callback, connecting the customer with a specialist, or ending the conversation because there is no fit.&lt;/p&gt;

&lt;p&gt;Not every cold lead should become a warm lead.&lt;/p&gt;

&lt;p&gt;The goal is to reduce wasted conversations and increase the number of interactions that move toward a useful sales outcome.&lt;/p&gt;

&lt;p&gt;That also changes how the system should be measured. Call volume and conversation duration are useful operational metrics, but they don't tell you whether the sales process is improving.&lt;/p&gt;

&lt;p&gt;Meaningful conversations, qualified leads, warm-lead progression, appropriate human handoffs, meetings that turn into opportunities, and revenue influenced by the system are much more useful measures.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Human handoff is part of the product
&lt;/h2&gt;

&lt;p&gt;The human handoff is another place where Voice AI systems often lose context.&lt;/p&gt;

&lt;p&gt;Transferring a call to a salesperson is easy.&lt;/p&gt;

&lt;p&gt;Transferring the customer's understanding is harder.&lt;/p&gt;

&lt;p&gt;If the customer has spent five minutes explaining their situation and the salesperson starts the conversation from zero, the customer has to repeat everything.&lt;/p&gt;

&lt;p&gt;A good handoff should carry the relevant context: the customer's problem, intent, objections, urgency, important questions, and why the AI decided that a human should take over.&lt;/p&gt;

&lt;p&gt;The salesperson should be able to continue the conversation instead of restarting it.&lt;/p&gt;

&lt;p&gt;This is especially important for technical products where the AI may identify an integration, security, or implementation concern that should be handled by a specialist.&lt;/p&gt;

&lt;p&gt;The handoff is not an exception to the product. It is part of the product design.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Build the feedback loop
&lt;/h2&gt;

&lt;p&gt;The call itself isn't the final outcome.&lt;/p&gt;

&lt;p&gt;What happens after the call tells you whether the conversation logic worked.&lt;/p&gt;

&lt;p&gt;Did the salesperson accept the lead? Was the customer qualified? Did the meeting happen? Did the opportunity progress? Was the lead rejected because there was no fit?&lt;/p&gt;

&lt;p&gt;Those outcomes should flow back into the system.&lt;/p&gt;

&lt;p&gt;Without that feedback, teams can end up optimizing activity instead of results.&lt;/p&gt;

&lt;p&gt;A Voice AI agent may handle thousands of calls and produce impressive transcripts, but if qualified opportunities aren't increasing, the system still has a sales problem.&lt;/p&gt;

&lt;p&gt;The feedback loop connects conversation behavior to actual business outcomes.&lt;br&gt;
&lt;br&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture matters more than the prompt
&lt;/h2&gt;

&lt;p&gt;From an engineering perspective, I would separate customer context, conversation state, decision logic, response generation, human handoff, and outcome feedback.&lt;/p&gt;

&lt;p&gt;The prompt should not become the place where the entire sales process lives.&lt;/p&gt;

&lt;p&gt;Once the prompt contains customer-state management, qualification rules, CRM behavior, escalation logic, and every possible objection, it becomes difficult to test and maintain.&lt;/p&gt;

&lt;p&gt;A state-driven design gives the team clearer boundaries. You can test whether the system understood the customer's intent, whether the next action was appropriate, and whether the handoff happened at the right time.&lt;/p&gt;

&lt;p&gt;You can then improve each part without rewriting the entire conversation.&lt;/p&gt;

&lt;p&gt;Voice quality will continue to improve. Voices will become more natural and models will handle more complex conversations.&lt;/p&gt;

&lt;p&gt;But the harder problem is deciding what the agent should do with everything it hears.&lt;/p&gt;

&lt;p&gt;That requires customer context, conversational state, adaptive decision-making, clear handoff boundaries, and a feedback loop connected to real sales outcomes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A Voice AI sales agent doesn't need a better script. It needs a better model of the conversation.&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;Building AI systems? I’d love to connect.&lt;/p&gt;

&lt;p&gt;Website - &lt;a href="https://gauravspace.in/" rel="noopener noreferrer"&gt;gauravspace.in&lt;/a&gt; &lt;br&gt;
&lt;a href="https://www.linkedin.com/in/gaurav-talesara-8099ba147/" rel="noopener noreferrer"&gt;Connect on LinkedIn&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Gaurav Talesara&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>architecture</category>
      <category>performance</category>
    </item>
    <item>
      <title>The Website Is No Longer the Center of Commerce</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Wed, 02 Sep 2026 18:59:10 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/the-website-is-no-longer-the-center-of-commerce-2agi</link>
      <guid>https://dev.to/gaurav_talesara/the-website-is-no-longer-the-center-of-commerce-2agi</guid>
      <description>&lt;p&gt;For years, digital commerce has been built around a simple assumption:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Get the user to the website.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Once the user arrives, guide them through a carefully designed journey.&lt;/p&gt;

&lt;p&gt;Search.&lt;/p&gt;

&lt;p&gt;Landing page.&lt;/p&gt;

&lt;p&gt;Product page.&lt;/p&gt;

&lt;p&gt;Cart.&lt;/p&gt;

&lt;p&gt;Checkout.&lt;/p&gt;

&lt;p&gt;The website was the center of the system.&lt;/p&gt;

&lt;p&gt;But that assumption is starting to change.&lt;/p&gt;

&lt;p&gt;AI-powered interfaces are creating a different model of interaction—one where users may discover products, evaluate options, and potentially complete transactions without following the traditional journey through a company's website.&lt;/p&gt;

&lt;p&gt;The journey is becoming something closer to:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Intent → AI → Discovery → Decision → Transaction&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The website may still exist.&lt;/p&gt;

&lt;p&gt;It may still be important.&lt;/p&gt;

&lt;p&gt;But it may no longer be the only interface between a business and its customers.&lt;/p&gt;

&lt;p&gt;As an engineering leader, I find the architectural implications of this shift more interesting than the interface itself.&lt;/p&gt;

&lt;p&gt;Because if AI becomes another layer between users and businesses, we need to ask a different question.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How do we build a better website?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;We may need to start asking:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;How does our product exist outside our website?&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  From page-centric products to capability-centric systems
&lt;/h2&gt;

&lt;p&gt;Traditional web architecture is often built around pages.&lt;/p&gt;

&lt;p&gt;A user visits a page.&lt;/p&gt;

&lt;p&gt;The page loads data.&lt;/p&gt;

&lt;p&gt;The user performs an action.&lt;/p&gt;

&lt;p&gt;The backend processes that action.&lt;/p&gt;

&lt;p&gt;The journey is designed around a human navigating an interface.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
  ↓
Website
  ↓
Product Page
  ↓
Cart
  ↓
Checkout
  ↓
Transaction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The interface controls the experience.&lt;/p&gt;

&lt;p&gt;But AI introduces another possibility.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
  ↓
Intent
  ↓
AI Interface / Agent
  ↓
Business Systems
  ↓
Decision
  ↓
Transaction
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a fundamentally different interaction model.&lt;/p&gt;

&lt;p&gt;The AI does not necessarily need to understand how your website looks.&lt;/p&gt;

&lt;p&gt;It needs to understand what your business can do.&lt;/p&gt;

&lt;p&gt;What products do you offer?&lt;/p&gt;

&lt;p&gt;What are the prices?&lt;/p&gt;

&lt;p&gt;What is available?&lt;/p&gt;

&lt;p&gt;What are the policies?&lt;/p&gt;

&lt;p&gt;What actions can be performed?&lt;/p&gt;

&lt;p&gt;What restrictions exist?&lt;/p&gt;

&lt;p&gt;What systems can be accessed safely?&lt;/p&gt;

&lt;p&gt;This moves part of the engineering problem away from &lt;strong&gt;pages&lt;/strong&gt; and toward &lt;strong&gt;capabilities&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Your product needs to become understandable
&lt;/h2&gt;

&lt;p&gt;Humans can navigate ambiguity.&lt;/p&gt;

&lt;p&gt;They can look at a product page, interpret an image, read descriptions, compare options, and understand context.&lt;/p&gt;

&lt;p&gt;Software systems cannot rely on that in the same way.&lt;/p&gt;

&lt;p&gt;AI systems and agents need structured information.&lt;/p&gt;

&lt;p&gt;If an external AI system needs to interact with your product, your underlying data becomes increasingly important.&lt;/p&gt;

&lt;p&gt;For commerce, that could include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Product information&lt;/li&gt;
&lt;li&gt;Categories and attributes&lt;/li&gt;
&lt;li&gt;Pricing&lt;/li&gt;
&lt;li&gt;Availability&lt;/li&gt;
&lt;li&gt;Inventory&lt;/li&gt;
&lt;li&gt;Offers&lt;/li&gt;
&lt;li&gt;Shipping information&lt;/li&gt;
&lt;li&gt;Return policies&lt;/li&gt;
&lt;li&gt;Product compatibility&lt;/li&gt;
&lt;li&gt;Restrictions&lt;/li&gt;
&lt;li&gt;Business rules&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The quality of the user interface still matters.&lt;/p&gt;

&lt;p&gt;But structured information behind the interface may become equally important.&lt;/p&gt;

&lt;p&gt;A beautifully designed website is not enough if an external system cannot understand what the business actually offers.&lt;/p&gt;




&lt;h2&gt;
  
  
  APIs may become part of the product experience
&lt;/h2&gt;

&lt;p&gt;For a long time, APIs were mostly considered technical infrastructure.&lt;/p&gt;

&lt;p&gt;Something used by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Mobile applications&lt;/li&gt;
&lt;li&gt;Internal systems&lt;/li&gt;
&lt;li&gt;Third-party integrations&lt;/li&gt;
&lt;li&gt;Partner platforms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But as AI agents become more capable of interacting with external systems, APIs and system capabilities could increasingly become part of the product experience itself.&lt;/p&gt;

&lt;p&gt;The question is no longer only:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can our frontend perform this action?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It may become:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Can a trusted external system understand and safely perform this action?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Check availability
Get price
Apply offer
Create order
Calculate delivery
Process payment
Track order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These capabilities should not exist only inside tightly coupled frontend flows.&lt;/p&gt;

&lt;p&gt;They need to be represented clearly within the system.&lt;/p&gt;

&lt;p&gt;This does not mean every business suddenly needs a public API for everything.&lt;/p&gt;

&lt;p&gt;Security, authentication, permissions, and business risk still matter.&lt;/p&gt;

&lt;p&gt;But the architectural mindset is changing.&lt;/p&gt;

&lt;p&gt;The system needs to expose capabilities in a way that can be understood and controlled.&lt;/p&gt;




&lt;h2&gt;
  
  
  The architecture behind AI commerce
&lt;/h2&gt;

&lt;p&gt;A useful way to think about the future architecture is through four layers.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Interface Layer
&lt;/h2&gt;

&lt;p&gt;This is where the user interacts.&lt;/p&gt;

&lt;p&gt;It could be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A website&lt;/li&gt;
&lt;li&gt;A mobile app&lt;/li&gt;
&lt;li&gt;Search&lt;/li&gt;
&lt;li&gt;An AI assistant&lt;/li&gt;
&lt;li&gt;A conversational interface&lt;/li&gt;
&lt;li&gt;A voice interface&lt;/li&gt;
&lt;li&gt;A third-party agent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important point is that the business should not depend entirely on one interface.&lt;/p&gt;

&lt;p&gt;Interfaces can change.&lt;/p&gt;

&lt;p&gt;The core capabilities should remain stable.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Intelligence Layer
&lt;/h2&gt;

&lt;p&gt;This layer understands intent.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"I need a laptop for software development under a certain budget."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The system needs to interpret:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;User requirements&lt;/li&gt;
&lt;li&gt;Budget&lt;/li&gt;
&lt;li&gt;Product preferences&lt;/li&gt;
&lt;li&gt;Constraints&lt;/li&gt;
&lt;li&gt;Available options&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An AI system may help translate human intent into structured actions.&lt;/p&gt;

&lt;p&gt;But AI should not be responsible for inventing the truth.&lt;/p&gt;

&lt;p&gt;It should retrieve information from reliable systems.&lt;/p&gt;

&lt;p&gt;That distinction is important.&lt;/p&gt;

&lt;p&gt;AI can interpret.&lt;/p&gt;

&lt;p&gt;Your systems should remain the source of truth.&lt;/p&gt;




&lt;h2&gt;
  
  
  3. Business Capability Layer
&lt;/h2&gt;

&lt;p&gt;This is where the actual business operations exist.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Product Catalog
Pricing Engine
Inventory System
Order Management
Payment System
Shipping System
Customer Management
Business Rules
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This layer should contain the actual capabilities of the business.&lt;/p&gt;

&lt;p&gt;The interface should consume these capabilities.&lt;/p&gt;

&lt;p&gt;An AI agent may consume them too.&lt;/p&gt;

&lt;p&gt;The website is just one consumer.&lt;/p&gt;

&lt;p&gt;That is the architectural shift I find particularly interesting.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Trust and Control Layer
&lt;/h2&gt;

&lt;p&gt;Once AI systems begin taking actions, trust becomes a core engineering problem.&lt;/p&gt;

&lt;p&gt;An AI should not simply have unrestricted access to a system.&lt;/p&gt;

&lt;p&gt;Businesses need to think about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication&lt;/li&gt;
&lt;li&gt;Authorization&lt;/li&gt;
&lt;li&gt;Permission boundaries&lt;/li&gt;
&lt;li&gt;Transaction limits&lt;/li&gt;
&lt;li&gt;Human confirmation&lt;/li&gt;
&lt;li&gt;Audit trails&lt;/li&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;li&gt;Error handling&lt;/li&gt;
&lt;li&gt;Rollbacks&lt;/li&gt;
&lt;li&gt;Fraud prevention&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The more autonomous the system becomes, the more important these controls become.&lt;/p&gt;

&lt;p&gt;AI agents may create a more flexible interface.&lt;/p&gt;

&lt;p&gt;But flexibility without control creates risk.&lt;/p&gt;




&lt;h2&gt;
  
  
  The website becomes one interface among many
&lt;/h2&gt;

&lt;p&gt;I don't think websites are disappearing.&lt;/p&gt;

&lt;p&gt;People will continue using websites.&lt;/p&gt;

&lt;p&gt;Brands will continue designing experiences.&lt;/p&gt;

&lt;p&gt;Companies will continue optimize their products for humans.&lt;/p&gt;

&lt;p&gt;But the website may gradually lose its position as the &lt;strong&gt;only center of the digital experience&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A product could increasingly exist across multiple interfaces.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;                 Website
                    │
                    │
Search ─────── Business Systems ─────── Mobile App
                    │
                    │
               AI Agents
                    │
                    │
             Voice Interfaces
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The business system becomes the foundation.&lt;/p&gt;

&lt;p&gt;Interfaces become ways of accessing it.&lt;/p&gt;

&lt;p&gt;This architecture is more resilient to changes in user behavior.&lt;/p&gt;

&lt;p&gt;If a new interface becomes important, you don't need to rebuild the entire business.&lt;/p&gt;

&lt;p&gt;You need to connect that interface to well-defined capabilities.&lt;/p&gt;




&lt;h2&gt;
  
  
  The engineering challenge is bigger than adding AI
&lt;/h2&gt;

&lt;p&gt;This is why I don't think the right response is simply:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Let's add an AI chatbot.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The bigger opportunity is to look at the architecture underneath the product.&lt;/p&gt;

&lt;p&gt;Can the system clearly represent the business?&lt;/p&gt;

&lt;p&gt;Can capabilities be accessed independently?&lt;/p&gt;

&lt;p&gt;Are APIs reliable?&lt;/p&gt;

&lt;p&gt;Is product data structured?&lt;/p&gt;

&lt;p&gt;Are business rules centralized?&lt;/p&gt;

&lt;p&gt;Can actions be safely executed?&lt;/p&gt;

&lt;p&gt;Can the system explain what happened?&lt;/p&gt;

&lt;p&gt;Can failures be detected and recovered?&lt;/p&gt;

&lt;p&gt;These are engineering questions.&lt;/p&gt;

&lt;p&gt;And they will become increasingly important as AI moves from generating answers to coordinating actions.&lt;/p&gt;




&lt;h2&gt;
  
  
  The real competitive advantage may move deeper into the stack
&lt;/h2&gt;

&lt;p&gt;For many years, companies competed heavily through the interface.&lt;/p&gt;

&lt;p&gt;Better UX.&lt;/p&gt;

&lt;p&gt;Better design.&lt;/p&gt;

&lt;p&gt;Better onboarding.&lt;/p&gt;

&lt;p&gt;Better conversion funnels.&lt;/p&gt;

&lt;p&gt;Those things will remain important.&lt;/p&gt;

&lt;p&gt;But AI-driven interfaces could make the underlying system architecture increasingly visible.&lt;/p&gt;

&lt;p&gt;If an AI agent cannot understand your product, access reliable information, or interact safely with your capabilities, you may become harder to discover and transact with outside your own interface.&lt;/p&gt;

&lt;p&gt;The competitive advantage may increasingly come from how well your business capabilities are represented inside the system.&lt;/p&gt;

&lt;p&gt;Not just how good the website looks.&lt;/p&gt;




&lt;h2&gt;
  
  
  A question I think engineering teams should start asking
&lt;/h2&gt;

&lt;p&gt;When building a new product, we often ask:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What screens do we need?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Maybe we should increasingly ask another question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;What capabilities does this business need to expose?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Instead of starting with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Product Page
↓
Cart Page
↓
Checkout Page
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start by understanding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Discover Product
Get Product Details
Check Availability
Calculate Price
Apply Offer
Create Order
Authorize Payment
Arrange Delivery
Track Order
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the capabilities are clear, different interfaces can be built around them.&lt;/p&gt;

&lt;p&gt;A website.&lt;/p&gt;

&lt;p&gt;A mobile app.&lt;/p&gt;

&lt;p&gt;An AI assistant.&lt;/p&gt;

&lt;p&gt;A partner integration.&lt;/p&gt;

&lt;p&gt;A voice interface.&lt;/p&gt;

&lt;p&gt;The interfaces may change.&lt;/p&gt;

&lt;p&gt;The business capabilities remain.&lt;/p&gt;




&lt;h2&gt;
  
  
  The interface is becoming more flexible. The system needs to become more structured.
&lt;/h2&gt;

&lt;p&gt;That is the part of this shift that interests me most.&lt;/p&gt;

&lt;p&gt;AI-powered commerce is not only a change in search, advertising, or checkout.&lt;/p&gt;

&lt;p&gt;It could push engineering teams toward a different way of designing products.&lt;/p&gt;

&lt;p&gt;One where products need to serve both:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Humans interacting through interfaces.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And increasingly:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI systems acting on behalf of humans.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The website is not disappearing.&lt;/p&gt;

&lt;p&gt;But it may become one interface among many.&lt;/p&gt;

&lt;p&gt;And the businesses best positioned for that future may be the ones that have built systems with clear data, well-defined capabilities, strong controls, and reliable foundations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The interface is becoming more flexible.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The system behind it needs to become more structured.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;And I think that's going to be one of the most interesting engineering challenges of the next few years.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>website</category>
      <category>google</category>
    </item>
    <item>
      <title>15 Things CA Firms Can Automate With AI — Starting With the Work Nobody Wants to Do</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Fri, 28 Aug 2026 17:26:27 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/15-things-ca-firms-can-automate-with-ai-starting-with-the-work-nobody-wants-to-do-499d</link>
      <guid>https://dev.to/gaurav_talesara/15-things-ca-firms-can-automate-with-ai-starting-with-the-work-nobody-wants-to-do-499d</guid>
      <description>&lt;p&gt;&lt;em&gt;WhatsApp follow-ups, missing documents, GST workflows, loan cases, bank queries, fee collection and more.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;A client sends a WhatsApp message:&lt;/p&gt;

&lt;p&gt;“Sir, GST ka kya hua?”&lt;/p&gt;

&lt;p&gt;Someone from the CA firm's team opens WhatsApp, searches for the client, checks an Excel sheet, looks through Google Drive, asks another team member, and finally replies.&lt;/p&gt;

&lt;p&gt;The answer may take two minutes.&lt;/p&gt;

&lt;p&gt;Finding the answer can take ten.&lt;/p&gt;

&lt;p&gt;Now multiply that by hundreds of clients.&lt;/p&gt;

&lt;p&gt;This is where I think the real opportunity for AI in CA firms begins.&lt;/p&gt;

&lt;p&gt;Not with another chatbot.&lt;/p&gt;

&lt;p&gt;With the repetitive work happening around the CA every single day.&lt;/p&gt;

&lt;h2&gt;
  
  
  AI Shouldn't Replace the CA. It Should Remove the Work Around the CA.
&lt;/h2&gt;

&lt;p&gt;A CA's time is valuable.&lt;/p&gt;

&lt;p&gt;Professional judgment, tax planning, financial advice, audits, client relationships and complex decisions require experience.&lt;/p&gt;

&lt;p&gt;But many activities surrounding that work are repetitive.&lt;/p&gt;

&lt;p&gt;Someone has to ask for documents.&lt;/p&gt;

&lt;p&gt;Someone has to send reminders.&lt;/p&gt;

&lt;p&gt;Someone has to check whether documents arrived.&lt;/p&gt;

&lt;p&gt;Someone has to update an Excel sheet.&lt;/p&gt;

&lt;p&gt;Someone has to follow up with a client.&lt;/p&gt;

&lt;p&gt;Someone has to track a bank query.&lt;/p&gt;

&lt;p&gt;Someone has to remember which loan case is waiting for what.&lt;/p&gt;

&lt;p&gt;These are exactly the kinds of workflows where automation can help.&lt;/p&gt;

&lt;p&gt;The question isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Where can we add AI?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;A better question is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“What does the team repeatedly do every day that doesn't actually require professional judgment?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Here are 15 areas worth exploring.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Client Document Collection
&lt;/h2&gt;

&lt;p&gt;This is probably one of the easiest places to start.&lt;/p&gt;

&lt;p&gt;CA firms constantly need documents from clients: bank statements, purchase registers, sales data, invoices, investment proofs, previous ITRs, GST information and more.&lt;/p&gt;

&lt;p&gt;The difficult part isn't asking once.&lt;/p&gt;

&lt;p&gt;It's keeping track of what's still missing across hundreds of clients.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client: Patel Textiles

Required: 10 documents
Received: 8
Verified: 7
Missing: 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An automation can send the request through WhatsApp or email.&lt;/p&gt;

&lt;p&gt;When the client uploads a document, the system can identify it, attach it to the correct client, update the checklist and request whatever is still missing.&lt;/p&gt;

&lt;p&gt;The staff no longer has to remember every outstanding document.&lt;/p&gt;

&lt;p&gt;The system does.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. AI Document Classification and Verification
&lt;/h2&gt;

&lt;p&gt;Receiving a document is only half the problem.&lt;/p&gt;

&lt;p&gt;The next question is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Did the client send the right document?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI can classify uploaded files and perform basic checks before they reach the CA team.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Uploaded: bank_statement_august.pdf

Document type: Bank Statement
Client: Patel Textiles
Period: August
Readable: Yes
Status: Verified
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Expected: Bank Statement

Received: Invoice

Status: Wrong Document
Action: Request correct document
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This can save staff from manually opening and sorting hundreds of files.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. WhatsApp Client Query Automation
&lt;/h2&gt;

&lt;p&gt;WhatsApp has become one of the most important communication channels for many Indian businesses.&lt;/p&gt;

&lt;p&gt;CA firms receive questions throughout the day:&lt;/p&gt;

&lt;p&gt;“What documents do I need?”&lt;/p&gt;

&lt;p&gt;“When is the GST deadline?”&lt;/p&gt;

&lt;p&gt;“Did you receive my statement?”&lt;/p&gt;

&lt;p&gt;“What is pending from my side?”&lt;/p&gt;

&lt;p&gt;“Has my return been filed?”&lt;/p&gt;

&lt;p&gt;An AI assistant can handle the first level of these questions.&lt;/p&gt;

&lt;p&gt;But there is an important difference between a generic chatbot and a useful CA assistant.&lt;/p&gt;

&lt;p&gt;A generic chatbot knows what GST is.&lt;/p&gt;

&lt;p&gt;A useful assistant knows that &lt;strong&gt;this particular client's GST filing is waiting for two documents&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Context is the real value.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Client Onboarding
&lt;/h2&gt;

&lt;p&gt;New clients create a surprising amount of administrative work.&lt;/p&gt;

&lt;p&gt;A new client may require:&lt;/p&gt;

&lt;p&gt;KYC.&lt;/p&gt;

&lt;p&gt;PAN.&lt;/p&gt;

&lt;p&gt;GST details.&lt;/p&gt;

&lt;p&gt;Previous returns.&lt;/p&gt;

&lt;p&gt;Accounting data.&lt;/p&gt;

&lt;p&gt;Bank statements.&lt;/p&gt;

&lt;p&gt;Engagement documentation.&lt;/p&gt;

&lt;p&gt;Service selection.&lt;/p&gt;

&lt;p&gt;Team assignment.&lt;/p&gt;

&lt;p&gt;Instead of managing this through spreadsheets and messages, an onboarding workflow can track everything.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client Onboarding: 72%

Completed:
KYC
PAN
GST

Pending:
Previous ITR
Bank Statements
Accounting Data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The client sees what is pending.&lt;/p&gt;

&lt;p&gt;The staff sees what is incomplete.&lt;/p&gt;

&lt;p&gt;The partner sees the overall status.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. GST Compliance Readiness
&lt;/h2&gt;

&lt;p&gt;Compliance automation shouldn't stop at sending deadline reminders.&lt;/p&gt;

&lt;p&gt;Knowing that a GST return is due in five days is useful.&lt;/p&gt;

&lt;p&gt;Knowing &lt;strong&gt;why the return isn't ready&lt;/strong&gt; is much more useful.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GSTR-3B

Sales Data        Complete
Purchase Data     Complete
GSTR-2B           Complete
Reconciliation    3 Exceptions
CA Review         Pending
Client Approval   Pending
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the system can identify exactly what is blocking the filing.&lt;/p&gt;

&lt;p&gt;That turns a simple reminder system into a workflow system.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. GST Reconciliation and Exception Detection
&lt;/h2&gt;

&lt;p&gt;AI doesn't need to review everything.&lt;/p&gt;

&lt;p&gt;Sometimes its most valuable role is to find the things that deserve human attention.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Duplicate invoices&lt;/li&gt;
&lt;li&gt;GST amount mismatches&lt;/li&gt;
&lt;li&gt;Purchase invoices missing from 2B&lt;/li&gt;
&lt;li&gt;Unusual ITC movements&lt;/li&gt;
&lt;li&gt;Sales data mismatches&lt;/li&gt;
&lt;li&gt;Unexpected changes in margins&lt;/li&gt;
&lt;li&gt;Credit-note anomalies&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead of asking someone to manually inspect everything, AI can create an exception list.&lt;/p&gt;

&lt;p&gt;The CA reviews the exceptions.&lt;/p&gt;

&lt;p&gt;This is a powerful pattern for professional services:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;AI prepares and prioritizes. Humans review and approve.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  7. ITR Document and Information Collection
&lt;/h2&gt;

&lt;p&gt;ITR preparation involves another predictable collection process.&lt;/p&gt;

&lt;p&gt;The firm may need Form 16, capital gains information, investment proofs, bank statements, interest certificates, previous returns and other supporting information.&lt;/p&gt;

&lt;p&gt;The workflow can become:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ITR Client
↓
Checklist Generated
↓
Documents Requested
↓
Documents Received
↓
AI Classification
↓
Missing Information Identified
↓
Staff Review
↓
CA Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The system isn't preparing tax advice on its own.&lt;/p&gt;

&lt;p&gt;It's reducing the administrative work required to get the case ready.&lt;/p&gt;

&lt;h2&gt;
  
  
  8. Compliance Deadline Management
&lt;/h2&gt;

&lt;p&gt;A calendar tells you when something is due.&lt;/p&gt;

&lt;p&gt;A workflow tells you whether you're actually ready.&lt;/p&gt;

&lt;p&gt;A better compliance system can automatically create tasks based on the client's services.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
↓
GST
↓
TDS
↓
ITR
↓
ROC
↓
Audit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each service creates its own workflow.&lt;/p&gt;

&lt;p&gt;Tasks can be assigned to staff.&lt;/p&gt;

&lt;p&gt;Missing information can trigger client reminders.&lt;/p&gt;

&lt;p&gt;Unresolved items can be escalated.&lt;/p&gt;

&lt;p&gt;The partner gets visibility without having to ask every employee for an update.&lt;/p&gt;

&lt;h2&gt;
  
  
  9. Internal Task Assignment
&lt;/h2&gt;

&lt;p&gt;One of the biggest problems in a growing CA firm is simply knowing:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who is doing what?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A compliance workflow could automatically create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client: ABC Industries

Task: GST Reconciliation
Owner: Rahul
Due: 25 August
Status: Waiting for Client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the documents arrive:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Status: Ready for Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the CA approves:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Status: Completed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates a clear operational trail.&lt;/p&gt;

&lt;p&gt;It also makes workload visible across the team.&lt;/p&gt;

&lt;h2&gt;
  
  
  10. Fee and Payment Follow-ups
&lt;/h2&gt;

&lt;p&gt;Another repetitive process is collecting fees.&lt;/p&gt;

&lt;p&gt;The basic workflow is simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invoice
↓
Due Date
↓
Reminder
↓
Client Response
↓
Payment
↓
Receipt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But automation can make this smarter.&lt;/p&gt;

&lt;p&gt;If a client says:&lt;/p&gt;

&lt;p&gt;“I'll pay on Friday.”&lt;/p&gt;

&lt;p&gt;The system can record the commitment.&lt;/p&gt;

&lt;p&gt;If Friday passes and payment hasn't arrived, the follow-up can happen automatically.&lt;/p&gt;

&lt;p&gt;Staff only needs to intervene when the normal workflow doesn't work.&lt;/p&gt;

&lt;h2&gt;
  
  
  11. Loan Lead Qualification
&lt;/h2&gt;

&lt;p&gt;This is where things get particularly interesting.&lt;/p&gt;

&lt;p&gt;Many CA firms also have finance or loan departments.&lt;/p&gt;

&lt;p&gt;They may help clients with business loans, working capital, project finance, machinery finance, home loans, vehicle loans and other financing requirements.&lt;/p&gt;

&lt;p&gt;A client might simply send:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Mare ₹2 crore ni working capital joiye che.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Instead of immediately handing the conversation to someone, an AI assistant can collect the basic information required for an initial assessment.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Business Type
Annual Turnover
Existing Loans
Current Bank
GST Status
ITR History
Loan Requirement
Purpose
Collateral
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The result becomes a structured finance lead instead of an unstructured WhatsApp conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  12. Loan Document Collection
&lt;/h2&gt;

&lt;p&gt;Loan processing involves a significant amount of documentation.&lt;/p&gt;

&lt;p&gt;Depending on the case, the team may need financial statements, ITRs, GST information, bank statements, existing loan details, KYC, property documents and other lender-specific information.&lt;/p&gt;

&lt;p&gt;A Loan Desk can show:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Loan Case #1047

Required: 24
Received: 18
Verified: 14
Missing: 6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The system can automatically follow up with the client.&lt;/p&gt;

&lt;p&gt;It can also identify which cases are blocked because of missing information.&lt;/p&gt;

&lt;p&gt;That makes the finance team's pipeline much easier to manage.&lt;/p&gt;

&lt;h2&gt;
  
  
  13. Bank Query Tracking
&lt;/h2&gt;

&lt;p&gt;This is one of the most interesting workflows in loan processing.&lt;/p&gt;

&lt;p&gt;A bank might ask:&lt;/p&gt;

&lt;p&gt;“Please explain the increase in debtor days.”&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;p&gt;“Please provide updated CMA.”&lt;/p&gt;

&lt;p&gt;Then:&lt;/p&gt;

&lt;p&gt;“Please submit promoter ITR.”&lt;/p&gt;

&lt;p&gt;These requests can arrive through email, WhatsApp or phone calls.&lt;/p&gt;

&lt;p&gt;A structured Loan Desk could turn each request into a task.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Loan Case #1047

Bank Query:
Explain increase in debtor days

Owner:
Rahul

Status:
Waiting for Client

Due:
Tomorrow

Required:
Debtor Ageing Report
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the team knows exactly what is pending and who owns it.&lt;/p&gt;

&lt;h2&gt;
  
  
  14. AI-Assisted CMA and Financial Analysis
&lt;/h2&gt;

&lt;p&gt;For CA firms involved in project finance and business loans, financial analysis can become another area for automation.&lt;/p&gt;

&lt;p&gt;A system could take historical financial information and help prepare a first-pass analysis.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Historical Financials
↓
P&amp;amp;L
Balance Sheet
Cash Flow
↓
Financial Ratios
↓
Projected Financials
↓
CMA / Project Analysis
↓
CA Review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AI could also help identify unusual changes or prepare explanations for financial movements.&lt;/p&gt;

&lt;p&gt;But the final assumptions, analysis and professional output should remain under human review.&lt;/p&gt;

&lt;p&gt;AI should accelerate the preparation.&lt;/p&gt;

&lt;p&gt;It shouldn't blindly make the decision.&lt;/p&gt;

&lt;h2&gt;
  
  
  15. An AI Daily Brief for the CA Partner
&lt;/h2&gt;

&lt;p&gt;Eventually, all of these workflows can come together.&lt;/p&gt;

&lt;p&gt;But I don't think the answer is another dashboard filled with charts.&lt;/p&gt;

&lt;p&gt;A CA partner doesn't necessarily need more information.&lt;/p&gt;

&lt;p&gt;They need to know:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What needs my attention today?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The system could provide a simple daily summary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;7 filings blocked
5 clients haven't submitted documents
3 bank queries pending
₹12.8L overdue
8 loan cases waiting for action
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's more useful than opening five different spreadsheets and asking five different people for updates.&lt;/p&gt;

&lt;p&gt;The goal isn't more information.&lt;/p&gt;

&lt;p&gt;It's &lt;strong&gt;less cognitive load&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Don't Build Everything at Once
&lt;/h2&gt;

&lt;p&gt;This is probably the most important point.&lt;/p&gt;

&lt;p&gt;It is tempting to build a complete CA platform with:&lt;/p&gt;

&lt;p&gt;CRM.&lt;/p&gt;

&lt;p&gt;WhatsApp.&lt;/p&gt;

&lt;p&gt;AI chatbot.&lt;/p&gt;

&lt;p&gt;GST.&lt;/p&gt;

&lt;p&gt;ITR.&lt;/p&gt;

&lt;p&gt;Documents.&lt;/p&gt;

&lt;p&gt;Loans.&lt;/p&gt;

&lt;p&gt;CMA.&lt;/p&gt;

&lt;p&gt;Payments.&lt;/p&gt;

&lt;p&gt;Analytics.&lt;/p&gt;

&lt;p&gt;That can quickly turn into a six-month or one-year project before anyone knows whether the product solves a real problem.&lt;/p&gt;

&lt;p&gt;A better approach is to pick one workflow.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Document Collection
↓
Document Verification
↓
Compliance Workflow
↓
GST Exceptions
↓
Loan Desk
↓
Financial Intelligence
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each module can solve a real problem independently.&lt;/p&gt;

&lt;p&gt;Then they can gradually connect.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Should Actually Be Automated?
&lt;/h2&gt;

&lt;p&gt;A simple rule can help.&lt;/p&gt;

&lt;p&gt;Automate work that is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Repetitive&lt;/li&gt;
&lt;li&gt;High-volume&lt;/li&gt;
&lt;li&gt;Time-consuming&lt;/li&gt;
&lt;li&gt;Rule-based&lt;/li&gt;
&lt;li&gt;Easy to verify&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep humans in control of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Professional judgment&lt;/li&gt;
&lt;li&gt;Tax positions&lt;/li&gt;
&lt;li&gt;Financial decisions&lt;/li&gt;
&lt;li&gt;Client advice&lt;/li&gt;
&lt;li&gt;Final approvals&lt;/li&gt;
&lt;li&gt;Lending decisions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The best AI system isn't the one that makes every decision.&lt;/p&gt;

&lt;p&gt;It's the one that handles the repetitive work and brings the important exceptions to the professional.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Bigger Opportunity
&lt;/h2&gt;

&lt;p&gt;The future of CA-firm software probably isn't another chatbot.&lt;/p&gt;

&lt;p&gt;It's a system that understands the state of every client and every workflow.&lt;/p&gt;

&lt;p&gt;Which documents are missing?&lt;/p&gt;

&lt;p&gt;Which filings are blocked?&lt;/p&gt;

&lt;p&gt;Which tasks are overdue?&lt;/p&gt;

&lt;p&gt;Which loan cases are waiting for the bank?&lt;/p&gt;

&lt;p&gt;Which clients haven't paid?&lt;/p&gt;

&lt;p&gt;Which cases need the CA's attention?&lt;/p&gt;

&lt;p&gt;And which tasks can simply be handled automatically?&lt;/p&gt;

&lt;p&gt;The CA remains the expert.&lt;/p&gt;

&lt;p&gt;The team remains responsible for professional execution.&lt;/p&gt;

&lt;p&gt;AI becomes the operational layer that connects the work.&lt;/p&gt;

&lt;p&gt;That's where I think the opportunity becomes much more interesting.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;If you run a CA firm, what is the one repetitive process your team does every day that you would happily never have to manage manually again?&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>productivity</category>
    </item>
    <item>
      <title>NOOA: What If an AI Agent Was Just a Python Object?</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Wed, 12 Aug 2026 17:32:29 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/nooa-what-if-an-ai-agent-was-just-a-python-object-453o</link>
      <guid>https://dev.to/gaurav_talesara/nooa-what-if-an-ai-agent-was-just-a-python-object-453o</guid>
      <description>&lt;p&gt;There is something interesting happening in the AI agent space.&lt;/p&gt;

&lt;p&gt;We have spent the last couple of years building agents using prompts, tools, function calling, workflows, graphs, memory systems, orchestration layers, and increasingly complicated frameworks.&lt;/p&gt;

&lt;p&gt;And now NVIDIA Labs has released something that made me stop and think:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if an AI agent was just a Python object?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the basic idea behind &lt;strong&gt;NOOA — NVIDIA Object-Oriented Agents&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;It is an open-source, model-agnostic Python framework that represents an agent as a Python class, where the object's fields represent state, methods represent capabilities, docstrings can define instructions, and type annotations define interfaces.&lt;/p&gt;

&lt;p&gt;The project is still very new and explicitly described as research software. But I think the idea behind it is worth paying attention to.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/NVIDIA-NeMo/labs-OO-Agents?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;NOOA on GitHub&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  So, what is actually different?
&lt;/h2&gt;

&lt;p&gt;Let's look at the traditional way we might build an agent.&lt;/p&gt;

&lt;p&gt;We could have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A system prompt&lt;/li&gt;
&lt;li&gt;A collection of tools&lt;/li&gt;
&lt;li&gt;Tool schemas&lt;/li&gt;
&lt;li&gt;A memory component&lt;/li&gt;
&lt;li&gt;An orchestration loop&lt;/li&gt;
&lt;li&gt;State management&lt;/li&gt;
&lt;li&gt;Some workflow engine&lt;/li&gt;
&lt;li&gt;Observability/tracing&lt;/li&gt;
&lt;li&gt;Retry logic&lt;/li&gt;
&lt;li&gt;Structured output handling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;All of these pieces are useful.&lt;/p&gt;

&lt;p&gt;But they also create a growing abstraction layer between the developer and the agent.&lt;/p&gt;

&lt;p&gt;NOOA takes a different approach.&lt;/p&gt;

&lt;p&gt;You define something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;nooa&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;Agent&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SupportAgent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;Agent&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;You are a customer support agent.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="n"&gt;order_db&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;OrderDB&lt;/span&gt;

    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;is_refund_eligible&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Order&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;bool&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;delivered&lt;/span&gt;
            &lt;span class="ow"&gt;and&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;days_since_delivery&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="mi"&gt;30&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;triage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Order&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;gt;&lt;/span&gt; &lt;span class="n"&gt;Ticket&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Create a typed support ticket.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;
        &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And suddenly the architecture feels very familiar.&lt;/p&gt;

&lt;p&gt;The agent is an object.&lt;/p&gt;

&lt;p&gt;Its state is on the object.&lt;/p&gt;

&lt;p&gt;Its capabilities are methods.&lt;/p&gt;

&lt;p&gt;Its interfaces are typed.&lt;/p&gt;

&lt;p&gt;Its instructions can live with the methods.&lt;/p&gt;

&lt;p&gt;And the interesting part is the &lt;code&gt;...&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A method with a real Python implementation behaves like normal deterministic Python.&lt;/p&gt;

&lt;p&gt;A method with an &lt;code&gt;...&lt;/code&gt; body becomes an LLM-driven method at runtime.&lt;/p&gt;

&lt;p&gt;That is a surprisingly simple idea.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does this matter?
&lt;/h2&gt;

&lt;p&gt;The thing that caught my attention isn't just the syntax.&lt;/p&gt;

&lt;p&gt;It is the &lt;strong&gt;mental model&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;We have traditionally thought about an AI agent as:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Prompt + Model + Tools + Memory + Loop&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;NOOA is suggesting another abstraction:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agent = Software Object + Model&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That is a meaningful shift.&lt;/p&gt;

&lt;p&gt;If the agent is a Python object, then many things software engineers already know how to do become natural again.&lt;/p&gt;

&lt;p&gt;Testing.&lt;/p&gt;

&lt;p&gt;Refactoring.&lt;/p&gt;

&lt;p&gt;Version control.&lt;/p&gt;

&lt;p&gt;Tracing.&lt;/p&gt;

&lt;p&gt;Dependency injection.&lt;/p&gt;

&lt;p&gt;Type checking.&lt;/p&gt;

&lt;p&gt;Composition.&lt;/p&gt;

&lt;p&gt;State management.&lt;/p&gt;

&lt;p&gt;Code review.&lt;/p&gt;

&lt;p&gt;Instead of learning another workflow DSL or another orchestration abstraction, developers can start with something they already understand: Python.&lt;/p&gt;

&lt;p&gt;NVIDIA's implementation goes further than simply wrapping an LLM in a class. The framework supports typed I/O, live Python objects passed by reference, model-generated Python as an action mechanism, programmable agent loops, context/event APIs, tracing, and long-term memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part I find especially interesting: code as action
&lt;/h2&gt;

&lt;p&gt;This is probably one of the most interesting pieces of NOOA.&lt;/p&gt;

&lt;p&gt;Instead of every capability necessarily becoming a traditional function/tool schema that gets serialized into the model context, the model can generate Python and operate within the agent's environment.&lt;/p&gt;

&lt;p&gt;In other words:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The model doesn't only call tools.&lt;br&gt;
The model can write code to use the agent's capabilities.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That changes the interface between the LLM and the application.&lt;/p&gt;

&lt;p&gt;Python methods and type annotations can become the interface the model works with.&lt;/p&gt;

&lt;p&gt;For developers, this could eventually mean less time maintaining huge collections of tool definitions and more time defining clean software interfaces.&lt;/p&gt;

&lt;p&gt;Of course, this also creates a very important security problem.&lt;/p&gt;

&lt;p&gt;If an LLM can generate and execute Python, we should treat that code as untrusted.&lt;/p&gt;

&lt;p&gt;NOOA's own documentation is very clear about this: its AST validation and module restrictions are defense-in-depth mechanisms, not a security boundary. The recommended containment boundary is OS-level sandboxing such as containers or NVIDIA OpenShell.&lt;/p&gt;

&lt;p&gt;And I think that distinction is extremely important.&lt;/p&gt;
&lt;h2&gt;
  
  
  Another interesting idea: agent state
&lt;/h2&gt;

&lt;p&gt;There is another architectural question that NOOA makes interesting.&lt;/p&gt;

&lt;p&gt;Where should an agent's state live?&lt;/p&gt;

&lt;p&gt;Today, a lot of agent state effectively lives inside the context window.&lt;/p&gt;

&lt;p&gt;The longer the conversation becomes, the more we start thinking about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;summarization&lt;/li&gt;
&lt;li&gt;context compression&lt;/li&gt;
&lt;li&gt;retrieval&lt;/li&gt;
&lt;li&gt;memory&lt;/li&gt;
&lt;li&gt;token optimization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;NOOA instead treats the agent as an object with state.&lt;/p&gt;

&lt;p&gt;That opens the door to a different model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
 ├── State
 ├── Capabilities
 ├── Memory
 ├── Methods
 ├── Context
 └── Model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The LLM becomes part of the object rather than the object being constructed around an LLM conversation.&lt;/p&gt;

&lt;p&gt;That is subtle, but I think it could become important as agents become more persistent and autonomous.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is NOOA going to replace other agent frameworks?
&lt;/h2&gt;

&lt;p&gt;I don't think we know that yet.&lt;/p&gt;

&lt;p&gt;And I would actually argue that this is the wrong question.&lt;/p&gt;

&lt;p&gt;NOOA is currently a &lt;strong&gt;0.x research preview&lt;/strong&gt;, and NVIDIA says its public API is not yet stable and can change between releases.&lt;/p&gt;

&lt;p&gt;So I wouldn't recommend looking at it today and saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"This is the new standard for AI agents."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It isn't.&lt;/p&gt;

&lt;p&gt;But I would definitely recommend watching it.&lt;/p&gt;

&lt;p&gt;Because research projects like this sometimes introduce an abstraction that looks unusual initially and becomes obvious later.&lt;/p&gt;

&lt;p&gt;Remember how strange some ideas looked before they became standard programming patterns?&lt;/p&gt;

&lt;p&gt;The interesting question here is whether &lt;strong&gt;object-oriented programming becomes a useful abstraction for agent engineering&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  My prediction
&lt;/h2&gt;

&lt;p&gt;I don't think the future will be one giant agent framework.&lt;/p&gt;

&lt;p&gt;I think we are going to see a convergence of several ideas.&lt;/p&gt;

&lt;p&gt;Agents will increasingly look like software components rather than chatbot conversations.&lt;/p&gt;

&lt;p&gt;They will have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Persistent state&lt;/li&gt;
&lt;li&gt;Typed interfaces&lt;/li&gt;
&lt;li&gt;Memory&lt;/li&gt;
&lt;li&gt;Capabilities&lt;/li&gt;
&lt;li&gt;Deterministic code&lt;/li&gt;
&lt;li&gt;Model-driven reasoning&lt;/li&gt;
&lt;li&gt;Observability&lt;/li&gt;
&lt;li&gt;Sandboxed execution&lt;/li&gt;
&lt;li&gt;Testable behavior&lt;/li&gt;
&lt;li&gt;Composable sub-agents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And the boundary between &lt;strong&gt;"AI logic"&lt;/strong&gt; and &lt;strong&gt;"software logic"&lt;/strong&gt; will become much thinner.&lt;/p&gt;

&lt;p&gt;That's what makes NOOA interesting to me.&lt;/p&gt;

&lt;p&gt;It isn't necessarily introducing another way to call an LLM.&lt;/p&gt;

&lt;p&gt;It is asking a more fundamental question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What should an AI agent look like from a software engineer's perspective?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And the answer from NVIDIA Labs is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Maybe it should just look like a Python object.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What I want to see next
&lt;/h2&gt;

&lt;p&gt;This is where things get really interesting.&lt;/p&gt;

&lt;p&gt;I'd love to see how this approach evolves around:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Production reliability&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;How do these agents behave under real workloads, failures, retries, concurrency, and partial state?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Security&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;If agents can generate and execute code, sandboxing will become a fundamental part of the architecture, not an optional feature.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Multi-agent systems&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What happens when Python objects representing agents start collaborating with each other?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;ResearchAgent
       ↓
PlanningAgent
       ↓
CodingAgent
       ↓
ReviewAgent
       ↓
DeploymentAgent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Can these simply become composable software objects?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Agent testing&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;This could be particularly powerful.&lt;/p&gt;

&lt;p&gt;Imagine being able to test an agent almost like any other Python component:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;test_refund_policy&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;
    &lt;span class="n"&gt;agent&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;SupportAgent&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
    &lt;span class="k"&gt;assert&lt;/span&gt; &lt;span class="n"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;is_refund_eligible&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;is&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And then separately evaluate the probabilistic behavior of its LLM-driven methods.&lt;/p&gt;

&lt;p&gt;That separation between deterministic logic and model-driven behavior could become a very useful engineering pattern.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. Agent observability&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;NOOA already traces LLM calls, code execution, and method invocations.&lt;/p&gt;

&lt;p&gt;I think this will become critical.&lt;/p&gt;

&lt;p&gt;As agents become more autonomous, "the model said something weird" won't be enough for debugging.&lt;/p&gt;

&lt;p&gt;We will need to know:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What did the agent know?
What state did it have?
What method did it invoke?
What code did it generate?
What did that code change?
What model decision happened?
What happened next?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is software observability applied to probabilistic systems.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bigger shift
&lt;/h2&gt;

&lt;p&gt;Maybe the most interesting thing about NOOA isn't NOOA itself.&lt;/p&gt;

&lt;p&gt;Maybe it is the direction it represents.&lt;/p&gt;

&lt;p&gt;The first generation of AI applications taught us how to integrate models into software.&lt;/p&gt;

&lt;p&gt;The next generation is teaching us how to make models behave like software components.&lt;/p&gt;

&lt;p&gt;And eventually, I think we will stop drawing such a sharp line between the two.&lt;/p&gt;

&lt;p&gt;An agent won't just be:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"an LLM with some tools."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It may become:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"a software component whose reasoning happens to be powered by a model."&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;NOOA is still early.&lt;/p&gt;

&lt;p&gt;The APIs will change.&lt;/p&gt;

&lt;p&gt;The patterns will evolve.&lt;/p&gt;

&lt;p&gt;There will be competing approaches.&lt;/p&gt;

&lt;p&gt;Some ideas will work. Some won't.&lt;/p&gt;

&lt;p&gt;But that's exactly why I think it is worth experimenting with now.&lt;/p&gt;

&lt;p&gt;Not because NOOA is already the answer.&lt;/p&gt;

&lt;p&gt;But because it might be pointing toward a different question.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What does software engineering look like when the software itself can reason?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That is the part I'm watching.&lt;/p&gt;

&lt;p&gt;And I have a feeling we are going to see some very interesting updates in this space over the next year.&lt;/p&gt;




&lt;p&gt;If you're building AI agents today, I'm curious:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Would you rather build your next agent as a workflow/graph — or as a Python object?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I'd love to hear what other engineers think.&lt;/p&gt;

&lt;h1&gt;
  
  
  AI #AIAgents #AgenticAI #NVIDIA #Python #LLM #GenerativeAI #SoftwareEngineering #ArtificialIntelligence #DeveloperExperience
&lt;/h1&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>nvidia</category>
      <category>agentskills</category>
    </item>
    <item>
      <title>Building an AI Workforce for Insurance with n8n, OpenAI, LangGraph and Supabase</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Tue, 16 Jun 2026 18:59:15 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/building-an-ai-workforce-for-insurance-with-n8n-openai-langgraph-and-supabase-4cj1</link>
      <guid>https://dev.to/gaurav_talesara/building-an-ai-workforce-for-insurance-with-n8n-openai-langgraph-and-supabase-4cj1</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;AI for Preparation. Humans for Judgment.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Most AI projects today are one of these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A chatbot&lt;/li&gt;
&lt;li&gt;A customer support bot&lt;/li&gt;
&lt;li&gt;A voice assistant&lt;/li&gt;
&lt;li&gt;A Q&amp;amp;A system&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But I wanted to explore something bigger:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What if businesses could build an AI Workforce?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Instead of one AI assistant,&lt;/p&gt;

&lt;p&gt;imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer

↓

AI Workforce

├── Discovery Agent

├── Research Agent

├── Policy Comparison Agent

├── Recommendation Agent

├── CRM Agent

└── Follow-up Agent

↓

Human Advisor

↓

Customer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This article explains the architecture and design decisions behind such a system.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Insurance?
&lt;/h2&gt;

&lt;p&gt;Insurance is an interesting industry for AI.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Research is repetitive.&lt;/li&gt;
&lt;li&gt;Recommendations are data-driven.&lt;/li&gt;
&lt;li&gt;Follow-ups are expensive.&lt;/li&gt;
&lt;li&gt;Trust is critical.&lt;/li&gt;
&lt;li&gt;Human judgment is still necessary.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This makes Insurance a perfect &lt;strong&gt;Human-in-the-Loop AI&lt;/strong&gt; use case.&lt;/p&gt;




&lt;h2&gt;
  
  
  Human In The Loop
&lt;/h2&gt;

&lt;p&gt;This is the core philosophy.&lt;/p&gt;

&lt;p&gt;I don't want AI to automatically sell insurance.&lt;/p&gt;

&lt;p&gt;I don't want AI replacing advisors.&lt;/p&gt;

&lt;p&gt;I want:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;AI prepares.

Humans decide.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The workflow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer

↓

AI Workforce

↓

Human Advisor Review

↓

Customer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This creates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Faster recommendations&lt;/li&gt;
&lt;li&gt;Better customer experience&lt;/li&gt;
&lt;li&gt;Safer AI adoption&lt;/li&gt;
&lt;li&gt;Human accountability&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  AI Workforce Architecture
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Customer

↓

WhatsApp
Phone Call
Website Chat
Email

↓

AI Workforce

├── Discovery Agent

├── Research Agent

├── Comparison Agent

├── Recommendation Agent

├── CRM Agent

└── Follow-up Agent

↓

Human Advisor

↓

Customer
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Discovery Agent
&lt;/h2&gt;

&lt;p&gt;The Discovery Agent understands the customer.&lt;/p&gt;

&lt;p&gt;Responsibilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Collect customer profile&lt;/li&gt;
&lt;li&gt;Understand goals&lt;/li&gt;
&lt;li&gt;Assess risk&lt;/li&gt;
&lt;li&gt;Understand existing insurance&lt;/li&gt;
&lt;li&gt;Identify gaps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example Output:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"risk_level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"family_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"married_with_children"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"insurance_goal"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"health_and_term"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recommended_health_cover"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"20L"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"recommended_term_cover"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"3Cr"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Research Agent
&lt;/h2&gt;

&lt;p&gt;The Research Agent acts like an insurance analyst.&lt;/p&gt;

&lt;p&gt;Responsibilities:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Analyze policies&lt;/li&gt;
&lt;li&gt;Compare waiting periods&lt;/li&gt;
&lt;li&gt;Review exclusions&lt;/li&gt;
&lt;li&gt;Evaluate premiums&lt;/li&gt;
&lt;li&gt;Generate recommendations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"customer_profile_summary"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"top_recommendations"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"risks"&lt;/span&gt;&lt;span class="p"&gt;:[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="s2"&gt;"..."&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"confidence_score"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="mf"&gt;0.92&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  Comparison Agent
&lt;/h2&gt;

&lt;p&gt;Creates structured comparisons:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Plan A&lt;/th&gt;
&lt;th&gt;Plan B&lt;/th&gt;
&lt;th&gt;Plan C&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Coverage&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Premium&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Waiting Period&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claim Process&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;td&gt;✓&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Output:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Best Overall&lt;/li&gt;
&lt;li&gt;Best Budget&lt;/li&gt;
&lt;li&gt;Best Family Plan&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Recommendation Agent
&lt;/h2&gt;

&lt;p&gt;Creates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Customer Summary&lt;/li&gt;
&lt;li&gt;Recommended Plan&lt;/li&gt;
&lt;li&gt;Alternatives&lt;/li&gt;
&lt;li&gt;Risk Analysis&lt;/li&gt;
&lt;li&gt;Advisor Notes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Everything before the advisor joins.&lt;/p&gt;




&lt;h2&gt;
  
  
  CRM Agent
&lt;/h2&gt;

&lt;p&gt;Updates:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Customer Records&lt;/li&gt;
&lt;li&gt;Recommendations&lt;/li&gt;
&lt;li&gt;Activities&lt;/li&gt;
&lt;li&gt;Opportunity Status&lt;/li&gt;
&lt;li&gt;Tasks&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Follow-up Agent
&lt;/h2&gt;

&lt;p&gt;Handles:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;WhatsApp reminders&lt;/li&gt;
&lt;li&gt;Renewal alerts&lt;/li&gt;
&lt;li&gt;Email follow-ups&lt;/li&gt;
&lt;li&gt;Call notes&lt;/li&gt;
&lt;li&gt;Engagement tracking&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Omnichannel AI
&lt;/h2&gt;

&lt;p&gt;One important decision:&lt;/p&gt;

&lt;p&gt;Customers should not install a new application.&lt;/p&gt;

&lt;p&gt;The AI Workforce should operate through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;WhatsApp&lt;/li&gt;
&lt;li&gt;Phone Calls&lt;/li&gt;
&lt;li&gt;Website Chat&lt;/li&gt;
&lt;li&gt;Email&lt;/li&gt;
&lt;li&gt;SMS&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Different channels.&lt;/p&gt;

&lt;p&gt;Same intelligence.&lt;/p&gt;




&lt;h2&gt;
  
  
  Technology Stack
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Frontend
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Next.js&lt;/li&gt;
&lt;li&gt;Tailwind&lt;/li&gt;
&lt;li&gt;Lovable AI&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Workflow Layer
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;n8n&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  AI Models
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;OpenAI&lt;/li&gt;
&lt;li&gt;Gemini&lt;/li&gt;
&lt;li&gt;Claude&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Multi-Agent Framework
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;LangGraph&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Database
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Supabase&lt;/li&gt;
&lt;li&gt;PostgreSQL&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Memory
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Pinecone&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Monitoring
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;LangSmith&lt;/li&gt;
&lt;li&gt;PostHog&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why n8n First?
&lt;/h2&gt;

&lt;p&gt;I intentionally started with n8n.&lt;/p&gt;

&lt;p&gt;Because:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fast prototyping&lt;/li&gt;
&lt;li&gt;Visual workflows&lt;/li&gt;
&lt;li&gt;Easy OpenAI integration&lt;/li&gt;
&lt;li&gt;Easy Supabase integration&lt;/li&gt;
&lt;li&gt;Easy WhatsApp integration&lt;/li&gt;
&lt;li&gt;Easy Email workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;After validation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;n8n

↓

NestJS

↓

LangGraph

↓

Production AI Workforce
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  The Bigger Vision
&lt;/h2&gt;

&lt;p&gt;I don't think AI will replace Insurance Advisors.&lt;/p&gt;

&lt;p&gt;I think every advisor may eventually have:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An AI Workforce working behind the scenes.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;AI provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Speed&lt;/li&gt;
&lt;li&gt;Consistency&lt;/li&gt;
&lt;li&gt;Scale&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Humans provide:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Trust&lt;/li&gt;
&lt;li&gt;Empathy&lt;/li&gt;
&lt;li&gt;Judgment&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The future is not:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human vs AI&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The future is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Human + AI Workforce&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;If you're building something similar, I'd love to hear your thoughts.&lt;/p&gt;

</description>
      <category>ai</category>
    </item>
    <item>
      <title>The Rise of Production-Grade AI Infrastructure</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Sat, 23 May 2026 07:01:21 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/the-rise-of-production-grade-ai-infrastructure-3h11</link>
      <guid>https://dev.to/gaurav_talesara/the-rise-of-production-grade-ai-infrastructure-3h11</guid>
      <description>&lt;p&gt;Most AI products today are impressive in demos.&lt;/p&gt;

&lt;p&gt;But the moment they hit production:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;workflows break&lt;/li&gt;
&lt;li&gt;context fails&lt;/li&gt;
&lt;li&gt;hallucinations appear&lt;/li&gt;
&lt;li&gt;costs explode&lt;/li&gt;
&lt;li&gt;observability disappears&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The AI industry does not really have an “intelligence” problem anymore.&lt;/p&gt;

&lt;p&gt;It has an infrastructure problem.&lt;/p&gt;

&lt;p&gt;For the last two years, the ecosystem focused heavily on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;chat interfaces&lt;/li&gt;
&lt;li&gt;prompt engineering&lt;/li&gt;
&lt;li&gt;copilots&lt;/li&gt;
&lt;li&gt;wrappers around foundation models&lt;/li&gt;
&lt;li&gt;“AI-powered” product features&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That phase accelerated adoption.&lt;/p&gt;

&lt;p&gt;But the market is now entering a different stage.&lt;/p&gt;

&lt;p&gt;The hard problem is no longer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Can AI generate something useful?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The hard problem is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Can AI systems operate reliably in real production environments?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And that is where the next major opportunity is emerging.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Demo Problem
&lt;/h2&gt;

&lt;p&gt;Most AI demos look incredible.&lt;/p&gt;

&lt;p&gt;They can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;generate code&lt;/li&gt;
&lt;li&gt;summarize documents&lt;/li&gt;
&lt;li&gt;automate workflows&lt;/li&gt;
&lt;li&gt;answer questions&lt;/li&gt;
&lt;li&gt;orchestrate tasks&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But production environments expose a completely different reality.&lt;/p&gt;

&lt;p&gt;Once real users, real workflows, and real operational constraints enter the system, problems begin to appear:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;hallucinations&lt;/li&gt;
&lt;li&gt;fragile context handling&lt;/li&gt;
&lt;li&gt;inconsistent outputs&lt;/li&gt;
&lt;li&gt;broken execution chains&lt;/li&gt;
&lt;li&gt;runaway costs&lt;/li&gt;
&lt;li&gt;poor observability&lt;/li&gt;
&lt;li&gt;unsafe automation&lt;/li&gt;
&lt;li&gt;missing governance&lt;/li&gt;
&lt;li&gt;unpredictable agent behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is why so many AI pilots never move beyond experimentation.&lt;/p&gt;

&lt;p&gt;The market today is filled with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI interfaces&lt;/li&gt;
&lt;li&gt;AI assistants&lt;/li&gt;
&lt;li&gt;AI wrappers&lt;/li&gt;
&lt;li&gt;AI copilots&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But what enterprises actually need are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reliable systems&lt;/li&gt;
&lt;li&gt;operational controls&lt;/li&gt;
&lt;li&gt;execution runtimes&lt;/li&gt;
&lt;li&gt;observability layers&lt;/li&gt;
&lt;li&gt;governance infrastructure&lt;/li&gt;
&lt;li&gt;context orchestration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is the real bottleneck now.&lt;/p&gt;




&lt;h2&gt;
  
  
  AI Systems Need a New Production Stack
&lt;/h2&gt;

&lt;p&gt;Traditional software engineering was built around deterministic systems.&lt;/p&gt;

&lt;p&gt;AI systems are different.&lt;/p&gt;

&lt;p&gt;They are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;probabilistic&lt;/li&gt;
&lt;li&gt;context-sensitive&lt;/li&gt;
&lt;li&gt;state-fragile&lt;/li&gt;
&lt;li&gt;operationally unpredictable&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means traditional software patterns are no longer enough.&lt;/p&gt;

&lt;p&gt;AI requires an entirely new operational layer.&lt;/p&gt;

&lt;p&gt;This feels very similar to earlier infrastructure shifts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Kubernetes standardized container orchestration&lt;/li&gt;
&lt;li&gt;Datadog transformed observability&lt;/li&gt;
&lt;li&gt;Stripe simplified payment infrastructure&lt;/li&gt;
&lt;li&gt;Temporal improved workflow reliability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AI is now reaching a similar stage.&lt;/p&gt;

&lt;p&gt;The next generation of products will not just be AI applications.&lt;/p&gt;

&lt;p&gt;They will be:&lt;/p&gt;

&lt;h3&gt;
  
  
  AI production infrastructure platforms.
&lt;/h3&gt;




&lt;h2&gt;
  
  
  The Real Layers of a Production-Grade AI System
&lt;/h2&gt;

&lt;p&gt;Most discussions about AI still focus only on models.&lt;/p&gt;

&lt;p&gt;But production-grade AI systems require much more than a model.&lt;/p&gt;

&lt;p&gt;Below are the infrastructure layers that are becoming increasingly important.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. Context Engineering
&lt;/h2&gt;

&lt;p&gt;This is becoming one of the most critical areas in AI engineering.&lt;/p&gt;

&lt;p&gt;Most AI systems fail not because the model is weak, but because the context is poor.&lt;/p&gt;

&lt;p&gt;Production systems need to manage:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;historical memory&lt;/li&gt;
&lt;li&gt;workflow state&lt;/li&gt;
&lt;li&gt;user intent&lt;/li&gt;
&lt;li&gt;permissions&lt;/li&gt;
&lt;li&gt;business logic&lt;/li&gt;
&lt;li&gt;external data&lt;/li&gt;
&lt;li&gt;codebase understanding&lt;/li&gt;
&lt;li&gt;semantic relationships&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This goes far beyond basic RAG.&lt;/p&gt;

&lt;p&gt;The future belongs to systems that can dynamically assemble the right context at the right moment.&lt;/p&gt;

&lt;p&gt;Prompt engineering is becoming commoditized.&lt;/p&gt;

&lt;p&gt;Context engineering is becoming the moat.&lt;/p&gt;




&lt;h2&gt;
  
  
  2. Agent Execution Runtime
&lt;/h2&gt;

&lt;p&gt;Most AI agents today are unreliable because they lack execution infrastructure.&lt;/p&gt;

&lt;p&gt;A production runtime needs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;retries&lt;/li&gt;
&lt;li&gt;rollback support&lt;/li&gt;
&lt;li&gt;checkpoints&lt;/li&gt;
&lt;li&gt;workflow state tracking&lt;/li&gt;
&lt;li&gt;timeout handling&lt;/li&gt;
&lt;li&gt;safe execution paths&lt;/li&gt;
&lt;li&gt;human approval systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without this, AI workflows become fragile very quickly.&lt;/p&gt;

&lt;p&gt;The market does not just need agents.&lt;/p&gt;

&lt;p&gt;It needs:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;workflow infrastructure for AI systems.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. Observability for AI Systems
&lt;/h2&gt;

&lt;p&gt;Debugging traditional software is already difficult.&lt;/p&gt;

&lt;p&gt;Debugging AI systems is significantly harder.&lt;/p&gt;

&lt;p&gt;Production AI requires visibility into:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;prompts&lt;/li&gt;
&lt;li&gt;memory retrieval&lt;/li&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;reasoning chains&lt;/li&gt;
&lt;li&gt;execution paths&lt;/li&gt;
&lt;li&gt;token usage&lt;/li&gt;
&lt;li&gt;latency&lt;/li&gt;
&lt;li&gt;hallucination patterns&lt;/li&gt;
&lt;li&gt;workflow failures&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most current systems still operate like black boxes.&lt;/p&gt;

&lt;p&gt;This creates a massive opportunity for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI observability&lt;/li&gt;
&lt;li&gt;AgentOps&lt;/li&gt;
&lt;li&gt;runtime tracing&lt;/li&gt;
&lt;li&gt;execution replay&lt;/li&gt;
&lt;li&gt;quality monitoring&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The industry will likely see a:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“Datadog for AI systems”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;category emerge.&lt;/p&gt;




&lt;h2&gt;
  
  
  4. Governance and Safety
&lt;/h2&gt;

&lt;p&gt;As AI systems become more autonomous, governance becomes mandatory.&lt;/p&gt;

&lt;p&gt;Enterprises need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;approval workflows&lt;/li&gt;
&lt;li&gt;audit trails&lt;/li&gt;
&lt;li&gt;permission systems&lt;/li&gt;
&lt;li&gt;policy enforcement&lt;/li&gt;
&lt;li&gt;data isolation&lt;/li&gt;
&lt;li&gt;secure execution environments&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Without operational controls, companies will struggle to trust autonomous systems at scale.&lt;/p&gt;

&lt;p&gt;This becomes especially important in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;healthcare&lt;/li&gt;
&lt;li&gt;finance&lt;/li&gt;
&lt;li&gt;enterprise automation&lt;/li&gt;
&lt;li&gt;internal copilots&lt;/li&gt;
&lt;li&gt;operational workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Governance is no longer optional infrastructure.&lt;/p&gt;

&lt;p&gt;It is foundational infrastructure.&lt;/p&gt;




&lt;h2&gt;
  
  
  5. Evaluation and Reliability Testing
&lt;/h2&gt;

&lt;p&gt;One of the biggest problems in AI today is silent degradation.&lt;/p&gt;

&lt;p&gt;An AI workflow may work perfectly today and fail tomorrow because of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;model updates&lt;/li&gt;
&lt;li&gt;prompt changes&lt;/li&gt;
&lt;li&gt;retrieval drift&lt;/li&gt;
&lt;li&gt;API schema changes&lt;/li&gt;
&lt;li&gt;edge cases&lt;/li&gt;
&lt;li&gt;workflow changes&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That means AI systems need continuous evaluation.&lt;/p&gt;

&lt;p&gt;Production-grade AI requires:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;regression testing&lt;/li&gt;
&lt;li&gt;scenario simulation&lt;/li&gt;
&lt;li&gt;adversarial testing&lt;/li&gt;
&lt;li&gt;replay systems&lt;/li&gt;
&lt;li&gt;benchmark scoring&lt;/li&gt;
&lt;li&gt;workflow validation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This category is still massively underdeveloped.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Infrastructure Will Matter More Than Interfaces
&lt;/h2&gt;

&lt;p&gt;The first AI wave rewarded:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;interfaces&lt;/li&gt;
&lt;li&gt;demos&lt;/li&gt;
&lt;li&gt;speed&lt;/li&gt;
&lt;li&gt;accessibility&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next AI wave will reward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reliability&lt;/li&gt;
&lt;li&gt;orchestration&lt;/li&gt;
&lt;li&gt;observability&lt;/li&gt;
&lt;li&gt;governance&lt;/li&gt;
&lt;li&gt;scalability&lt;/li&gt;
&lt;li&gt;operational maturity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That changes where the real value gets created.&lt;/p&gt;

&lt;p&gt;The winning companies may not be the ones with the best chat interface.&lt;/p&gt;

&lt;p&gt;They may be the ones building:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context runtimes&lt;/li&gt;
&lt;li&gt;orchestration layers&lt;/li&gt;
&lt;li&gt;observability platforms&lt;/li&gt;
&lt;li&gt;execution infrastructure&lt;/li&gt;
&lt;li&gt;repo intelligence systems&lt;/li&gt;
&lt;li&gt;AI governance tooling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The real opportunity is shifting downward into the infrastructure layer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Repo Intelligence Might Become a Major Category
&lt;/h2&gt;

&lt;p&gt;One particularly interesting opportunity is repo intelligence.&lt;/p&gt;

&lt;p&gt;Current AI coding tools can generate code.&lt;/p&gt;

&lt;p&gt;But they often lack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;architectural understanding&lt;/li&gt;
&lt;li&gt;dependency awareness&lt;/li&gt;
&lt;li&gt;service relationships&lt;/li&gt;
&lt;li&gt;domain knowledge&lt;/li&gt;
&lt;li&gt;operational context&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That creates problems in large production codebases.&lt;/p&gt;

&lt;p&gt;A smarter system would:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;scan repositories&lt;/li&gt;
&lt;li&gt;understand architecture&lt;/li&gt;
&lt;li&gt;build dependency graphs&lt;/li&gt;
&lt;li&gt;map services&lt;/li&gt;
&lt;li&gt;infer business domains&lt;/li&gt;
&lt;li&gt;track workflows&lt;/li&gt;
&lt;li&gt;generate contextual intelligence for AI systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This could dramatically improve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI coding reliability&lt;/li&gt;
&lt;li&gt;automated refactoring&lt;/li&gt;
&lt;li&gt;debugging&lt;/li&gt;
&lt;li&gt;onboarding&lt;/li&gt;
&lt;li&gt;workflow automation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The future of AI-assisted engineering may depend heavily on systems that deeply understand software architecture.&lt;/p&gt;




&lt;h2&gt;
  
  
  What This Means for Builders
&lt;/h2&gt;

&lt;p&gt;If you are building in AI today, this shift matters.&lt;/p&gt;

&lt;p&gt;The market is getting saturated with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;wrappers&lt;/li&gt;
&lt;li&gt;chat interfaces&lt;/li&gt;
&lt;li&gt;generic copilots&lt;/li&gt;
&lt;li&gt;shallow automation tools&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But infrastructure gaps are still massively underbuilt.&lt;/p&gt;

&lt;p&gt;That means opportunities are emerging in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;context orchestration&lt;/li&gt;
&lt;li&gt;observability&lt;/li&gt;
&lt;li&gt;evaluation systems&lt;/li&gt;
&lt;li&gt;governance tooling&lt;/li&gt;
&lt;li&gt;repo intelligence&lt;/li&gt;
&lt;li&gt;workflow runtimes&lt;/li&gt;
&lt;li&gt;execution reliability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The next major AI products may come from engineering pain, not prompt creativity.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Market Is Moving from Apps to Systems
&lt;/h2&gt;

&lt;p&gt;This is the transition happening right now.&lt;/p&gt;

&lt;p&gt;We are moving from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI apps → AI infrastructure&lt;/li&gt;
&lt;li&gt;prompts → context systems&lt;/li&gt;
&lt;li&gt;copilots → execution runtimes&lt;/li&gt;
&lt;li&gt;experimentation → operational maturity&lt;/li&gt;
&lt;li&gt;wrappers → production platforms&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The companies that win in AI will likely be the ones that solve:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;reliability&lt;/li&gt;
&lt;li&gt;orchestration&lt;/li&gt;
&lt;li&gt;observability&lt;/li&gt;
&lt;li&gt;governance&lt;/li&gt;
&lt;li&gt;context management&lt;/li&gt;
&lt;li&gt;execution safety&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not just generation.&lt;/p&gt;

&lt;p&gt;The biggest AI companies of the next decade may not even look like AI companies.&lt;/p&gt;

&lt;p&gt;They may look like infrastructure companies.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Thoughts
&lt;/h2&gt;

&lt;p&gt;AI will absolutely transform software.&lt;/p&gt;

&lt;p&gt;But models alone are not enough.&lt;/p&gt;

&lt;p&gt;The next major challenge is building systems that AI can operate inside reliably.&lt;/p&gt;

&lt;p&gt;That means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;better infrastructure&lt;/li&gt;
&lt;li&gt;better orchestration&lt;/li&gt;
&lt;li&gt;better context systems&lt;/li&gt;
&lt;li&gt;better observability&lt;/li&gt;
&lt;li&gt;better governance&lt;/li&gt;
&lt;li&gt;better operational tooling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The future of AI does not belong only to model providers.&lt;/p&gt;

&lt;p&gt;It also belongs to the companies building the operational layer around those models.&lt;/p&gt;

&lt;p&gt;And that may become one of the biggest infrastructure opportunities of the next decade.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>infrastructure</category>
      <category>machinelearning</category>
      <category>systemdesign</category>
    </item>
    <item>
      <title>Why Software Isn’t Built for AI Agents</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Sat, 02 May 2026 12:52:03 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/why-software-isnt-built-for-ai-agents-3ik5</link>
      <guid>https://dev.to/gaurav_talesara/why-software-isnt-built-for-ai-agents-3ik5</guid>
      <description>&lt;p&gt;The next users of your software won’t be humans.&lt;br&gt;
They’ll be agents.&lt;/p&gt;

&lt;p&gt;And most software today is completely unprepared for that.&lt;/p&gt;

&lt;p&gt;Right now, AI agents are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Browsing websites&lt;/li&gt;
&lt;li&gt;Filling forms&lt;/li&gt;
&lt;li&gt;Clicking buttons&lt;/li&gt;
&lt;li&gt;Navigating dashboards&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That’s not scale. That’s a workaround.&lt;/p&gt;

&lt;p&gt;We’re forcing machines to behave like humans—because our systems were never designed for anything else.&lt;/p&gt;


&lt;h2&gt;
  
  
  The Core Problem
&lt;/h2&gt;

&lt;p&gt;Modern software is built around a simple assumption:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;A human will be sitting in front of a screen.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That assumption drives everything:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;UI-heavy workflows&lt;/li&gt;
&lt;li&gt;Step-by-step interactions&lt;/li&gt;
&lt;li&gt;Documentation meant to be read, not executed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But agents don’t need interfaces.&lt;br&gt;
They need &lt;strong&gt;interfaces they can reason about and execute against programmatically&lt;/strong&gt;.&lt;/p&gt;


&lt;h2&gt;
  
  
  Where Current Systems Break for Agents
&lt;/h2&gt;

&lt;p&gt;Let’s break this down from a systems perspective.&lt;/p&gt;
&lt;h3&gt;
  
  
  1. UI-First Architecture
&lt;/h3&gt;

&lt;p&gt;Most SaaS products expose functionality through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dashboards&lt;/li&gt;
&lt;li&gt;Forms&lt;/li&gt;
&lt;li&gt;Buttons&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents interacting with these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rely on scraping or automation layers&lt;/li&gt;
&lt;li&gt;Break when UI changes&lt;/li&gt;
&lt;li&gt;Lack reliability&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  2. Non-Deterministic Outputs
&lt;/h3&gt;

&lt;p&gt;Agents need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured responses&lt;/li&gt;
&lt;li&gt;Predictable schemas&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Instead, they get:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HTML pages&lt;/li&gt;
&lt;li&gt;Inconsistent API responses&lt;/li&gt;
&lt;li&gt;Unstructured data&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  3. Human-Centric Authentication
&lt;/h3&gt;

&lt;p&gt;Current flows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;OAuth screens&lt;/li&gt;
&lt;li&gt;Email verification&lt;/li&gt;
&lt;li&gt;CAPTCHA&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These are friction points for agents trying to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Discover tools&lt;/li&gt;
&lt;li&gt;Authenticate&lt;/li&gt;
&lt;li&gt;Execute tasks autonomously&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  4. Documentation Isn’t Machine-Readable
&lt;/h3&gt;

&lt;p&gt;Docs today are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Written for humans&lt;/li&gt;
&lt;li&gt;Scattered across pages&lt;/li&gt;
&lt;li&gt;Hard to parse programmatically&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured capability descriptions&lt;/li&gt;
&lt;li&gt;Executable contracts&lt;/li&gt;
&lt;li&gt;Clear input/output expectations&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  APIs Alone Are Not the Answer
&lt;/h2&gt;

&lt;p&gt;A common assumption is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“We already have APIs, so we’re agent-ready.”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That’s not true.&lt;/p&gt;

&lt;p&gt;APIs are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Too generic&lt;/li&gt;
&lt;li&gt;Often inconsistent&lt;/li&gt;
&lt;li&gt;Not designed for autonomous decision-making&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents need more than endpoints.&lt;/p&gt;

&lt;p&gt;They need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Action schemas&lt;/strong&gt; (what can be done, not just how)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic contracts&lt;/strong&gt; (guaranteed outputs)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability discovery&lt;/strong&gt; (what tools exist and when to use them)&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  What “Agent-First Software” Actually Looks Like
&lt;/h2&gt;

&lt;p&gt;If we design systems for agents as first-class users, the architecture changes.&lt;/p&gt;
&lt;h3&gt;
  
  
  1. Machine-Readable Interfaces
&lt;/h3&gt;

&lt;p&gt;Instead of UI-first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Structured APIs with strict schemas&lt;/li&gt;
&lt;li&gt;Tool definitions with clear contracts&lt;/li&gt;
&lt;li&gt;Standardized input/output formats&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  2. Programmatic Onboarding
&lt;/h3&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Signup → verify → explore&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Agents should:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Discover → authenticate → execute&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This requires:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Auto-provisioned credentials&lt;/li&gt;
&lt;li&gt;Machine-readable pricing/limits&lt;/li&gt;
&lt;li&gt;Capability endpoints&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  3. Permissioned Execution
&lt;/h3&gt;

&lt;p&gt;Agents need controlled autonomy:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Scoped access tokens&lt;/li&gt;
&lt;li&gt;Role-based permissions&lt;/li&gt;
&lt;li&gt;Execution boundaries&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  4. Deterministic Execution Layer
&lt;/h3&gt;

&lt;p&gt;Every action should be:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Predictable&lt;/li&gt;
&lt;li&gt;Retry-safe&lt;/li&gt;
&lt;li&gt;Observable&lt;/li&gt;
&lt;/ul&gt;


&lt;h3&gt;
  
  
  5. Observability for Agents
&lt;/h3&gt;

&lt;p&gt;Traditional logs aren’t enough.&lt;/p&gt;

&lt;p&gt;We need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Decision tracing&lt;/li&gt;
&lt;li&gt;Tool-call lineage&lt;/li&gt;
&lt;li&gt;Cost per execution&lt;/li&gt;
&lt;li&gt;Latency breakdowns&lt;/li&gt;
&lt;/ul&gt;


&lt;h2&gt;
  
  
  A Practical Agent-System Architecture
&lt;/h2&gt;

&lt;p&gt;A simplified flow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Agent
  ↓
Planner (decides what to do)
  ↓
Tool Registry (what tools are available)
  ↓
Execution Layer (calls APIs/tools)
  ↓
Response Validator (ensures correctness)
  ↓
Memory (stores context + learnings)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each layer is critical:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Planner&lt;/strong&gt; → reasoning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool registry&lt;/strong&gt; → discoverability&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Execution&lt;/strong&gt; → action&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validator&lt;/strong&gt; → reliability&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory&lt;/strong&gt; → continuity&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is very different from traditional request-response systems.&lt;/p&gt;




&lt;h2&gt;
  
  
  Where the Opportunity Is
&lt;/h2&gt;

&lt;p&gt;Most people today are focused on:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“How do we build better agents?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;But the bigger opportunity is:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“How do we build better systems for agents to operate on?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Every major category is open:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;CRM → agent-native workflows&lt;/li&gt;
&lt;li&gt;Payments → programmable financial actions&lt;/li&gt;
&lt;li&gt;Support → autonomous resolution systems&lt;/li&gt;
&lt;li&gt;Analytics → queryable, structured insights&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Not as add-ons.&lt;br&gt;
But as &lt;strong&gt;core design principles&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Most People Get Wrong
&lt;/h2&gt;

&lt;h3&gt;
  
  
  ❌ “APIs are enough”
&lt;/h3&gt;

&lt;p&gt;They’re not.&lt;br&gt;
Agents need structured, reliable, discoverable systems.&lt;/p&gt;




&lt;h3&gt;
  
  
  ❌ “Just add AI on top”
&lt;/h3&gt;

&lt;p&gt;That creates brittle layers, not scalable systems.&lt;/p&gt;




&lt;h3&gt;
  
  
  ❌ “Agents will replace software”
&lt;/h3&gt;

&lt;p&gt;No.&lt;br&gt;
Agents will &lt;strong&gt;consume software differently&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Shift That’s Coming
&lt;/h2&gt;

&lt;p&gt;We’re moving from:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Human-first software
→ to&lt;/li&gt;
&lt;li&gt;Agent-first systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn’t a feature upgrade.&lt;/p&gt;

&lt;p&gt;It’s a &lt;strong&gt;paradigm shift in how software is designed and consumed&lt;/strong&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;The companies that win won’t be the ones with the smartest agents.&lt;/p&gt;

&lt;p&gt;They’ll be the ones:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Agents prefer to use.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  👋 If You’re Building in This Space
&lt;/h2&gt;

&lt;p&gt;I’m currently working on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agent-based systems&lt;/li&gt;
&lt;li&gt;Automation architectures&lt;/li&gt;
&lt;li&gt;AI-native SaaS workflows&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you’re exploring similar problems or thinking about building agent-first products, I’d be interested to exchange ideas.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>api</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Engineers Won’t Just Have Salaries - They’ll Have Token Budgets</title>
      <dc:creator>Gaurav Talesara</dc:creator>
      <pubDate>Mon, 13 Apr 2026 18:40:27 +0000</pubDate>
      <link>https://dev.to/gaurav_talesara/engineers-wont-just-have-salaries-theyll-have-token-budgets-3ag0</link>
      <guid>https://dev.to/gaurav_talesara/engineers-wont-just-have-salaries-theyll-have-token-budgets-3ag0</guid>
      <description>&lt;h2&gt;
  
  
  Introduction
&lt;/h2&gt;

&lt;p&gt;There’s a subtle shift happening in how software is being built.&lt;/p&gt;

&lt;p&gt;It’s not loud.&lt;br&gt;
It’s not fully standardized.&lt;br&gt;
But it’s already visible if you look closely.&lt;/p&gt;

&lt;p&gt;We are moving from a world where:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Engineering output was limited by human effort&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;To a world where:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Output is increasingly limited by how much AI you can effectively use&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And that introduces a new concept most teams are not yet fully prepared for:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Token budgets.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What’s Changing Right Now
&lt;/h2&gt;

&lt;p&gt;If you zoom into how modern engineering teams are working:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI tools are no longer optional — they’re embedded in daily workflows&lt;/li&gt;
&lt;li&gt;Engineers are generating, reviewing, and iterating faster than ever&lt;/li&gt;
&lt;li&gt;The bottleneck is shifting from &lt;em&gt;writing code&lt;/em&gt; to &lt;em&gt;orchestrating systems&lt;/em&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Some early signals:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Companies are beginning to track &lt;strong&gt;AI usage per employee&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;AI costs are becoming a &lt;strong&gt;visible line item in engineering budgets&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Token consumption is growing at an &lt;strong&gt;unpredictable pace&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This isn’t theoretical.&lt;/p&gt;

&lt;p&gt;It’s already happening in pockets of the industry.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Real Constraint Has Changed
&lt;/h2&gt;

&lt;p&gt;Traditionally, engineering constraints looked like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Developer bandwidth&lt;/li&gt;
&lt;li&gt;System architecture&lt;/li&gt;
&lt;li&gt;Infrastructure scaling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now there’s a new constraint emerging:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Effective AI utilization&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Two engineers today are no longer equal if:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One uses AI occasionally&lt;/li&gt;
&lt;li&gt;The other builds workflows, agents, and automation around it&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The second engineer is operating with &lt;strong&gt;leverage&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;And that leverage is powered by tokens.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why Token Budgets Will Emerge
&lt;/h2&gt;

&lt;p&gt;Right now, most companies are in an &lt;strong&gt;experimental phase&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pay-as-you-go AI usage&lt;/li&gt;
&lt;li&gt;No clear limits&lt;/li&gt;
&lt;li&gt;Costs that are hard to predict&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This doesn’t scale.&lt;/p&gt;

&lt;p&gt;As usage increases, companies will need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cost control&lt;/li&gt;
&lt;li&gt;Predictability&lt;/li&gt;
&lt;li&gt;Fair distribution of resources&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The natural evolution?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Allocated token budgets per engineer or team&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Just like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Cloud budgets&lt;/li&gt;
&lt;li&gt;API rate limits&lt;/li&gt;
&lt;li&gt;SaaS seat allocations&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Tokens = The New Productivity Unit
&lt;/h2&gt;

&lt;p&gt;We’re used to measuring productivity through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Output (features shipped)&lt;/li&gt;
&lt;li&gt;Velocity (story points, sprints)&lt;/li&gt;
&lt;li&gt;Efficiency (time to deliver)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But AI introduces a different layer.&lt;/p&gt;

&lt;p&gt;Now, productivity is increasingly tied to:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How effectively you can convert tokens into outcomes&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Not all token usage is equal.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Some engineers waste tokens on low-value prompts&lt;/li&gt;
&lt;li&gt;Others build reusable systems that compound output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where the real differentiation will happen.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Rise of the “AI-Orchestrating Engineer”
&lt;/h2&gt;

&lt;p&gt;The best engineers in the next phase won’t just:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write clean code&lt;/li&gt;
&lt;li&gt;Design scalable systems&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;They will:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Design &lt;strong&gt;agent workflows&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Optimize &lt;strong&gt;token usage vs output&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Build systems that &lt;strong&gt;act, not just respond&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In other words:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;They will orchestrate intelligence.&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  What This Means for Engineering Leaders
&lt;/h2&gt;

&lt;p&gt;If you’re leading teams today, this shift has implications:&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Budgeting will change
&lt;/h3&gt;

&lt;p&gt;AI costs will move from “tools” to &lt;strong&gt;core infrastructure spend&lt;/strong&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Hiring signals will change
&lt;/h3&gt;

&lt;p&gt;You won’t just evaluate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Coding ability&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You’ll evaluate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;AI leverage&lt;/li&gt;
&lt;li&gt;System thinking&lt;/li&gt;
&lt;li&gt;Automation mindset&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Internal tooling will evolve
&lt;/h3&gt;

&lt;p&gt;Teams will build:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Internal agents&lt;/li&gt;
&lt;li&gt;Workflow automation systems&lt;/li&gt;
&lt;li&gt;Token-efficient pipelines&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Why This Isn’t Mainstream Yet
&lt;/h2&gt;

&lt;p&gt;It’s important to stay grounded.&lt;/p&gt;

&lt;p&gt;Most companies today:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Do NOT have formal token budgets&lt;/li&gt;
&lt;li&gt;Are still figuring out pricing and limits&lt;/li&gt;
&lt;li&gt;Are experimenting without clear governance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is still early-stage behavior.&lt;/p&gt;

&lt;p&gt;But the direction is clear.&lt;/p&gt;




&lt;h2&gt;
  
  
  The Bigger Shift: From Software to Systems
&lt;/h2&gt;

&lt;p&gt;Today’s companies are built on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;APIs&lt;/li&gt;
&lt;li&gt;Databases&lt;/li&gt;
&lt;li&gt;Services&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Tomorrow’s companies will increasingly rely on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Agents&lt;/li&gt;
&lt;li&gt;Workflows&lt;/li&gt;
&lt;li&gt;Token pipelines&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And that changes how value is created.&lt;/p&gt;




&lt;h2&gt;
  
  
  Final Thought
&lt;/h2&gt;

&lt;p&gt;We’re not just adding AI to existing systems.&lt;/p&gt;

&lt;p&gt;We’re redefining how work gets done.&lt;/p&gt;

&lt;p&gt;The question is no longer:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;“How fast can your team build?”&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It’s:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;“How effectively can your team deploy intelligence?”&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And in that world—&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Tokens become leverage.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  Closing
&lt;/h2&gt;

&lt;p&gt;This shift isn’t fully visible yet.&lt;/p&gt;

&lt;p&gt;But it’s already in motion.&lt;/p&gt;

&lt;p&gt;The teams that understand it early will have an advantage that compounds over time.&lt;/p&gt;




&lt;p&gt;If you're building or leading engineering teams right now—&lt;/p&gt;

&lt;p&gt;How are you thinking about AI usage?&lt;/p&gt;

&lt;p&gt;As a tool…&lt;/p&gt;

&lt;p&gt;Or as infrastructure?&lt;/p&gt;




</description>
      <category>ai</category>
      <category>career</category>
      <category>productivity</category>
      <category>softwareengineering</category>
    </item>
  </channel>
</rss>
