<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Agrima Gupta</title>
    <description>The latest articles on DEV Community by Agrima Gupta (@agrima-06).</description>
    <link>https://dev.to/agrima-06</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4125872%2Fe7a3bcb7-fe2b-4e47-94f4-65badc467977.png</url>
      <title>DEV Community: Agrima Gupta</title>
      <link>https://dev.to/agrima-06</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/agrima-06"/>
    <language>en</language>
    <item>
      <title>How Caching Saves Your Website From Thousands of Requests</title>
      <dc:creator>Agrima Gupta</dc:creator>
      <pubDate>Tue, 06 Oct 2026 11:18:00 +0000</pubDate>
      <link>https://dev.to/agrima-06/how-caching-saves-your-website-from-thousands-of-requests-4nda</link>
      <guid>https://dev.to/agrima-06/how-caching-saves-your-website-from-thousands-of-requests-4nda</guid>
      <description>&lt;h1&gt;
  
  
  How Does Netflix Load So Fast? The Secret Is Caching.
&lt;/h1&gt;

&lt;p&gt;Have you ever wondered why Netflix can serve millions of users without making its database work insanely hard for every single request?&lt;/p&gt;

&lt;p&gt;Or why YouTube thumbnails appear almost instantly?&lt;/p&gt;

&lt;p&gt;Or why a website you've already visited sometimes loads noticeably faster the second time?&lt;/p&gt;

&lt;p&gt;One of the biggest reasons is &lt;strong&gt;caching&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;You might have heard this word a hundred times, but the idea behind it is actually pretty simple:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;If you've already done the expensive work once, why do it again?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Let's see how this works in real applications.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;So, What Exactly Is Caching?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine you go to a restaurant and order coffee.&lt;/p&gt;

&lt;p&gt;The waiter doesn't go to the farm, collect coffee beans, roast them, grind them, and make the coffee from scratch every time you ask for another cup.&lt;/p&gt;

&lt;p&gt;The restaurant already has the ingredients ready.&lt;/p&gt;

&lt;p&gt;Caching follows a similar idea.&lt;/p&gt;

&lt;p&gt;Instead of repeatedly fetching or calculating the same data, we temporarily store the result somewhere that can be accessed much faster.&lt;/p&gt;

&lt;p&gt;Without caching, a request might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
 ↓
Server
 ↓
Database
 ↓
Get Data
 ↓
Server
 ↓
User
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User
 ↓
Server
 ↓
Cache ⚡
 ↓
Data
 ↓
User
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database doesn't have to do the same work over and over again.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Let's Take a Simple Example&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine your website has an API:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api/products
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every time someone visits your website, the backend might query the database:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now imagine 10,000 people visit your website.&lt;/p&gt;

&lt;p&gt;You could end up with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10,000 requests
      ↓
10,000 database queries
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But what if the product list only changes occasionally?&lt;/p&gt;

&lt;p&gt;There's no reason to ask the database the exact same question 10,000 times.&lt;/p&gt;

&lt;p&gt;Instead, we can cache the result.&lt;/p&gt;

&lt;p&gt;The first request goes to the database:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Backend
   ↓
Database
   ↓
Products
   ↓
Cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next requests can simply use the cached result:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Backend
   ↓
Cache ⚡
   ↓
Products
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the basic idea behind caching.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Cache Hit vs Cache Miss&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;You'll often hear developers use the terms &lt;strong&gt;cache hit&lt;/strong&gt; and &lt;strong&gt;cache miss&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A cache hit means the data you're looking for is already in the cache.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Cache
   ↓
Found ✅
   ↓
Return Data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A cache miss means the cache doesn't have the data.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Cache
   ↓
Not Found ❌
   ↓
Database
   ↓
Store Result in Cache
   ↓
Return Data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is generally to have a high cache hit rate for data that makes sense to cache.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Your Browser Is Already Using Caching&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Caching isn't something you only encounter in backend development.&lt;/p&gt;

&lt;p&gt;Your browser does it all the time.&lt;/p&gt;

&lt;p&gt;When you visit a website, your browser may store things like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Images
CSS
JavaScript
Fonts
Other static resources
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Suppose you download:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;logo.png
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next time you visit the same website, your browser may already have that image.&lt;/p&gt;

&lt;p&gt;Instead of downloading it again:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Server → Browser
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;the browser can use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser Cache → logo.png
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is one of the reasons websites can feel faster when you revisit them.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Then There Are CDNs&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Now imagine your main server is located in the US, but your users are all over the world.&lt;/p&gt;

&lt;p&gt;A user in India requesting a large image might have to communicate with a server thousands of kilometres away.&lt;/p&gt;

&lt;p&gt;That's where a &lt;strong&gt;CDN — Content Delivery Network&lt;/strong&gt; — becomes useful.&lt;/p&gt;

&lt;p&gt;A CDN stores copies of content across servers distributed around the world.&lt;/p&gt;

&lt;p&gt;Without a CDN:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;India
  ↓
US Server
  ↓
Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With a CDN:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;India
  ↓
Nearby CDN
  ↓
Response ⚡
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CDNs are commonly used for images, videos, JavaScript, CSS, static pages, and sometimes API responses.&lt;/p&gt;

&lt;p&gt;This is one of the techniques that allows large platforms to deliver content efficiently to users around the world.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What About Backend Caching?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Let's move one level deeper.&lt;/p&gt;

&lt;p&gt;Suppose your backend has this endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api/products
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Without caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
  ↓
Backend
  ↓
Database
  ↓
Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
  ↓
Backend
  ↓
Cache
 ↙     ↘
Hit    Miss
 ↓       ↓
Data   Database
         ↓
       Cache
         ↓
      Response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is where technologies like &lt;strong&gt;Redis&lt;/strong&gt; are commonly used.&lt;/p&gt;

&lt;p&gt;Redis is an in-memory data store, which makes it extremely useful for storing frequently accessed data.&lt;/p&gt;

&lt;p&gt;For example, we might store:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Key:
products

Value:
[
  { id: 1, name: "Laptop" },
  { id: 2, name: "Mouse" }
]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Instead of hitting the database for every request, the backend can first check Redis.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;A Simple Backend Example&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine a Python backend.&lt;/p&gt;

&lt;p&gt;Without caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_products&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;

    &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;database&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_products&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every request goes directly to the database.&lt;/p&gt;

&lt;p&gt;With caching:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_products&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;

    &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;

    &lt;span class="n"&gt;products&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;database&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_products&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;products&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;products&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now the logic is basically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Request
   ↓
Check Cache
   ↓
Is data there?
  ↙       ↘
YES       NO
 ↓         ↓
Return   Database
           ↓
         Cache
           ↓
        Return
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Simple idea, but extremely powerful when an application gets large.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;But Here's the Problem With Caching&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Imagine your website stores:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Product price = ₹999
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;in the cache.&lt;/p&gt;

&lt;p&gt;Then the actual price changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;₹999 → ₹899
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the cache still contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;₹999
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now your users are seeing old information.&lt;/p&gt;

&lt;p&gt;This is called &lt;strong&gt;stale data&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;So caches usually have an expiration time, commonly called &lt;strong&gt;TTL — Time To Live&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cache Data
    ↓
TTL = 60 seconds
    ↓
60 seconds pass
    ↓
Cache expires
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The next request can fetch fresh data and store it again.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Cache Invalidation Is Where Things Get Interesting&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;There's a famous saying in software:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"There are only two hard things in Computer Science: cache invalidation and naming things."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;And cache invalidation really can become complicated.&lt;/p&gt;

&lt;p&gt;Suppose we have:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database:
₹899

Cache:
₹999
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the product changes, we need to make sure the cache is updated or removed.&lt;/p&gt;

&lt;p&gt;One approach is to delete the cached value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;cache&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;delete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;product:123&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then the next request fetches fresh data from the database.&lt;/p&gt;

&lt;p&gt;Another approach is to update the cache whenever the database changes.&lt;/p&gt;

&lt;p&gt;The right strategy depends on the application.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Why Not Cache Everything?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;If caching makes things faster, why not cache the entire application?&lt;/p&gt;

&lt;p&gt;Because caching comes with trade-offs.&lt;/p&gt;

&lt;p&gt;Caches consume memory.&lt;/p&gt;

&lt;p&gt;Cached data can become stale.&lt;/p&gt;

&lt;p&gt;And now you have another system to manage.&lt;/p&gt;

&lt;p&gt;Instead of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Backend → Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;your architecture becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Backend → Cache → Database
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now you have to think about what should be cached, how long it should live, and what should happen when the cache is unavailable.&lt;/p&gt;

&lt;p&gt;Caching is powerful, but it isn't free.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;What Makes Good Data for Caching?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;A good candidate is usually data that is:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Frequently requested, expensive to calculate or retrieve, and doesn't change constantly.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Product catalog
Popular posts
API responses
Configuration
Frequently accessed database queries
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Data that changes constantly or must always be completely real-time may be less suitable.&lt;/p&gt;

&lt;p&gt;The important question isn't:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Can I cache this?"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;It's:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Does caching this actually improve my application?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;How Much Faster Can Caching Be?&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Let's use a simplified example.&lt;/p&gt;

&lt;p&gt;Imagine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Database query = 200ms
Cache lookup   = 5ms
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If 1,000 requests all hit the database, that's a lot of unnecessary work.&lt;/p&gt;

&lt;p&gt;But if most of those requests can be served from the cache, the database can focus on requests that actually need it.&lt;/p&gt;

&lt;p&gt;The exact numbers will vary depending on your architecture, network, database, query, and cache setup.&lt;/p&gt;

&lt;p&gt;But the fundamental idea remains:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Avoid expensive work when you already have the answer.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Caching Exists at Multiple Levels&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;One interesting thing about caching is that there isn't just one cache.&lt;/p&gt;

&lt;p&gt;A modern application might look something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;              USER
                ↓
         Browser Cache
                ↓
               CDN
                ↓
        Application Cache
                ↓
             Redis
                ↓
            Database
                ↓
         Database Cache
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Different layers solve different problems.&lt;/p&gt;

&lt;p&gt;And together, they can dramatically reduce the amount of work your backend and database need to perform.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;The Mental Model I Use for Caching&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Whenever you hear the word &lt;strong&gt;cache&lt;/strong&gt;, think about one simple question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;"Do I already have this?"&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If yes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Return it ⚡
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If no:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Fetch / Calculate
       ↓
Store
       ↓
Return
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's caching at its core.&lt;/p&gt;

&lt;p&gt;Everything else — Redis, CDNs, TTLs, cache invalidation, cache hit rates — is built around this basic idea.&lt;/p&gt;




&lt;h2&gt;
  
  
  &lt;strong&gt;Final Takeaway&lt;/strong&gt;
&lt;/h2&gt;

&lt;p&gt;Caching isn't some complicated trick that only companies like Netflix or YouTube use.&lt;/p&gt;

&lt;p&gt;It's a simple idea that becomes incredibly powerful at scale:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Don't do expensive work repeatedly when you can reuse the result.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Your browser caches images.&lt;/p&gt;

&lt;p&gt;CDNs cache content.&lt;/p&gt;

&lt;p&gt;Backends cache API responses.&lt;/p&gt;

&lt;p&gt;Redis stores frequently accessed data.&lt;/p&gt;

&lt;p&gt;Databases have their own caching mechanisms.&lt;/p&gt;

&lt;p&gt;Together, these layers help modern applications serve huge amounts of traffic without forcing the database to handle every single request.&lt;/p&gt;

&lt;p&gt;So the next time Netflix loads your homepage almost instantly, remember:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;It probably isn't asking the database to figure everything out from scratch.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Somewhere along the way, someone already did the work — and the result is waiting in a cache. ⚡&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>database</category>
      <category>redis</category>
    </item>
    <item>
      <title>How API Calls Actually Work: From Button Click to Server Response</title>
      <dc:creator>Agrima Gupta</dc:creator>
      <pubDate>Mon, 05 Oct 2026 16:12:45 +0000</pubDate>
      <link>https://dev.to/agrima-06/how-api-calls-actually-work-from-button-click-to-server-response-1j2h</link>
      <guid>https://dev.to/agrima-06/how-api-calls-actually-work-from-button-click-to-server-response-1j2h</guid>
      <description>&lt;p&gt;If you've worked with React, Next.js, or any frontend framework, you've probably written something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Looks simple, right?&lt;/p&gt;

&lt;p&gt;But I used to wonder:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What actually happens after I call &lt;code&gt;fetch()&lt;/code&gt;?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Where does the request go?&lt;br&gt;&lt;br&gt;
How does the backend know what I want?&lt;br&gt;&lt;br&gt;
Where does the database fit in?&lt;br&gt;&lt;br&gt;
And how does the response come back to my UI?&lt;/p&gt;

&lt;p&gt;Let's break down the entire journey of one API call.&lt;/p&gt;


&lt;h2&gt;
  
  
  The Big Picture
&lt;/h2&gt;

&lt;p&gt;Imagine we have a button:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;onclick=&lt;/span&gt;&lt;span class="s"&gt;"getUsers()"&lt;/span&gt;&lt;span class="nt"&gt;&amp;gt;&lt;/span&gt;
  Get Users
&lt;span class="nt"&gt;&amp;lt;/button&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the user clicks it, the flow looks roughly like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Button Click
     ↓
Frontend JavaScript
     ↓
HTTP Request
     ↓
Backend API
     ↓
Database
     ↓
Backend Response
     ↓
Frontend
     ↓
UI Update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it.&lt;/p&gt;

&lt;p&gt;Now let's see what's actually happening at each step.&lt;/p&gt;




&lt;h2&gt;
  
  
  1. The Button Click
&lt;/h2&gt;

&lt;p&gt;Let's say we want to fetch a list of users.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getUsers&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://localhost:8000/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important line is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;http://localhost:8000/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We're basically telling the browser:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Hey, send a request to this API and give me whatever comes back."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  2. The Browser Sends an HTTP Request
&lt;/h2&gt;

&lt;p&gt;Behind that simple &lt;code&gt;fetch()&lt;/code&gt; call, the browser creates an HTTP request.&lt;/p&gt;

&lt;p&gt;Something like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;GET /api/users
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;GET&lt;/code&gt; is the HTTP method.&lt;/p&gt;

&lt;p&gt;Some common methods are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET     → Get data
POST    → Create data
PUT     → Replace data
PATCH   → Update data
DELETE  → Delete data
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api/users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;basically means:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Give me the users."&lt;/p&gt;
&lt;/blockquote&gt;




&lt;h2&gt;
  
  
  3. The Backend Receives It
&lt;/h2&gt;

&lt;p&gt;Now imagine our backend is built using FastAPI.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_users&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Agrima&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rahul&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nd"&gt;@app.get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/api/users&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;is basically telling the backend:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Whenever someone sends a GET request to &lt;code&gt;/api/users&lt;/code&gt;, run this function."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;So our request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /api/users
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;finds this function:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="nf"&gt;get_users&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  4. What About the Database?
&lt;/h2&gt;

&lt;p&gt;In a real application, we're obviously not going to hardcode users.&lt;/p&gt;

&lt;p&gt;The backend would probably query a database.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;SELECT&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="k"&gt;FROM&lt;/span&gt; &lt;span class="n"&gt;users&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database might return:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1 | Agrima
2 | Rahul
3 | Priya
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The backend then converts this into a response, usually JSON:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Agrima"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Rahul"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;So now our flow becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Frontend
   ↓
API Request
   ↓
Backend
   ↓
Database
   ↓
Backend
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  5. The Response Comes Back
&lt;/h2&gt;

&lt;p&gt;The backend sends an HTTP response.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;200 OK
Content-Type: application/json
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Agrima"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Rahul"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And this is where our frontend gets the data.&lt;/p&gt;

&lt;p&gt;Remember this?&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Now we do:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And &lt;code&gt;data&lt;/code&gt; becomes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Agrima&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Rahul&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We can finally use it in our UI.&lt;/p&gt;




&lt;h2&gt;
  
  
  6. What Does &lt;code&gt;await&lt;/code&gt; Actually Do?
&lt;/h2&gt;

&lt;p&gt;This was another thing that confused me when I started.&lt;/p&gt;

&lt;p&gt;An API request takes time.&lt;/p&gt;

&lt;p&gt;Maybe:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;50ms
200ms
1 second
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;JavaScript can't just assume the response will instantly arrive.&lt;/p&gt;

&lt;p&gt;That's why we use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;We're basically saying:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Send the request and wait for the response before continuing this function."&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That's why API calls are commonly written with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt;&lt;span class="sr"&gt;/awai&lt;/span&gt;&lt;span class="err"&gt;t
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h2&gt;
  
  
  7. What If the API Fails?
&lt;/h2&gt;

&lt;p&gt;Things go wrong. A lot.&lt;/p&gt;

&lt;p&gt;Maybe the server is down.&lt;/p&gt;

&lt;p&gt;Maybe the URL is wrong.&lt;/p&gt;

&lt;p&gt;Maybe the user isn't authenticated.&lt;/p&gt;

&lt;p&gt;Maybe the backend crashes.&lt;/p&gt;

&lt;p&gt;So don't write API calls assuming everything will always work.&lt;/p&gt;

&lt;p&gt;A basic pattern is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getUsers&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

        &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="s2"&gt;`Request failed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;
            &lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;

        &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
            &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;

        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Something went wrong:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="nx"&gt;error&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Some status codes you'll see often:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;200 → Success
201 → Created
400 → Bad Request
401 → Unauthorized
403 → Forbidden
404 → Not Found
500 → Server Error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;






&lt;h1&gt;
  
  
  The Complete Flow
&lt;/h1&gt;

&lt;p&gt;So that tiny line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;actually triggers a much bigger process:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;        USER
          │
          ↓
    Clicks Button
          │
          ↓
     JavaScript
          │
          ↓
     fetch()
          │
          ↓
    HTTP Request
          │
          ↓
      Backend
          │
          ↓
      Database
          │
          ↓
      Backend
          │
          ↓
    HTTP Response
          │
          ↓
     JavaScript
          │
          ↓
       UI Update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;And that's basically the &lt;strong&gt;request-response cycle&lt;/strong&gt;.&lt;/p&gt;




&lt;h1&gt;
  
  
  One More Real-World Example
&lt;/h1&gt;

&lt;p&gt;Think about an AI notes app.&lt;/p&gt;

&lt;p&gt;You type:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;"Summarize these notes"&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;and click &lt;strong&gt;Summarize&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Behind the scenes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You click "Summarize"
        ↓
Next.js frontend
        ↓
POST /api/summarize
        ↓
Backend
        ↓
AI API
        ↓
AI generates summary
        ↓
Backend sends response
        ↓
Frontend receives it
        ↓
Summary appears on screen
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You only see a button and a result.&lt;/p&gt;

&lt;p&gt;But underneath, multiple systems are talking to each other.&lt;/p&gt;

&lt;p&gt;That's the part I find most interesting about APIs.&lt;/p&gt;




&lt;h1&gt;
  
  
  Final Takeaway
&lt;/h1&gt;

&lt;p&gt;An API call isn't just:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/api/users&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It's a conversation between different parts of an application.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client → Request → Server → Process → Response → Client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once you understand this flow, things like &lt;strong&gt;REST APIs, Axios, FastAPI, Express, authentication, CORS, and even AI APIs&lt;/strong&gt; start making a lot more sense.&lt;/p&gt;

&lt;p&gt;So next time you write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(...)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;you'll know what's actually happening behind that one line. 🚀&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;If you're learning APIs right now, try this:&lt;/strong&gt; build a tiny frontend + backend with just three endpoints:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET    /users
POST   /users
DELETE /users/:id
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll learn more by building those three endpoints than by reading ten API tutorials.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>api</category>
      <category>beginners</category>
    </item>
    <item>
      <title>An LLM Is Not Your Backend — Here's What I Learned</title>
      <dc:creator>Agrima Gupta</dc:creator>
      <pubDate>Wed, 16 Sep 2026 04:53:23 +0000</pubDate>
      <link>https://dev.to/agrima-06/an-llm-is-not-your-backend-heres-what-i-learned-3p0f</link>
      <guid>https://dev.to/agrima-06/an-llm-is-not-your-backend-heres-what-i-learned-3p0f</guid>
      <description>&lt;p&gt;When I first started working with AI, I used to think of an LLM as something like a super-smart backend. You give it some input, it understands it, processes it, and gives you an answer. So naturally, I started thinking, "Why do I need so much backend logic? Can't I just tell the LLM what my application does and let it handle everything?"&lt;/p&gt;

&lt;p&gt;Turns out, no.&lt;br&gt;
And understanding why completely changed the way I think about building AI applications.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;So, What Even Is an LLM?&lt;/strong&gt;&lt;br&gt;
LLM stands for Large Language Model. In simple terms, an LLM is a system trained on huge amounts of text that learns patterns in language and uses those patterns to generate text.&lt;/p&gt;

&lt;p&gt;Imagine you've spent your entire life reading books, articles, conversations, documentation, stories, emails, questions, answers, and code. After seeing enough examples, you become really good at understanding how language works and predicting what kind of sentence would make sense next. If someone says, "The sun rises in the..." your brain immediately expects the word "east." If someone says, "Can you pass me the..." you might expect "salt" or "water."&lt;/p&gt;

&lt;p&gt;An LLM does something conceptually similar, except it does this using mathematical representations and neural networks at a massive scale.This is also why LLMs can do so many different things. They can explain concepts, write code, summarize documents, translate languages, generate stories, analyze text, and have conversations. It can feel like you're talking to something that understands everything. But underneath all of that, the model is still working with patterns it learned during training and generating tokens based on those patterns.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is an LLM Just Autocomplete?&lt;/strong&gt;&lt;br&gt;
In a way, yes. But modern LLMs are obviously much more complicated than the autocomplete on your phone.Your keyboard might see "I'll see you" and predict that "tomorrow" could come next. That's a very simple example of predicting what comes next.&lt;/p&gt;

&lt;p&gt;An LLM takes this basic idea to a completely different scale. Instead of looking at just a few words, it can process much larger contexts and has learned incredibly complicated patterns across language and code.That's why you can give it a paragraph of context and ask it to summarize it, give it a programming problem and ask for an implementation, or give it a messy explanation and ask it to turn that into structured information.&lt;/p&gt;

&lt;p&gt;The important thing to understand is that the model doesn't work exactly like a human brain. It processes language mathematically and predicts what should come next based on the context it has received.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How Does an LLM Learn?&lt;/strong&gt;&lt;br&gt;
This is probably the most important part to understand.&lt;/p&gt;

&lt;p&gt;Imagine teaching a child a language. You don't simply give them a dictionary and ask them to memorize every word. They hear people speaking, read things, see patterns, make mistakes, receive feedback, and gradually become better at understanding and producing language.Training an LLM is obviously much more complicated than this, but the basic intuition is somewhat similar.&lt;/p&gt;

&lt;p&gt;The model is exposed to enormous amounts of training data. It processes pieces of text and repeatedly tries to predict what comes next. When its prediction is different from the actual data, the model's parameters are adjusted. This process happens again and again across huge amounts of data.&lt;/p&gt;

&lt;p&gt;Over time, the model becomes extremely good at recognizing patterns in language.&lt;/p&gt;

&lt;p&gt;This doesn't mean it stores every sentence it has seen like a giant database. Instead, the training process changes the model's internal parameters so that it becomes better at representing and generating language.&lt;/p&gt;

&lt;p&gt;Here's a Simple Example,&lt;br&gt;
Imagine I say, "I went to the restaurant and ordered..." Your brain might expect words like "food", "pizza", "dinner", or "burger."&lt;/p&gt;

&lt;p&gt;An LLM does something conceptually similar, but mathematically. It calculates probabilities for possible next tokens based on the context. For example, the probabilities might conceptually look something like 32% for "pizza", 21% for "food", 15% for "dinner", and so on. These numbers are only an illustration, not actual probabilities from a specific model. The important idea is that the model is constantly predicting what should come next based on the context it has.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Then Why Does It Feel Like It's Thinking?&lt;/strong&gt;&lt;br&gt;
This is where things get really interesting.&lt;/p&gt;

&lt;p&gt;If I ask an LLM, "Explain recursion like I'm five," I get a completely different kind of response than if I ask, "Write a Java solution for this problem."&lt;/p&gt;

&lt;p&gt;If I ask it to rewrite an email professionally, it changes its tone. If I give it a long document, it can summarize the important parts. If I give it a programming problem, it can reason through possible approaches.&lt;/p&gt;

&lt;p&gt;It can feel like there is an actual person sitting behind the screen thinking through my question. Modern models can perform surprisingly sophisticated reasoning and problem-solving tasks. But we shouldn't confuse that capability with an LLM being a complete software system.&lt;/p&gt;

&lt;p&gt;The model generates an output.&lt;br&gt;
Your application still needs to decide what happens after that output.&lt;/p&gt;

&lt;p&gt;And that's where the title of this article comes in.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;An LLM Is Not Your Backend&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Let's imagine you're building an online food delivery application with an AI assistant.&lt;/p&gt;

&lt;p&gt;A user says, "Order me a pizza."&lt;/p&gt;

&lt;p&gt;You send that message to an LLM, and the LLM replies, "Sure! Your pizza has been ordered."&lt;/p&gt;

&lt;p&gt;Sounds great.&lt;br&gt;
But wait.&lt;br&gt;
Did it actually order anything?&lt;/p&gt;

&lt;p&gt;No.&lt;br&gt;
The LLM generated a sentence. That's it.&lt;/p&gt;

&lt;p&gt;It didn't check which restaurants are open. It didn't check whether the pizza is available. It didn't verify your address. It didn't process your payment. It didn't create an order in the database. It didn't send the order to the restaurant.&lt;/p&gt;

&lt;p&gt;Those are &lt;em&gt;backend responsibilities.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;This distinction became very clear to me when I started thinking about AI applications as actual software products instead of just chatbots.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;LLM vs Backend&lt;/strong&gt;&lt;br&gt;
The LLM is great at understanding natural language, extracting intent, generating responses, summarizing information, working with context, and deciding which tool might be useful.&lt;/p&gt;

&lt;p&gt;The backend is responsible for things like authentication, authorization, database operations, business rules, transactions, validation, payments, APIs, security, and other deterministic operations.&lt;/p&gt;

&lt;p&gt;These aren't competing responsibilities.&lt;br&gt;
They're different responsibilities.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The LLM provides flexibility and intelligence around human language, while the backend provides control and reliability.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>llm</category>
      <category>ai</category>
      <category>backend</category>
      <category>softwaredevelopment</category>
    </item>
    <item>
      <title>I Built an AI Voice Sales Agent — Here’s the Architecture Behind It</title>
      <dc:creator>Agrima Gupta</dc:creator>
      <pubDate>Tue, 15 Sep 2026 08:55:51 +0000</pubDate>
      <link>https://dev.to/agrima-06/i-built-an-ai-voice-sales-agent-heres-the-architecture-behind-it-g4g</link>
      <guid>https://dev.to/agrima-06/i-built-an-ai-voice-sales-agent-heres-the-architecture-behind-it-g4g</guid>
      <description>&lt;p&gt;I Built an AI Voice Sales Agent — Here’s the Architecture Behind It&lt;/p&gt;

&lt;p&gt;What if a dealer could simply call a number, tell an AI what they need, and place an order through a normal conversation?&lt;/p&gt;

&lt;p&gt;No app.&lt;br&gt;
No searching through products.&lt;br&gt;
No filling out forms.&lt;/p&gt;

&lt;p&gt;Just talk.&lt;/p&gt;

&lt;p&gt;That was the idea behind a project I've been working on: an AI Voice Sales Agent for dealers.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;"I need 50 boxes of Product X and 20 of Product Y."&lt;/p&gt;

&lt;p&gt;The AI should understand the request, check whether the products are available, figure out the dealer's pricing, apply any applicable schemes, confirm the order, and eventually push it into the company's ERP.&lt;/p&gt;

&lt;p&gt;This isn't a finished enterprise product yet. I'm still building and figuring things out, but I wanted to document how I'm approaching the architecture and some of the decisions I've made along the way.&lt;/p&gt;




&lt;p&gt;The Architecture&lt;/p&gt;

&lt;p&gt;The basic flow looks something like this:&lt;/p&gt;

&lt;p&gt;Dealer&lt;br&gt;
   ↓&lt;br&gt;
Phone Call&lt;br&gt;
   ↓&lt;br&gt;
Twilio&lt;br&gt;
   ↓&lt;br&gt;
OpenAI Realtime API&lt;br&gt;
   ↓&lt;br&gt;
LangGraph&lt;br&gt;
   ↓&lt;br&gt;
Business Tools&lt;br&gt;
   ↓&lt;br&gt;
FastAPI Backend&lt;br&gt;
   ↓&lt;br&gt;
PostgreSQL / Redis&lt;br&gt;
   ↓&lt;br&gt;
ERP&lt;/p&gt;

&lt;p&gt;The interesting part isn't just getting an LLM to talk.&lt;/p&gt;

&lt;p&gt;The AI actually needs to do things.&lt;/p&gt;

&lt;p&gt;It should be able to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Understand what the dealer wants&lt;/li&gt;
&lt;li&gt;Search for products&lt;/li&gt;
&lt;li&gt;Check inventory&lt;/li&gt;
&lt;li&gt;Get dealer-specific pricing&lt;/li&gt;
&lt;li&gt;Check applicable schemes&lt;/li&gt;
&lt;li&gt;Remember useful customer context&lt;/li&gt;
&lt;li&gt;Create draft orders&lt;/li&gt;
&lt;li&gt;Eventually sync with an ERP&lt;/li&gt;
&lt;li&gt;Handle failures without making things up&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Let's break down how I'm thinking about each part.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;The Voice Layer&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The first challenge is simple:&lt;/p&gt;

&lt;p&gt;How does the dealer actually talk to the system?&lt;/p&gt;

&lt;p&gt;I'm using Twilio as the telephony layer.&lt;/p&gt;

&lt;p&gt;The basic flow is:&lt;/p&gt;

&lt;p&gt;Dealer calls&lt;br&gt;
    ↓&lt;br&gt;
Twilio receives the call&lt;br&gt;
    ↓&lt;br&gt;
Audio goes to the AI system&lt;br&gt;
    ↓&lt;br&gt;
AI processes the conversation&lt;br&gt;
    ↓&lt;br&gt;
Response is generated&lt;br&gt;
    ↓&lt;br&gt;
Dealer hears the response&lt;/p&gt;

&lt;p&gt;Initially, I thought of voice as simply:&lt;/p&gt;

&lt;p&gt;Speech → Text → LLM → Text → Speech&lt;/p&gt;

&lt;p&gt;But once you start thinking about an actual conversation, latency becomes a huge deal.&lt;/p&gt;

&lt;p&gt;Imagine saying something to an AI and waiting 4–5 seconds for every response.&lt;/p&gt;

&lt;p&gt;Technically, it works.&lt;/p&gt;

&lt;p&gt;As a conversation?&lt;/p&gt;

&lt;p&gt;Not great.&lt;/p&gt;

&lt;p&gt;That's why I'm looking at real-time voice capabilities rather than treating the system like a normal chatbot with a microphone attached to it.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;The AI Layer&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For the conversational intelligence, I'm using the OpenAI Realtime API.&lt;/p&gt;

&lt;p&gt;But here's something I realized pretty quickly:&lt;/p&gt;

&lt;p&gt;The LLM shouldn't be responsible for everything.&lt;/p&gt;

&lt;p&gt;For example, if a dealer says:&lt;/p&gt;

&lt;p&gt;"Give me 50 of the blue ones."&lt;/p&gt;

&lt;p&gt;The AI needs to understand what "blue ones" refers to.&lt;/p&gt;

&lt;p&gt;That requires conversation context.&lt;/p&gt;

&lt;p&gt;The system might have something like:&lt;/p&gt;

&lt;p&gt;Dealer:&lt;br&gt;
ABC Distributors&lt;/p&gt;

&lt;p&gt;Current conversation:&lt;br&gt;
Product: Product X&lt;br&gt;
Variant: Blue&lt;br&gt;
Requested quantity: 50&lt;/p&gt;

&lt;p&gt;Dealer preferences:&lt;br&gt;
Preferred warehouse: Pune&lt;br&gt;
Preferred language: English&lt;/p&gt;

&lt;p&gt;So the AI understands the conversation instead of treating every sentence as an isolated question.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;LangGraph — Making the AI Actually Do Things&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is probably one of the parts I'm most interested in.&lt;/p&gt;

&lt;p&gt;I don't want the LLM to directly interact with my database and randomly decide what to do.&lt;/p&gt;

&lt;p&gt;Instead, I'm giving the agent specific tools.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;search_product()&lt;br&gt;
check_inventory()&lt;br&gt;
get_dealer_price()&lt;br&gt;
get_scheme()&lt;br&gt;
get_customer_history()&lt;br&gt;
create_draft_order()&lt;/p&gt;

&lt;p&gt;So a conversation could look something like:&lt;/p&gt;

&lt;p&gt;Dealer:&lt;br&gt;
"I need 50 units of Product X."&lt;/p&gt;

&lt;p&gt;AI:&lt;br&gt;
"Let me check the availability."&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    ↓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;check_inventory("Product X")&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    ↓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Inventory:&lt;br&gt;
72 units available&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    ↓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;AI:&lt;br&gt;
"We have 72 units available. Would you like me to create the order for 50?"&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    ↓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;Dealer:&lt;br&gt;
"Yes."&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;    ↓
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;

&lt;p&gt;create_draft_order(...)&lt;/p&gt;

&lt;p&gt;This is where LangGraph becomes useful.&lt;/p&gt;

&lt;p&gt;Instead of having one massive prompt trying to handle the entire business workflow, the agent can move through different states and use specific tools.&lt;/p&gt;

&lt;p&gt;Conceptually:&lt;/p&gt;

&lt;p&gt;START&lt;br&gt;
  ↓&lt;br&gt;
Understand Request&lt;br&gt;
  ↓&lt;br&gt;
Identify Product&lt;br&gt;
  ↓&lt;br&gt;
Check Inventory&lt;br&gt;
  ↓&lt;br&gt;
Check Dealer Pricing&lt;br&gt;
  ↓&lt;br&gt;
Apply Scheme&lt;br&gt;
  ↓&lt;br&gt;
Confirm Order&lt;br&gt;
  ↓&lt;br&gt;
Create Draft Order&lt;br&gt;
  ↓&lt;br&gt;
END&lt;/p&gt;

&lt;p&gt;This also makes debugging much easier.&lt;/p&gt;

&lt;p&gt;If something goes wrong, I can ask:&lt;/p&gt;

&lt;p&gt;Which step failed?&lt;/p&gt;

&lt;p&gt;Instead of:&lt;/p&gt;

&lt;p&gt;Why did the AI randomly do that?&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;PostgreSQL — Where the Real Data Lives&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;One thing I definitely don't want is for the AI to become the source of truth.&lt;/p&gt;

&lt;p&gt;The LLM can understand things.&lt;/p&gt;

&lt;p&gt;It can reason.&lt;/p&gt;

&lt;p&gt;It can communicate.&lt;/p&gt;

&lt;p&gt;But it shouldn't be the database.&lt;/p&gt;

&lt;p&gt;The system needs structured data for things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dealers&lt;/li&gt;
&lt;li&gt;Dealer contacts&lt;/li&gt;
&lt;li&gt;Dealer addresses&lt;/li&gt;
&lt;li&gt;Products&lt;/li&gt;
&lt;li&gt;Product variants&lt;/li&gt;
&lt;li&gt;SKUs&lt;/li&gt;
&lt;li&gt;Inventory&lt;/li&gt;
&lt;li&gt;Orders&lt;/li&gt;
&lt;li&gt;Pricing&lt;/li&gt;
&lt;li&gt;Schemes&lt;/li&gt;
&lt;li&gt;Customer preferences&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;Dealer&lt;br&gt;
 ├── Contacts&lt;br&gt;
 ├── Addresses&lt;br&gt;
 ├── Credit Limit&lt;br&gt;
 ├── Language Preference&lt;br&gt;
 └── Preferred Warehouse&lt;/p&gt;

&lt;p&gt;Product&lt;br&gt;
 ├── Product Variant&lt;br&gt;
 ├── SKU&lt;br&gt;
 ├── Pricing&lt;br&gt;
 └── Inventory&lt;/p&gt;

&lt;p&gt;I'm using PostgreSQL for this persistent data.&lt;/p&gt;

&lt;p&gt;I'm also using UUIDs for primary keys and keeping things like timestamps, foreign keys, indexes, and soft-delete support in the domain design.&lt;/p&gt;

&lt;p&gt;I'm still refining the schema, but getting this foundation right is important because adding AI on top of a messy data model isn't going to magically fix it.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;Redis — The Fast Stuff&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Not everything needs to be stored permanently in PostgreSQL.&lt;/p&gt;

&lt;p&gt;A voice conversation can have a lot of temporary state.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Current session&lt;/li&gt;
&lt;li&gt;Conversation state&lt;/li&gt;
&lt;li&gt;Temporary context&lt;/li&gt;
&lt;li&gt;Caching&lt;/li&gt;
&lt;li&gt;Rate limits&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's where Redis comes in.&lt;/p&gt;

&lt;p&gt;The mental model I'm using is basically:&lt;/p&gt;

&lt;p&gt;PostgreSQL&lt;br&gt;
= Persistent business data&lt;/p&gt;

&lt;p&gt;Redis&lt;br&gt;
= Fast temporary state + caching&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;Inventory Is Where Things Get Real&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Let's say a dealer says:&lt;/p&gt;

&lt;p&gt;"I need 100 units."&lt;/p&gt;

&lt;p&gt;The AI can't just say:&lt;/p&gt;

&lt;p&gt;"Sure, I've placed the order."&lt;/p&gt;

&lt;p&gt;It needs to check what's actually available.&lt;/p&gt;

&lt;p&gt;Something like:&lt;/p&gt;

&lt;p&gt;Dealer Request&lt;br&gt;
      ↓&lt;br&gt;
Identify SKU&lt;br&gt;
      ↓&lt;br&gt;
Inventory Service&lt;br&gt;
      ↓&lt;br&gt;
Available Quantity&lt;br&gt;
      ↓&lt;br&gt;
AI Response&lt;/p&gt;

&lt;p&gt;If the system only has 60 units, the AI should say:&lt;/p&gt;

&lt;p&gt;"We currently have 60 units available. Would you like me to create the order for 60?"&lt;/p&gt;

&lt;p&gt;This is a principle I'm trying to stick to throughout the project:&lt;/p&gt;

&lt;p&gt;The AI should reason about business data, not invent business data.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;Dealer-Specific Pricing&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is another place where a normal chatbot approach isn't enough.&lt;/p&gt;

&lt;p&gt;In B2B, everyone doesn't necessarily get the same price.&lt;/p&gt;

&lt;p&gt;Different dealers might have:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Different price lists&lt;/li&gt;
&lt;li&gt;Different discounts&lt;/li&gt;
&lt;li&gt;Different schemes&lt;/li&gt;
&lt;li&gt;Different credit limits&lt;/li&gt;
&lt;li&gt;Different warehouses&lt;/li&gt;
&lt;li&gt;Different purchasing histories&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;So the flow could look like:&lt;/p&gt;

&lt;p&gt;Base Product Price&lt;br&gt;
        ↓&lt;br&gt;
Dealer-specific Price&lt;br&gt;
        ↓&lt;br&gt;
Applicable Scheme&lt;br&gt;
        ↓&lt;br&gt;
Discount&lt;br&gt;
        ↓&lt;br&gt;
Final Price&lt;/p&gt;

&lt;p&gt;But here's an important architectural decision:&lt;/p&gt;

&lt;p&gt;The LLM shouldn't calculate the final business-critical price itself.&lt;/p&gt;

&lt;p&gt;The backend should do that.&lt;/p&gt;

&lt;p&gt;The AI can explain the result to the dealer.&lt;/p&gt;

&lt;p&gt;The actual calculation should come from deterministic business logic.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;Customer Memory&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is probably one of the coolest parts of the idea.&lt;/p&gt;

&lt;p&gt;A useful sales agent shouldn't feel like it has amnesia after every phone call.&lt;/p&gt;

&lt;p&gt;Suppose a dealer usually orders a particular product or prefers a particular warehouse.&lt;/p&gt;

&lt;p&gt;The system could remember useful information such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Preferred warehouse&lt;/li&gt;
&lt;li&gt;Preferred products&lt;/li&gt;
&lt;li&gt;Typical order quantity&lt;/li&gt;
&lt;li&gt;Preferred language&lt;/li&gt;
&lt;li&gt;Previous orders&lt;/li&gt;
&lt;li&gt;Communication preferences&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then a future conversation could be much smoother.&lt;/p&gt;

&lt;p&gt;Instead of asking:&lt;/p&gt;

&lt;p&gt;"Which warehouse do you want?"&lt;/p&gt;

&lt;p&gt;every single time, the system could already know the dealer's preferred warehouse and simply confirm it when needed.&lt;/p&gt;

&lt;p&gt;But there's an important distinction here.&lt;/p&gt;

&lt;p&gt;Not everything the dealer says should automatically become permanent memory.&lt;/p&gt;

&lt;p&gt;Memory needs rules.&lt;/p&gt;

&lt;p&gt;Some information is temporary conversation context.&lt;/p&gt;

&lt;p&gt;Some information is actual customer data.&lt;/p&gt;

&lt;p&gt;And some information probably shouldn't be stored at all.&lt;/p&gt;

&lt;p&gt;That's something I want to handle carefully as the project develops.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;ERP Integration&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Eventually, the AI needs to connect with the systems that the business already uses.&lt;/p&gt;

&lt;p&gt;That's where ERP integration comes in.&lt;/p&gt;

&lt;p&gt;The architecture I'm aiming for is:&lt;/p&gt;

&lt;p&gt;AI Agent&lt;br&gt;
    ↓&lt;br&gt;
Backend&lt;br&gt;
    ↓&lt;br&gt;
Business Logic&lt;br&gt;
    ↓&lt;br&gt;
ERP APIs&lt;br&gt;
    ↓&lt;br&gt;
Orders / Inventory / Customers&lt;/p&gt;

&lt;p&gt;I don't want the AI agent directly modifying ERP data.&lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;p&gt;AI&lt;br&gt;
 ↓&lt;br&gt;
Tool&lt;br&gt;
 ↓&lt;br&gt;
Backend Validation&lt;br&gt;
 ↓&lt;br&gt;
ERP&lt;/p&gt;

&lt;p&gt;That gives us a much safer boundary between the unpredictable nature of AI and the deterministic nature of enterprise systems.&lt;/p&gt;




&lt;ol&gt;
&lt;li&gt;Why I'm NOT Giving the LLM Direct Database Access&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This is probably one of the biggest things I've learned while designing this.&lt;/p&gt;

&lt;p&gt;It can be tempting to just give an LLM database access and say:&lt;/p&gt;

&lt;p&gt;"Do whatever you need."&lt;/p&gt;

&lt;p&gt;But for a system dealing with real orders, pricing and customer information, that's a very bad idea.&lt;/p&gt;

&lt;p&gt;I'd rather have:&lt;/p&gt;

&lt;p&gt;❌ LLM → Database&lt;/p&gt;

&lt;p&gt;✅ LLM → Tool → Backend → Database&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;p&gt;LLM&lt;br&gt;
 ↓&lt;br&gt;
get_inventory("SKU123")&lt;br&gt;
 ↓&lt;br&gt;
Backend validates request&lt;br&gt;
 ↓&lt;br&gt;
Database query&lt;br&gt;
 ↓&lt;br&gt;
Structured result&lt;br&gt;
 ↓&lt;br&gt;
LLM&lt;/p&gt;

&lt;p&gt;Now I have a proper place for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication&lt;/li&gt;
&lt;li&gt;Authorization&lt;/li&gt;
&lt;li&gt;Validation&lt;/li&gt;
&lt;li&gt;Logging&lt;/li&gt;
&lt;li&gt;Security&lt;/li&gt;
&lt;li&gt;Error handling&lt;/li&gt;
&lt;li&gt;Auditing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;And most importantly, I know exactly what the AI is allowed to do.&lt;/p&gt;




&lt;p&gt;The Tech Stack&lt;/p&gt;

&lt;p&gt;Frontend: Next.js&lt;/p&gt;

&lt;p&gt;Styling: Tailwind CSS + shadcn/ui&lt;/p&gt;

&lt;p&gt;Backend: FastAPI&lt;/p&gt;

&lt;p&gt;Database: PostgreSQL&lt;/p&gt;

&lt;p&gt;Cache / State: Redis&lt;/p&gt;

&lt;p&gt;ORM: SQLAlchemy&lt;/p&gt;

&lt;p&gt;Migrations: Alembic&lt;/p&gt;

&lt;p&gt;Voice: Twilio&lt;/p&gt;

&lt;p&gt;Real-time AI: OpenAI Realtime API&lt;/p&gt;

&lt;p&gt;Agent Orchestration: LangGraph&lt;/p&gt;

&lt;p&gt;Vector Search: pgvector&lt;/p&gt;

&lt;p&gt;Version Control: Git + GitHub&lt;/p&gt;

&lt;p&gt;Monitoring: Sentry + OpenTelemetry&lt;/p&gt;

&lt;p&gt;I'm trying to avoid choosing technologies just because they're popular.&lt;/p&gt;

&lt;p&gt;I want every piece of the stack to have a reason for being there.&lt;/p&gt;




&lt;p&gt;What I'm Still Figuring Out&lt;/p&gt;

&lt;p&gt;The project is still a work in progress, so there are a lot of things I'm actively figuring out.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Voice latency&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A voice agent has to feel like a conversation.&lt;/p&gt;

&lt;p&gt;Even a technically correct answer feels bad if it takes too long.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Hallucinations&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent absolutely cannot randomly invent:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Product availability&lt;/li&gt;
&lt;li&gt;Prices&lt;/li&gt;
&lt;li&gt;Discounts&lt;/li&gt;
&lt;li&gt;Order status&lt;/li&gt;
&lt;li&gt;Credit information&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those things need to come from actual systems.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Conversation recovery&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What happens when the dealer says:&lt;/p&gt;

&lt;p&gt;"No, not that one. The other blue one."&lt;/p&gt;

&lt;p&gt;The agent needs enough context to understand what they're referring to.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Tool failures&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What if the inventory service is down?&lt;/p&gt;

&lt;p&gt;The AI shouldn't pretend that everything worked.&lt;/p&gt;

&lt;p&gt;It needs to understand that the tool failed and communicate that properly.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Security&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Once you're dealing with actual business transactions, security becomes a major part of the architecture.&lt;/p&gt;

&lt;p&gt;Things like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication&lt;/li&gt;
&lt;li&gt;Authorization&lt;/li&gt;
&lt;li&gt;PII protection&lt;/li&gt;
&lt;li&gt;Call verification&lt;/li&gt;
&lt;li&gt;Audit logs&lt;/li&gt;
&lt;li&gt;Tool permissions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;can't just be an afterthought.&lt;/p&gt;




&lt;p&gt;The Biggest Thing I've Learned&lt;/p&gt;

&lt;p&gt;The biggest lesson from this project so far is:&lt;/p&gt;

&lt;p&gt;Building an AI application isn't the same as putting an LLM inside an application.&lt;/p&gt;

&lt;p&gt;The LLM is only one part of the system.&lt;/p&gt;

&lt;p&gt;A useful AI product needs:&lt;/p&gt;

&lt;p&gt;AI&lt;br&gt;
+&lt;br&gt;
Business Logic&lt;br&gt;
+&lt;br&gt;
Data&lt;br&gt;
+&lt;br&gt;
Tools&lt;br&gt;
+&lt;br&gt;
State&lt;br&gt;
+&lt;br&gt;
Security&lt;br&gt;
+&lt;br&gt;
Observability&lt;br&gt;
+&lt;br&gt;
Reliable Infrastructure&lt;/p&gt;

&lt;p&gt;The model provides the intelligence.&lt;/p&gt;

&lt;p&gt;But the architecture around it provides the reliability.&lt;/p&gt;

&lt;p&gt;And I think that's an important distinction, especially when moving from AI demos to actual products.&lt;/p&gt;




&lt;p&gt;What's Next?&lt;/p&gt;

&lt;p&gt;Right now, I'm focusing on building the foundation properly before trying to make everything "smart."&lt;/p&gt;

&lt;p&gt;The next things on my list are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Building the core backend APIs&lt;/li&gt;
&lt;li&gt;Implementing the voice pipeline&lt;/li&gt;
&lt;li&gt;Connecting the agent to business tools&lt;/li&gt;
&lt;li&gt;Building the order workflow&lt;/li&gt;
&lt;li&gt;Adding customer memory&lt;/li&gt;
&lt;li&gt;Connecting inventory&lt;/li&gt;
&lt;li&gt;Designing ERP synchronization&lt;/li&gt;
&lt;li&gt;Adding observability&lt;/li&gt;
&lt;li&gt;Testing real-world conversations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The end goal is pretty simple:&lt;/p&gt;

&lt;p&gt;A dealer should be able to pick up a phone and complete a business transaction through a natural conversation.&lt;/p&gt;

&lt;p&gt;Something like:&lt;/p&gt;

&lt;p&gt;Call&lt;br&gt;
 ↓&lt;br&gt;
Talk&lt;br&gt;
 ↓&lt;br&gt;
Confirm&lt;br&gt;
 ↓&lt;br&gt;
Order&lt;/p&gt;

&lt;p&gt;That's the experience I'm trying to build.&lt;/p&gt;




&lt;p&gt;Final Thoughts&lt;/p&gt;

&lt;p&gt;This project has also changed the way I think about software engineering.&lt;/p&gt;

&lt;p&gt;Earlier, I mostly thought about applications as:&lt;/p&gt;

&lt;p&gt;Frontend&lt;br&gt;
+&lt;br&gt;
Backend&lt;br&gt;
+&lt;br&gt;
Database&lt;/p&gt;

&lt;p&gt;Now I'm asking a lot more questions:&lt;/p&gt;

&lt;p&gt;What should the AI be allowed to decide?&lt;/p&gt;

&lt;p&gt;What should the backend decide?&lt;/p&gt;

&lt;p&gt;Where does the actual source of truth live?&lt;/p&gt;

&lt;p&gt;What happens when the AI is wrong?&lt;/p&gt;

&lt;p&gt;What happens when a tool fails?&lt;/p&gt;

&lt;p&gt;How do we recover from a misunderstood request?&lt;/p&gt;

&lt;p&gt;How do we make the whole thing observable?&lt;/p&gt;

&lt;p&gt;And honestly, I'm still figuring out many of these answers.&lt;/p&gt;

&lt;p&gt;That's probably my favorite part of building this.&lt;/p&gt;

&lt;p&gt;I'm not trying to pretend I have the perfect architecture figured out.&lt;/p&gt;

&lt;p&gt;I'm building it, breaking things, learning, and improving it as I go.&lt;/p&gt;

&lt;p&gt;Build → Break → Learn → Improve.&lt;/p&gt;

&lt;p&gt;If you're also building AI agents, voice applications, or AI-powered SaaS products, I'd love to hear what you're working on and what problems you've run into.&lt;/p&gt;

&lt;p&gt;Let's learn from each other. 🚀&lt;/p&gt;

</description>
      <category>ai</category>
      <category>voiceai</category>
      <category>systemdesign</category>
      <category>softwaredevelopment</category>
    </item>
  </channel>
</rss>
