<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Naresh Chandra Lohani</title>
    <description>The latest articles on DEV Community by Naresh Chandra Lohani (@naresh_chandralohani).</description>
    <link>https://dev.to/naresh_chandralohani</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3893656%2F7c19497e-daa2-45b5-a85d-8b6e2b15430a.jpeg</url>
      <title>DEV Community: Naresh Chandra Lohani</title>
      <link>https://dev.to/naresh_chandralohani</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/naresh_chandralohani"/>
    <language>en</language>
    <item>
      <title>Zoho Integration Services: Fixing CRM Sync Failures</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Fri, 25 Sep 2026 06:51:14 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/zoho-integration-services-fixing-crm-sync-failures-3ibd</link>
      <guid>https://dev.to/naresh_chandralohani/zoho-integration-services-fixing-crm-sync-failures-3ibd</guid>
      <description>&lt;p&gt;A Zoho CRM sync can run perfectly for weeks and then start returning:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;INVALID_OAUTHTOKEN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The obvious response is to request another token. That is often the wrong fix.&lt;/p&gt;

&lt;p&gt;We ran into this class of problem while designing &lt;a href="https://www.oodles.com/zoho?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_17" rel="noopener noreferrer"&gt;Zoho Integration Services&lt;/a&gt; around a Node.js backend and PostgreSQL. The integration had background workers, scheduled synchronization, and multiple API requests running concurrently.&lt;/p&gt;

&lt;p&gt;The difficult part was not calling the Zoho API. It was deciding where OAuth state, datacenter configuration, retries, and API-credit consumption belonged.&lt;/p&gt;

&lt;p&gt;This article walks through that failure mode. The implementation uses Node.js 20+ and PostgreSQL, with Zoho CRM API v8 as the integration boundary.&lt;/p&gt;

&lt;p&gt;The key change is simple: treat Zoho authentication and API limits as shared infrastructure, not as properties of individual HTTP requests.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Start with the failure, not the SDK
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;INVALID_OAUTHTOKEN&lt;/code&gt; response gave us the first useful clue. Zoho documents several causes, including using a token against the wrong datacenter and generating too many active access tokens from the same refresh token. Access tokens are valid for one hour.&lt;/p&gt;

&lt;p&gt;Our naive implementation looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Naive: refreshing independently can create competing access tokens.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getZohoToken&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;refreshToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://accounts.zoho.in/oauth/v2/token&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/x-www-form-urlencoded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;URLSearchParams&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="na"&gt;refresh_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;refreshToken&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;client_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ZOHO_CLIENT_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;client_secret&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ZOHO_CLIENT_SECRET&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;grant_type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;refresh_token&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
      &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The problem appears when several workers discover an expired token simultaneously.&lt;/p&gt;

&lt;p&gt;Worker A refreshes.&lt;/p&gt;

&lt;p&gt;Worker B refreshes again.&lt;/p&gt;

&lt;p&gt;Worker C does the same.&lt;/p&gt;

&lt;p&gt;Now different workers can hold different access tokens while the database still contains stale authentication state. Zoho specifically recommends saving and reusing access tokens rather than repeatedly generating new ones.&lt;/p&gt;

&lt;p&gt;That made token acquisition our first shared resource.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Put OAuth state behind one database record
&lt;/h2&gt;

&lt;p&gt;Once multiple workers can refresh credentials, the token belongs in shared state. PostgreSQL gives us a convenient place to coordinate that state.&lt;/p&gt;

&lt;p&gt;We used a table shaped like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="c1"&gt;-- One credential row represents one Zoho organization and datacenter.&lt;/span&gt;
&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;zoho_credentials&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;organization_id&lt;/span&gt; &lt;span class="nb"&gt;bigint&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_domain&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;access_token&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;refresh_token&lt;/span&gt; &lt;span class="nb"&gt;text&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;expires_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;updated_at&lt;/span&gt; &lt;span class="n"&gt;timestamptz&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important field is &lt;code&gt;api_domain&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Zoho's documentation distinguishes datacenters such as US, EU, and IN. A token generated for one domain must not simply be sent to another. The Node.js SDK documentation makes the same environment and domain distinction.&lt;/p&gt;

&lt;p&gt;We therefore store the API domain with the credential instead of hard-coding:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Use the domain returned by Zoho instead of assuming www.zohoapis.com.&lt;/span&gt;
&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;zohoApiUrl&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;apiDomain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;apiDomain&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/crm/v8&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;path&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For an Indian Zoho organization, for example, the authentication server may be &lt;code&gt;accounts.zoho.in&lt;/code&gt;. The exact domain must come from the organization's Zoho configuration.&lt;/p&gt;

&lt;p&gt;There is another OAuth trap worth testing explicitly. Zoho says an authorization code is single-use and valid for only two minutes. A redirect URI must also exactly match the registered URI.&lt;/p&gt;

&lt;p&gt;That means an authorization callback should exchange the code once, then persist the resulting refresh token. It should not become a general-purpose token endpoint.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.zoho.com/crm/developer/docs/api/v8/?utm_source=chatgpt.com" rel="noopener noreferrer"&gt;Zoho CRM API v8 documentation&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Stop spending API credits on avoidable requests
&lt;/h2&gt;

&lt;p&gt;The token problem was only half of the production failure.&lt;/p&gt;

&lt;p&gt;Our next constraint was API consumption.&lt;/p&gt;

&lt;p&gt;Zoho's current CRM API documentation supports batches of up to 100 records for insert operations. Its current platform also provides COQL and Bulk APIs for larger data retrieval workloads.&lt;/p&gt;

&lt;p&gt;A common synchronization loop looks harmless:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Naive: one remote request per local record.&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;contact&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;contacts&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createZohoContact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For 10,000 contacts, that can become 10,000 remote operations.&lt;/p&gt;

&lt;p&gt;The better design is to construct batches at the integration boundary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Zoho CRM API v8 accepts up to 100 records in this insert request.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;insertBatch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;records&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;apiDomain&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nf"&gt;zohoApiUrl&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;apiDomain&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/Contacts&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Zoho-oauthtoken &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Zoho returned HTTP &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The difference is not merely latency.&lt;/p&gt;

&lt;p&gt;Zoho's API-credit model means request volume can become an account-level constraint. Zoho documents &lt;code&gt;TOO_MANY_REQUESTS&lt;/code&gt; when the allowed API usage is exhausted, and its credit model varies by API operation.&lt;/p&gt;

&lt;p&gt;That changes the architecture.&lt;/p&gt;

&lt;p&gt;Retries cannot simply mean "try the same request again." A retry policy must know whether the failed operation is safe to repeat and whether repeating it consumes more quota.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Make retries selective
&lt;/h2&gt;

&lt;p&gt;That API-credit constraint changes how we handle errors.&lt;/p&gt;

&lt;p&gt;For transient HTTP failures, exponential backoff is reasonable. For authentication failures, we refresh once. For validation failures, retrying is useless.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Only authentication failures trigger token refresh; validation errors do not.&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;requestZoho&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;makeRequest&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;refreshToken&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;makeRequest&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;refreshToken&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;makeRequest&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Zoho authentication failed after token refresh&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail is the single retry.&lt;/p&gt;

&lt;p&gt;Without that boundary, a worker can enter an authentication retry loop and turn one failed synchronization into dozens of requests.&lt;/p&gt;

&lt;p&gt;We also separate retryable failures from permanent failures in the queue. A malformed phone number should reach a dead-letter path. A temporary upstream failure should remain retryable.&lt;/p&gt;

&lt;p&gt;This distinction becomes especially important when a synchronization job can process thousands of records.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. We hit the scaling problem at the worker boundary
&lt;/h2&gt;

&lt;p&gt;That selective retry policy exposed the trade-off we had been avoiding: batching reduces API calls, but large batches increase the amount of work that can fail together.&lt;/p&gt;

&lt;p&gt;We implemented this in an anonymized CRM synchronization service with Node.js workers and PostgreSQL. We initially processed records independently because it made error handling straightforward. That approach created unnecessary API calls and made synchronization time proportional to individual records.&lt;/p&gt;

&lt;p&gt;We changed the worker to claim records from PostgreSQL, build Zoho-sized batches, and persist the result of each batch before claiming more work. OAuth credentials were stored centrally, and token refresh was serialized around the shared credential row.&lt;/p&gt;

&lt;p&gt;The important result is intentionally left as a measurement placeholder because it depends on the actual workload rather than a reproducible benchmark:&lt;/p&gt;

&lt;p&gt;Result: [VERIFY: replace with the measured before/after synchronization duration and API-call reduction from the production job.]&lt;/p&gt;

&lt;p&gt;We did not use a benchmark to claim that batching always produces a specific percentage improvement. Zoho's API mix, payload size, CRM edition, network latency, and worker concurrency all affect the result.&lt;/p&gt;

&lt;p&gt;The architectural result was more concrete: one worker no longer had to own authentication state, and a transient Zoho failure no longer caused every worker to independently refresh credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. The production boundary becomes easier to reason about
&lt;/h2&gt;

&lt;p&gt;Once token state, batching, and retries were separated, the worker itself became smaller.&lt;/p&gt;

&lt;p&gt;The worker's job became:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Claim pending records.&lt;/li&gt;
&lt;li&gt;Load shared Zoho credentials.&lt;/li&gt;
&lt;li&gt;Refresh only when required.&lt;/li&gt;
&lt;li&gt;Build API-sized batches.&lt;/li&gt;
&lt;li&gt;Send the batch.&lt;/li&gt;
&lt;li&gt;Persist success or failure.&lt;/li&gt;
&lt;li&gt;Release the claimed records.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That separation also makes observability more useful.&lt;/p&gt;

&lt;p&gt;We log the Zoho organization identifier, operation type, batch size, HTTP status, retry count, and synchronization job ID. We do not log access tokens or refresh tokens.&lt;/p&gt;

&lt;p&gt;The Node.js SDK can manage OAuth details for applications that choose the SDK route, and Zoho publishes the current API v8 SDK separately from its older archived Node SDK. The current v8 repository documents OAuth handling and environment/domain-specific tokens.&lt;/p&gt;

&lt;p&gt;For a small integration, the SDK can reduce boilerplate. For a service with its own queue, credential store, retry policy, and observability, direct REST calls can make those boundaries easier to control.&lt;/p&gt;

&lt;p&gt;That is an architecture decision, not a rule that applies to every &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles Zoho Integration Services&lt;/a&gt; project.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Store Zoho OAuth state centrally when multiple workers can access the same CRM account.&lt;/li&gt;
&lt;li&gt;Keep the Zoho datacenter with the credential, because authentication and API domains are not interchangeable.&lt;/li&gt;
&lt;li&gt;Batch CRM operations instead of creating one remote request per local record.&lt;/li&gt;
&lt;li&gt;Refresh tokens once per authentication failure, then fail rather than entering a retry loop.&lt;/li&gt;
&lt;li&gt;Measure API calls and synchronization time from your own workload instead of assuming a universal performance improvement.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have handled &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Zoho CRM synchronization&lt;/a&gt; at higher worker concurrency, share how you coordinate token refresh and API limits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FAQ&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;What are Zoho Integration Services?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Zoho Integration Services connect Zoho CRM and other Zoho applications with external systems such as PostgreSQL databases, backend services, ERPs, CRMs, and custom applications. A production integration typically handles authentication, data synchronization, API limits, retries, and error recovery.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How do I fix INVALID_OAUTHTOKEN in a Zoho integration?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;First, verify that the access token has not expired and that you are using the correct Zoho datacenter domain. Zoho access tokens are short-lived, so production integrations should store the refresh token securely and refresh access tokens when required rather than repeatedly requesting new authorization flows.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How can I reduce API calls in Zoho CRM integrations?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Batch records wherever the API operation supports it instead of sending one request per record. For large synchronization jobs, also consider whether COQL or the Bulk APIs are more appropriate than repeatedly querying individual records.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Should I use the Zoho CRM SDK or direct REST APIs?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The SDK can reduce authentication and request-handling boilerplate. Direct REST calls can provide more control when your application already has its own queue, retry policy, credential store, logging, and API-rate management. The choice depends on where you want those responsibilities to live.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;How should I handle Zoho API errors in production?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Separate retryable failures from permanent failures. Authentication failures should trigger a controlled token refresh, transient upstream failures can use bounded exponential backoff, and validation or malformed-data errors should be recorded for correction rather than retried indefinitely.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Chatbot Development Services: Fixing Stateful AI Chats</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Thu, 24 Sep 2026 05:41:09 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/chatbot-development-services-fixing-stateful-ai-chats-4i1n</link>
      <guid>https://dev.to/naresh_chandralohani/chatbot-development-services-fixing-stateful-ai-chats-4i1n</guid>
      <description>&lt;p&gt;A chatbot can answer the first message perfectly and still fail as a product.&lt;/p&gt;

&lt;p&gt;We saw the failure mode while building a multi-turn chatbot with Node.js, TypeScript, and the OpenAI Responses API. The first request succeeded. The second request started carrying conversation state. Then TypeScript produced an overload error around &lt;code&gt;previous_response_id&lt;/code&gt;, while our manual-history fallback kept sending more context on every turn.&lt;/p&gt;

&lt;p&gt;That created two separate problems. The compiler issue slowed development, while the history strategy increased request size as conversations grew.&lt;/p&gt;

&lt;p&gt;This is where &lt;a href="https://www.oodles.com/chat-bot?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_16" rel="noopener noreferrer"&gt;chatbot development services&lt;/a&gt; become an architecture problem rather than an API integration exercise.&lt;/p&gt;

&lt;p&gt;The fix is to separate three concerns: conversation identity, model context, and application data. We will build that boundary with Node.js 22+, TypeScript, the official &lt;code&gt;openai&lt;/code&gt; SDK, and PostgreSQL as the application-side store.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Start with the failure, not the framework
&lt;/h2&gt;

&lt;p&gt;The TypeScript error is easy to dismiss because the API request itself looks reasonable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Chatbot Development Services example: the problematic inferred type is the important part.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;instructions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;previous_response_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;previousResponseId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The official OpenAI Node repository has documented a TypeScript inference issue around this pattern. In the reported case, TypeScript could not determine the correct overloaded &lt;code&gt;responses.create()&lt;/code&gt; signature when &lt;code&gt;previous_response_id&lt;/code&gt; was optional. The issue showed &lt;code&gt;response&lt;/code&gt; being inferred as &lt;code&gt;any&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That matters because a chatbot usually has exactly this shape:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;No previous response on the first turn.&lt;/li&gt;
&lt;li&gt;A response ID after the first turn.&lt;/li&gt;
&lt;li&gt;The same code path on every later turn.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;We should make the request type explicit instead of hiding the problem with &lt;code&gt;any&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Explicitly typing the request avoids the optional previous_response_id inference trap.&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;ResponseCreateParamsNonStreaming&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai/resources/responses/responses&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;ResponseCreateParamsNonStreaming&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;instructions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;previous_response_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;previousResponseId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The SDK's current documentation also exposes &lt;code&gt;previous_response_id&lt;/code&gt; specifically for continuing a response-based conversation.&lt;/p&gt;

&lt;p&gt;The important decision is not the type annotation. It is deciding who owns conversation state.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Stop treating the database as the model's conversation buffer
&lt;/h2&gt;

&lt;p&gt;Once &lt;code&gt;previous_response_id&lt;/code&gt; works, the next temptation is to store every assistant message in PostgreSQL and resend the complete transcript.&lt;/p&gt;

&lt;p&gt;That approach looks simple:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Naive approach: rebuilding the complete transcript makes every turn carry old context again.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getMessages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;conversationId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The database still needs the conversation record. It may also need messages for auditing, search, analytics, or compliance.&lt;/p&gt;

&lt;p&gt;But those are application requirements, not necessarily the model's context-management mechanism.&lt;/p&gt;

&lt;p&gt;OpenAI documents several ways to manage conversation state and recommends the Responses API for stateful interactions. The Conversations API can also persist conversation state across sessions and devices.&lt;/p&gt;

&lt;p&gt;For a straightforward request-response chatbot, we can instead persist the provider response identifier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Keep application ownership of the conversation while letting the API carry model context.&lt;/span&gt;
&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;Conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;previousResponseId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;conversation&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getConversation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;conversationId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;instructions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;...(&lt;/span&gt;&lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;previousResponseId&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;previous_response_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;conversation&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;previousResponseId&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{}),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;updateConversation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;conversationId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;previousResponseId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This changes the data model from "our database contains the entire prompt" to "our database knows which conversation state to continue."&lt;/p&gt;

&lt;p&gt;That distinction becomes important when traffic increases.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Keep tool calls inside the same state machine
&lt;/h2&gt;

&lt;p&gt;Once conversation state is externalized, tool calling becomes the next failure point.&lt;/p&gt;

&lt;p&gt;A chatbot that can check orders, create tickets, or query customer records should not treat tool calls as separate conversations. The model can request a function, your application executes it, and the result goes back into the same response chain.&lt;/p&gt;

&lt;p&gt;The official Node SDK documents this loop explicitly. It also warns that a response can contain multiple function calls, so application code must inspect every function-call item.&lt;/p&gt;

&lt;p&gt;A simplified implementation looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Non-obvious part: match function results using call_id, not the optional output-item id.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.5&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;instructions&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outputs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;

&lt;span class="k"&gt;for &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt; &lt;span class="k"&gt;of&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;output&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;function_call&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;continue&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;args&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;arguments&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;get_order&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;order&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getOrder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;args&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;orderId&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="nx"&gt;outputs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;function_call_output&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;call_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;call_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;order&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The application then submits those outputs using &lt;code&gt;previous_response_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This matters because tool execution is where chatbot code stops being a prompt wrapper. The application becomes responsible for authorization, validation, idempotency, and failure handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Protect the context you paid to build
&lt;/h2&gt;

&lt;p&gt;The previous steps solve correctness. They do not automatically solve cost.&lt;/p&gt;

&lt;p&gt;Prompt caching becomes relevant when the chatbot sends a large, stable instruction prefix on repeated requests. OpenAI currently documents prompt caching for supported models and says GPT-5.6 and later have a minimum cacheable prefix of 1,024 visible input tokens. Cached input is priced at a lower rate than uncached input.&lt;/p&gt;

&lt;p&gt;That gives us a concrete ordering rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Keep stable instructions and tool definitions before dynamic user-specific content.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-5.6&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;instructions&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`
    You are the support assistant for Acme.
    Follow the escalation policy.
    Use tools only when required.
  `&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;userMessage&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not casually rebuild the instruction prefix on every turn.&lt;/p&gt;

&lt;p&gt;Changing tool definitions, ordering, schemas, or earlier context can change the reusable prefix. OpenAI's deployment guidance specifically recommends keeping stable instructions, examples, reference material, and tool definitions consistent when optimizing caching.&lt;/p&gt;

&lt;p&gt;For production chatbot development, we therefore monitor cached tokens rather than guessing whether caching is working.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. The production trade-off is state ownership
&lt;/h2&gt;

&lt;p&gt;That trade-off became clear in our implementation: keeping the complete transcript in PostgreSQL gave us maximum application control, while provider-managed state reduced the amount of context our application had to reconstruct.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; initially favored the transcript because it made debugging easy. The failure appeared when every turn required another history reconstruction, while the application also had to preserve tool-call ordering correctly.&lt;/p&gt;

&lt;p&gt;We moved the active model state to the response chain and kept PostgreSQL for durable application records. That removed the repeated history assembly from the hot path.&lt;/p&gt;

&lt;p&gt;The important architectural result is measurable even before that number is inserted: the application no longer needs to rebuild the entire active model context for every turn.&lt;/p&gt;

&lt;p&gt;For high-volume systems, this boundary also makes rate limiting easier to reason about. OpenAI rate limits can apply to requests per minute and tokens per minute, so a chatbot can hit a token limit even when request volume looks acceptable.&lt;/p&gt;

&lt;p&gt;We therefore record at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;conversation ID&lt;/li&gt;
&lt;li&gt;response ID&lt;/li&gt;
&lt;li&gt;model&lt;/li&gt;
&lt;li&gt;input and output token counts&lt;/li&gt;
&lt;li&gt;cached input tokens&lt;/li&gt;
&lt;li&gt;tool calls&lt;/li&gt;
&lt;li&gt;API request ID&lt;/li&gt;
&lt;li&gt;retry count&lt;/li&gt;
&lt;li&gt;total application latency&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Node SDK exposes request IDs and supports configurable retries and timeouts. Its documented defaults should still be checked against the exact SDK version deployed by your application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: What actually matters
&lt;/h2&gt;

&lt;p&gt;The failure was not "the chatbot API broke." The architecture had unclear ownership of conversation state.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use &lt;code&gt;previous_response_id&lt;/code&gt; or the Conversations API when provider-managed conversation state fits the product.&lt;/li&gt;
&lt;li&gt;Keep PostgreSQL responsible for durable business data, audit records, and application-level conversation metadata.&lt;/li&gt;
&lt;li&gt;Type the Responses API request explicitly when optional state produces TypeScript overload inference problems.&lt;/li&gt;
&lt;li&gt;Treat tool calls as part of the same conversation state machine and match results using &lt;code&gt;call_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Measure cached tokens, token usage, retries, and p95 latency before changing prompts or models for performance reasons.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are building a production chatbot, the interesting engineering question is usually not how to make the first response work. It is where conversation state should live after the hundredth response.&lt;/p&gt;

&lt;p&gt;For examples of production-oriented chatbot development services and implementation patterns, Contact us &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;chatbot development service overview&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;What state-management strategy are you using for multi-turn chats: provider-managed state, application-managed history, or a hybrid?&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;FAQ&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;What are Chatbot Development Services?&lt;/p&gt;

&lt;p&gt;Chatbot Development Services cover the engineering required to build, integrate, deploy, and maintain conversational applications. In a production system, that can include conversation state, LLM integration, tool calling, authentication, databases, observability, rate limiting, and failure handling.&lt;/p&gt;

&lt;p&gt;Should chatbot conversation history live in PostgreSQL?&lt;/p&gt;

&lt;p&gt;Not necessarily. PostgreSQL is useful for durable application data, audit records, analytics, and conversation metadata. For active model context, provider-managed state such as previous_response_id or a conversation API can avoid reconstructing the entire transcript on every request.&lt;/p&gt;

&lt;p&gt;When should I use previous_response_id?&lt;/p&gt;

&lt;p&gt;Use previous_response_id when you want a later Responses API request to continue from an earlier response. It is useful for straightforward multi-turn conversations where the model's previous response should remain part of the conversation state.&lt;/p&gt;

&lt;p&gt;For more complex requirements involving durable conversations across sessions or devices, evaluate the Conversations API instead.&lt;/p&gt;

&lt;p&gt;How do I prevent chatbot API costs from growing with conversation length?&lt;/p&gt;

&lt;p&gt;First, measure input and output tokens rather than assuming where the cost comes from. Then evaluate provider-managed state, prompt caching, summarization, and selective retrieval.&lt;/p&gt;

&lt;p&gt;Prompt caching can reduce the cost of repeated stable prefixes when the request structure satisfies the provider's caching requirements.&lt;/p&gt;

&lt;p&gt;Why does TypeScript sometimes report an overload error with responses.create()?&lt;/p&gt;

&lt;p&gt;Optional properties such as previous_response_id can interact with TypeScript's overload resolution. Instead of suppressing the error with any, explicitly type the request using the SDK's ResponseCreateParamsNonStreaming type and keep the request shape consistent.&lt;/p&gt;

</description>
      <category>powerplatform</category>
      <category>chatgpt</category>
      <category>ai</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How to Optimize Odoo Implementation Services</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Wed, 23 Sep 2026 05:42:46 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-optimize-odoo-implementation-services-kef</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-optimize-odoo-implementation-services-kef</guid>
      <description>&lt;p&gt;An Odoo deployment can become slow even when the server has enough CPU and memory. The common causes are usually inside the application layer: repeated ORM queries, unbatched record creation, inefficient computed fields, missing indexes, or synchronous work inside user requests. These issues become visible when transaction volume grows.&lt;/p&gt;

&lt;p&gt;This is where Odoo Implementation Services require more than module configuration. Developers need to treat the ERP as an application stack involving Python, the Odoo ORM, PostgreSQL, workers, integrations, and scheduled jobs. A useful starting point is to define performance budgets before customization begins and then validate them with profiling and query measurements.&lt;/p&gt;

&lt;p&gt;For teams planning a production rollout, &lt;a href="https://www.oodles.com/odoo-implementation?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_15" rel="noopener noreferrer"&gt;Odoo implementation and integration services&lt;/a&gt; can also include architecture planning, customization, data migration, and post-deployment optimization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;A typical Odoo architecture contains an Odoo application layer, PostgreSQL, background jobs, external integrations, and one or more worker processes. Custom modules sit directly on the ORM, so inefficient Python code can create database bottlenecks without changing the infrastructure.&lt;/p&gt;

&lt;p&gt;Odoo's own documentation provides a useful performance example: iterating over 1,000 partner records without effective prefetching can result in 2,000 database queries, while the ORM's prefetch mechanism can reduce that example to a single query.&lt;/p&gt;

&lt;p&gt;Before implementing custom workflows, establish:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Expected transaction volume and concurrent users.&lt;/li&gt;
&lt;li&gt;Maximum acceptable response time for critical operations.&lt;/li&gt;
&lt;li&gt;Expected daily batch size for imports and scheduled jobs.&lt;/li&gt;
&lt;li&gt;External APIs that participate in synchronous transactions.&lt;/li&gt;
&lt;li&gt;Query-count and database-load baselines.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Odoo also provides an integrated profiler with SQL and periodic collectors, making it possible to identify expensive queries and Python execution paths before changing infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Odoo Implementation Services: A Performance-First Approach
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Profile the ORM Before Changing Infrastructure
&lt;/h3&gt;

&lt;p&gt;Start by measuring the application rather than increasing server resources.&lt;/p&gt;

&lt;p&gt;Odoo's profiler can record SQL queries and stack traces. The SQL collector is useful for finding excessive query counts, while the periodic collector samples execution stacks with relatively low overhead.&lt;/p&gt;

&lt;p&gt;A practical investigation looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reproduce the slow workflow with representative data.&lt;/li&gt;
&lt;li&gt;Record SQL queries and execution traces.&lt;/li&gt;
&lt;li&gt;Identify repeated searches inside loops.&lt;/li&gt;
&lt;li&gt;Check computed fields and relational field access.&lt;/li&gt;
&lt;li&gt;Measure again after each code change.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This avoids treating PostgreSQL, workers, or cloud infrastructure as the default solution to an application-level problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Batch ORM Operations
&lt;/h3&gt;

&lt;p&gt;The next step is to eliminate unnecessary database round trips.&lt;/p&gt;

&lt;p&gt;For example, avoid creating records individually when the ORM can process them as a batch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;values&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;append&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;item&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;})&lt;/span&gt;

&lt;span class="c1"&gt;# Why: batch creation lets Odoo optimize field computation.
&lt;/span&gt;&lt;span class="n"&gt;records&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom.order&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;values&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same principle applies to searches and computed fields. Odoo recommends batch operations because executing SQL-producing methods inside recordset loops can multiply database queries.&lt;/p&gt;

&lt;p&gt;Prefetching is another important consideration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;orders&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale.order&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;browse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order_ids&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Why: browsing the full recordset allows ORM prefetching.
&lt;/span&gt;    &lt;span class="n"&gt;customer_name&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;partner_id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;
    &lt;span class="nf"&gt;process_order&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;customer_name&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Avoid repeatedly browsing individual IDs when the complete recordset is already available.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Move Long Operations Out of Requests
&lt;/h3&gt;

&lt;p&gt;A user-facing HTTP request should not perform a large import, external synchronization, or multi-thousand-record recalculation when the operation can run asynchronously.&lt;/p&gt;

&lt;p&gt;Use scheduled actions or background processing for workloads that do not need an immediate response. Odoo recommends processing scheduled actions in batches so a worker is not blocked for an extended period and timeouts are less likely.&lt;/p&gt;

&lt;p&gt;The trade-off is operational complexity. Asynchronous jobs need retry handling, idempotency, progress tracking, and failure visibility. However, keeping long-running work outside the request path makes response-time behavior easier to control.&lt;/p&gt;

&lt;p&gt;For database-heavy workloads, indexes should also be introduced selectively. Odoo notes that indexes can accelerate searches but consume storage and add overhead to inserts, updates, and deletes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our Odoo Implementation Services projects at Oodles, Virbac India required a centralized planning platform covering three supply-chain functions: sales forecasting, production planning, and procurement. The solution was implemented using Odoo with forecasting, production scheduling, raw-material requirement calculations, role-based access controls, and audit capabilities. The planning system also incorporated five years of historical sales data for forecasting workflows.&lt;/p&gt;

&lt;p&gt;The measurable architectural outcome was consolidation of three previously distinct planning functions into one governed Odoo platform, with forecasting outputs feeding production and procurement workflows. This design reduced the need for disconnected planning processes and established a shared data model for downstream calculations.&lt;/p&gt;

&lt;p&gt;The implementation illustrates why performance work should begin with workflow boundaries. If forecasting, inventory, procurement, and reporting each implement independent database access patterns, query volume can grow rapidly. A shared Odoo data model, batch processing, targeted indexes, and profiling provide a more controlled foundation.&lt;/p&gt;

&lt;p&gt;You can explore more engineering and ERP work from &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion / Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Profile before scaling: identify expensive Python paths and SQL queries before adding infrastructure.&lt;/li&gt;
&lt;li&gt;Batch ORM work: recordsets and batch creation reduce unnecessary database round trips.&lt;/li&gt;
&lt;li&gt;Protect request latency: move imports, synchronization, and heavy calculations into controlled background jobs.&lt;/li&gt;
&lt;li&gt;Index selectively: index fields used by important search domains, but account for write overhead.&lt;/li&gt;
&lt;li&gt;Measure query behavior: Odoo supports query-count testing, including &lt;code&gt;assertQueryCount()&lt;/code&gt;, for regression checks.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Start a Technical Discussion
&lt;/h2&gt;

&lt;p&gt;If you are implementing Odoo and have questions about ORM performance, PostgreSQL tuning, integrations, data migration, or custom module architecture, share your scenario in the comments.&lt;/p&gt;

&lt;p&gt;For project-specific requirements, discuss &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Odoo Implementation Services&lt;/a&gt; with the Oodles engineering team.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What are Odoo Implementation Services?
&lt;/h3&gt;

&lt;p&gt;Odoo Implementation Services cover activities required to deploy and adapt Odoo for business operations, including requirements analysis, module configuration, customization, integrations, data migration, testing, deployment, training, and post-launch support. Performance engineering can be included when custom workflows or transaction volumes require application-level optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. How do I diagnose slow Odoo operations?
&lt;/h3&gt;

&lt;p&gt;Start with Odoo's built-in profiler and inspect SQL queries, Python execution traces, computed fields, and ORM access patterns. Reproduce the operation with realistic data, measure query counts, identify repeated work, apply one optimization, and profile again to verify the change.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Why does Odoo become slow with large recordsets?
&lt;/h3&gt;

&lt;p&gt;Large recordsets can expose inefficient loops, repeated searches, poor algorithmic complexity, or missing indexes. Odoo's ORM uses caching and prefetching, but custom code can bypass those benefits. Batch operations, recordset-aware logic, appropriate indexes, and profiling help isolate the actual bottleneck.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Should Odoo custom code use raw SQL?
&lt;/h3&gt;

&lt;p&gt;Raw SQL can be appropriate for complex queries or specific performance requirements, but it bypasses Odoo's ORM security and behavior. Odoo recommends using ORM utilities where practical and requires developers to account for flushing before querying data that may still be pending in the ORM.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. When should I use Odoo Implementation Services instead of configuring Odoo myself?
&lt;/h3&gt;

&lt;p&gt;Odoo Implementation Services are particularly relevant when a deployment involves multiple business modules, custom workflows, external integrations, large data migrations, or production-scale infrastructure. A technical implementation process can establish architecture, testing, performance baselines, deployment procedures, and post-launch support alongside functional configuration.&lt;/p&gt;

</description>
      <category>odoo</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>CRM Software Development Company: Building Fast APIs</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Tue, 22 Sep 2026 02:29:29 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/crm-software-development-company-building-fast-apis-1pdg</link>
      <guid>https://dev.to/naresh_chandralohani/crm-software-development-company-building-fast-apis-1pdg</guid>
      <description>&lt;p&gt;A CRM API can become slow long before the database is technically overloaded. The usual cause is architectural: one request fetches customer data, activities, deals, permissions, notifications, and analytics synchronously. A CRM Software Development Company building for high-concurrency workloads needs to separate transactional paths from secondary work instead of adding more database capacity.&lt;/p&gt;

&lt;p&gt;This article shows a practical Node.js and AWS architecture for reducing API contention in CRM systems. The approach applies to customer profiles, sales pipelines, activity timelines, and tenant-specific dashboards. For teams evaluating implementation options, see Oodles' &lt;a href="https://www.oodles.com/video/crm-applications?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_14" rel="noopener noreferrer"&gt;CRM application development services&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;A CRM backend typically has four latency-sensitive layers: API processing, authorization, data access, and asynchronous business events.&lt;/p&gt;

&lt;p&gt;A useful baseline architecture is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
API Gateway / Load Balancer
  |
Node.js API
  |
  +-- Authentication / Tenant Context
  |
  +-- CRM Service
  |      |
  |      +-- PostgreSQL / DynamoDB
  |      +-- Redis
  |
  +-- Event Publisher
           |
           +-- Queue
                 |
                 +-- Notifications
                 +-- Search indexing
                 +-- Analytics
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important design decision is that creating a contact should not wait for every downstream operation to finish.&lt;/p&gt;

&lt;p&gt;Node.js is designed around an event loop and non-blocking I/O, but expensive callbacks can still block other requests. Node.js documentation specifically warns that blocking the Event Loop reduces throughput because incoming requests share that execution path.&lt;/p&gt;

&lt;p&gt;For workloads using DynamoDB, AWS documents single-digit millisecond latency for singleton operations when the primary key is fully specified. That measurement applies to the DynamoDB service itself and does not include application or network overhead.&lt;/p&gt;

&lt;h2&gt;
  
  
  CRM Software Development Company Architecture for Write-Heavy APIs
&lt;/h2&gt;

&lt;p&gt;The core solution is to keep the synchronous transaction small and move non-critical work into an event-driven pipeline.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Separate the Command Path
&lt;/h3&gt;

&lt;p&gt;A CRM write endpoint should validate the request, authorize the tenant, persist the primary record, and return.&lt;/p&gt;

&lt;p&gt;Do not send emails, rebuild search indexes, calculate reports, or call multiple third-party APIs before responding.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/contacts&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Why: authenticate before touching tenant-specific CRM data.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;tenantId&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: validation prevents malformed records from entering the write path.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;contact&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;validateContact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: the primary transaction completes without waiting for secondary systems.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;saved&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;contactService&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: publish follow-up work after the business record exists.&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;eventBus&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;publish&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;contact.created&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;contactId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;saved&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;201&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;saved&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The event can then trigger independent workers for email, search indexing, audit records, or analytics.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Design Reads Around Access Patterns
&lt;/h3&gt;

&lt;p&gt;A CRM Software Development Company should model storage around actual queries rather than beginning with entities alone.&lt;/p&gt;

&lt;p&gt;For example, a sales dashboard may repeatedly request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tenant + pipeline + status + updated_at
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That access pattern should influence the index or partition design.&lt;/p&gt;

&lt;p&gt;With DynamoDB, AWS recommends data modeling that minimizes joins and structures data around application access patterns.&lt;/p&gt;

&lt;p&gt;For relational systems, the same principle can be applied through carefully selected composite indexes, query-specific projections, and avoiding unnecessary joins.&lt;/p&gt;

&lt;p&gt;A simple Node.js repository method might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getOpenDeals&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Why: query only the fields required by the pipeline screen.&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;deals&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findMany&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;where&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;tenantId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="nx"&gt;ownerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;OPEN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;select&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;updatedAt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The goal is not to make every query fast through caching. The goal is to make the query itself appropriate for the screen requesting it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Push Secondary Work to Workers
&lt;/h3&gt;

&lt;p&gt;Notifications and integrations are good candidates for queues.&lt;/p&gt;

&lt;p&gt;A worker can consume &lt;code&gt;deal.updated&lt;/code&gt; events and independently update search indexes or notify account managers.&lt;/p&gt;

&lt;p&gt;This approach introduces eventual consistency. A newly updated deal might appear in the primary CRM view immediately while its search representation updates a moment later.&lt;/p&gt;

&lt;p&gt;That trade-off is usually acceptable for secondary projections, but not for operations such as authorization or financial state transitions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our CRM-focused implementations at Oodles, the architectural pattern centers on separating customer-facing transactional operations from background processing. The system uses API services for core CRM operations while secondary activities such as notifications and integrations are handled asynchronously.&lt;/p&gt;

&lt;p&gt;For measurable performance validation, teams should capture p50, p95, and p99 API latency before and after the architectural change rather than relying on averages alone.&lt;/p&gt;

&lt;p&gt;AWS provides a useful external benchmark for this design choice: DynamoDB reports single-digit millisecond service latency for singleton operations, while AWS also notes that client-side processing and network transport contribute additional latency.&lt;/p&gt;

&lt;p&gt;For implementation guidance and architecture discussions, &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; works across CRM application architecture, backend services, cloud infrastructure, and integration requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Keep CRM write transactions focused on business-critical persistence.&lt;/li&gt;
&lt;li&gt;Move notifications, indexing, analytics, and integrations to asynchronous workers.&lt;/li&gt;
&lt;li&gt;Model database indexes around real CRM access patterns.&lt;/li&gt;
&lt;li&gt;Measure p50, p95, and p99 latency instead of using average response time alone.&lt;/li&gt;
&lt;li&gt;Prevent CPU-heavy operations from blocking the Node.js Event Loop.&lt;/li&gt;
&lt;li&gt;Treat eventual consistency as an explicit architectural decision, not an accidental side effect.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Building a CRM backend with high request volume, multiple integrations, or tenant isolation? Share your architecture or bottleneck in the comments. The most useful details are request volume, current database, p95 latency, and the slowest API path.&lt;/p&gt;

&lt;p&gt;For technical consultation, discuss your requirements with a &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;CRM Software Development Company&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What does a CRM Software Development Company build?
&lt;/h3&gt;

&lt;p&gt;A CRM Software Development Company designs and develops systems for managing customer records, sales pipelines, activities, communications, permissions, integrations, reporting, and workflow automation. Depending on requirements, the platform may use monolithic, modular, microservice, or event-driven architecture.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Why should CRM applications use asynchronous processing?
&lt;/h3&gt;

&lt;p&gt;Asynchronous processing prevents non-critical tasks from extending the main API transaction. Operations such as email delivery, search indexing, analytics processing, and webhook delivery can run through queues and workers while the CRM API responds after completing the primary business operation.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Is Node.js suitable for CRM backend development?
&lt;/h3&gt;

&lt;p&gt;Node.js is suitable for CRM backends with many concurrent I/O operations because its event-driven architecture handles network and database operations without blocking the main execution path. CPU-heavy work should be moved to workers or separate services to protect request throughput.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Should a CRM use SQL or NoSQL?
&lt;/h3&gt;

&lt;p&gt;The choice depends on access patterns and consistency requirements. SQL databases fit highly relational CRM workflows and complex transactional queries. NoSQL can fit predictable, high-volume access patterns where horizontal scaling and low-latency key-based operations are priorities.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How does a CRM Software Development Company improve API performance?
&lt;/h3&gt;

&lt;p&gt;A CRM Software Development Company can improve API performance by reducing synchronous work, optimizing database access patterns, introducing appropriate indexes or partitions, caching repeated reads, moving background tasks to queues, and monitoring p95 and p99 latency across production traffic.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How a Video Streaming App Development Company Builds a Low-Latency, Scalable Streaming Pipeline</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Mon, 21 Sep 2026 01:09:27 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-a-video-streaming-app-development-company-builds-a-low-latency-scalable-streaming-pipeline-4j51</link>
      <guid>https://dev.to/naresh_chandralohani/how-a-video-streaming-app-development-company-builds-a-low-latency-scalable-streaming-pipeline-4j51</guid>
      <description>&lt;p&gt;A video platform can have fast APIs and still deliver a poor viewing experience. The common failure appears after the user presses Play: the player waits too long, the first segment arrives late, bitrate switches are unstable, or playback repeatedly stalls. These problems usually originate in the media pipeline rather than the application server.&lt;/p&gt;

&lt;p&gt;A Video Streaming App Development Company must therefore design the player, encoding workflow, object storage, CDN, APIs, and observability layer as one system. For teams evaluating a &lt;a href="https://www.oodles.com/video-streaming?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_12" rel="noopener noreferrer"&gt;video streaming development solution&lt;/a&gt;, the important engineering question is not simply how to stream a file, but how to make every stage measurable and independently scalable.&lt;/p&gt;

&lt;p&gt;This article walks through an AWS-oriented architecture using HLS, S3, MediaConvert, CloudFront, and a Node.js API layer.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The recommended architecture separates control-plane traffic from media delivery.&lt;/p&gt;

&lt;p&gt;A typical request path looks like this:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Client → Node.js API → Authentication / Metadata&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;while the media path becomes:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Client → CloudFront → S3 → HLS Segments&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;For uploaded VOD content, the processing path can be:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Upload → S3 → Event → MediaConvert → HLS ABR → S3 → CloudFront&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;AWS documents a similar VOD architecture using S3 for source and destination media, MediaConvert for transcoding, Lambda for workflow tasks, and CloudFront for distribution.&lt;/p&gt;

&lt;p&gt;HLS is useful here because it breaks media into HTTP-delivered segments and supports adaptive playback across changing network conditions. Apple describes HLS as supporting both live and on-demand delivery while dynamically adapting to available network speed.&lt;/p&gt;

&lt;p&gt;Performance targets should be based on observed playback data rather than arbitrary API latency goals. Akamai research found that abandonment begins rising when video startup exceeds roughly two seconds, with its analysis estimating about a 5.8% increase in abandonment for each additional second of startup delay in the studied data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing the Pipeline as a Video Streaming App Development Company
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Separate Video Processing From the API
&lt;/h3&gt;

&lt;p&gt;The first design decision is to keep video encoding out of the request-response path.&lt;/p&gt;

&lt;p&gt;A Node.js API should create upload sessions, validate metadata, authorize users, and return pre-signed S3 upload URLs. The client can then upload directly to object storage.&lt;/p&gt;

&lt;p&gt;The sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Client requests an upload session.&lt;/li&gt;
&lt;li&gt;Node.js validates authentication and file metadata.&lt;/li&gt;
&lt;li&gt;API generates a pre-signed S3 URL.&lt;/li&gt;
&lt;li&gt;Client uploads the source file directly to S3.&lt;/li&gt;
&lt;li&gt;An S3 event starts the processing workflow.&lt;/li&gt;
&lt;li&gt;MediaConvert generates the required HLS renditions.&lt;/li&gt;
&lt;li&gt;The resulting manifest and segments are published to the delivery bucket.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This avoids keeping application servers occupied while multi-gigabyte files are uploaded.&lt;/p&gt;

&lt;p&gt;AWS's reference implementation follows the same event-driven principle, using S3 uploads to initiate MediaConvert processing and CloudWatch/EventBridge components to track job completion.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Generate Adaptive Bitrate Renditions
&lt;/h3&gt;

&lt;p&gt;The second step is to produce multiple bitrate and resolution variants instead of serving one large MP4.&lt;/p&gt;

&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;270p for constrained connections&lt;/li&gt;
&lt;li&gt;360p for low bandwidth&lt;/li&gt;
&lt;li&gt;540p for moderate bandwidth&lt;/li&gt;
&lt;li&gt;720p for HD playback&lt;/li&gt;
&lt;li&gt;1080p for high-bandwidth devices&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The master HLS playlist allows the player to select an appropriate rendition.&lt;/p&gt;

&lt;p&gt;A simplified Node.js endpoint might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/videos/:id/playback&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;video&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;videoRepository&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;params&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;video&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;404&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Video not found&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: keep media delivery behind the CDN instead of the API server.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;playbackUrl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CDN_URL&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;video&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;hlsPath&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/master.m3u8`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;playbackUrl&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API returns metadata and authorization information. It does not proxy video bytes.&lt;/p&gt;

&lt;p&gt;AWS's VOD Foundation creates HLS adaptive-bitrate outputs and documents a default configuration containing five renditions, including 1080p, 720p, 540p, 360p, and 270p.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Tune CDN Caching and Measure Playback
&lt;/h3&gt;

&lt;p&gt;The third step is controlling what CloudFront caches and measuring what happens when requests miss the cache.&lt;/p&gt;

&lt;p&gt;Video segments are generally strong CDN candidates because many viewers can request the same immutable objects. The manifest requires more careful cache policies because its contents can change.&lt;/p&gt;

&lt;p&gt;Use these metrics:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Startup time: Play request to first rendered frame.&lt;/li&gt;
&lt;li&gt;Rebuffer ratio: Stalled playback time divided by total playback time.&lt;/li&gt;
&lt;li&gt;Bitrate switches: Frequency and direction of rendition changes.&lt;/li&gt;
&lt;li&gt;CDN cache hit rate: Percentage of cacheable requests served from edge locations.&lt;/li&gt;
&lt;li&gt;Origin latency: Time CloudFront spends waiting for the origin.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;AWS specifically exposes CloudFront cache hit rate and origin latency as monitoring metrics.&lt;/p&gt;

&lt;p&gt;Cache-key design also matters. AWS recommends including only the request values that actually affect the response because unnecessary headers, cookies, or query parameters can create duplicate cache objects and reduce cache efficiency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our Video Streaming App Development Company projects at Oodles, the team worked on Streamly, a US streaming platform supporting more than 100 live channels across Roku, Fire TV, Apple TV, Android, iOS, and web. The architecture included Wowza and Flussonic for live delivery, a Drupal CMS for content operations, and Gracenote integration for electronic program guide data. Oodles reports that the platform crossed 90,000 downloads and reached 24,000 monthly recurring users.&lt;/p&gt;

&lt;p&gt;The engineering lesson is architectural: multi-device delivery requires the media pipeline to remain independent from business APIs. Stream ingestion, transcoding, DRM, CDN delivery, CMS operations, and client playback each have different scaling characteristics.&lt;/p&gt;

&lt;p&gt;More details about &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; are available for teams researching similar streaming architectures.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Do not stream media through application servers. Use direct object-storage uploads and CDN-based playback.&lt;/li&gt;
&lt;li&gt;Use adaptive bitrate packaging. Multiple HLS renditions let the player react to changing network conditions.&lt;/li&gt;
&lt;li&gt;Treat CDN configuration as application architecture. Cache keys, TTLs, manifests, and segments affect origin load and playback latency.&lt;/li&gt;
&lt;li&gt;Measure QoE separately from API performance. Startup time and rebuffer ratio describe the viewer's actual experience.&lt;/li&gt;
&lt;li&gt;Make media processing event-driven. Encoding jobs should scale independently from authentication, catalog, and user APIs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Building a streaming platform requires decisions across encoding, storage, CDN delivery, playback, security, and observability. If you are working through a specific architecture or performance problem, share the details in the DEV.to comments.&lt;/p&gt;

&lt;p&gt;For a technical discussion with the engineering team, contact a &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Video Streaming App Development Company&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What architecture is commonly used for video streaming applications?
&lt;/h3&gt;

&lt;p&gt;A common VOD architecture stores source media in object storage, transcodes it into HLS or DASH renditions, stores the outputs separately, and delivers segments through a CDN. Application APIs handle authentication, metadata, entitlements, and playback authorization rather than transferring video bytes.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Why is adaptive bitrate streaming important?
&lt;/h3&gt;

&lt;p&gt;Adaptive bitrate streaming lets a player switch between encoded renditions according to available bandwidth and playback conditions. Instead of forcing every viewer to receive the highest bitrate, the player can select a lower representation when network capacity drops, reducing the probability of playback stalls.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. What does a Video Streaming App Development Company optimize first?
&lt;/h3&gt;

&lt;p&gt;A Video Streaming App Development Company should first establish measurable playback KPIs such as startup time, rebuffer ratio, bitrate stability, CDN cache hit rate, and origin latency. These measurements help engineers identify whether the bottleneck is encoding, player behavior, network delivery, CDN configuration, or backend infrastructure.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Should video files be served directly from an application server?
&lt;/h3&gt;

&lt;p&gt;Usually, no. Application servers are better suited to authentication, authorization, catalog operations, and metadata APIs. Video segments can be stored in object storage and distributed through a CDN, reducing application-server bandwidth consumption and allowing media delivery to scale independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How can streaming performance be debugged systematically?
&lt;/h3&gt;

&lt;p&gt;Start by measuring startup time and rebuffer ratio on real devices and networks. Then correlate playback sessions with CDN cache hits, origin latency, bitrate switches, HTTP errors, and encoding profiles. This separates player-side problems from CDN, origin, encoding, and network problems instead of treating every buffering event as an API issue.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Build Production-Ready Generative AI Development Services with RAG</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Fri, 18 Sep 2026 06:29:47 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-build-production-ready-generative-ai-development-services-with-rag-3pik</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-build-production-ready-generative-ai-development-services-with-rag-3pik</guid>
      <description>&lt;p&gt;A production LLM application can generate fluent answers and still fail when users ask about private documents, rapidly changing data, or domain-specific rules. The problem usually appears when a model is treated as the database instead of as the reasoning layer.&lt;/p&gt;

&lt;p&gt;This is where Generative AI Development Services need a different architecture. A practical approach is to combine an LLM with retrieval, application-level validation, observability, and controlled data access. In this guide, we will build that architecture around Python, FastAPI, a vector database, and an LLM API. If you are evaluating implementation options, Oodles' &lt;a href="https://www.oodles.com/generative-ai?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_11" rel="noopener noreferrer"&gt;Generative AI solutions&lt;/a&gt; cover similar LLM, RAG, and AI application patterns.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The direct answer is to keep knowledge retrieval separate from text generation. The application should first determine what information is relevant, then give that context to the model, rather than asking the model to answer from its pretrained knowledge alone.&lt;/p&gt;

&lt;p&gt;A typical request path looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  |
  v
API Gateway
  |
  v
FastAPI Service
  |
  +----&amp;gt; Query validation
  |
  +----&amp;gt; Embedding model
  |          |
  |          v
  |      Vector DB
  |          |
  |          v
  |      Relevant chunks
  |
  +----&amp;gt; LLM
             |
             v
       Validated response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The architecture becomes particularly useful for internal knowledge assistants, support applications, document search, research tools, and domain-specific copilots.&lt;/p&gt;

&lt;p&gt;There is also a practical reason to design around retrieval and validation. The 2024 Stack Overflow Developer Survey reported that 62% of respondents were already using AI tools in their development process, while 43% said they trusted the accuracy of AI output. The same survey found that 45% of professional developers considered AI tools bad or very bad at handling complex tasks. That gap makes application architecture important rather than treating model output as inherently reliable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generative AI Development Services: Designing the RAG Request Pipeline
&lt;/h2&gt;

&lt;p&gt;The core design has three stages: retrieve the right context, generate a constrained answer, and validate the result before returning it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Build a Retrieval Boundary
&lt;/h3&gt;

&lt;p&gt;The first step is to turn documents into searchable knowledge instead of sending entire files to the model.&lt;/p&gt;

&lt;p&gt;A basic ingestion pipeline is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Extract text from PDFs, HTML, DOCX, or database records.&lt;/li&gt;
&lt;li&gt;Normalize whitespace and remove irrelevant metadata.&lt;/li&gt;
&lt;li&gt;Split documents into chunks with controlled overlap.&lt;/li&gt;
&lt;li&gt;Generate embeddings for each chunk.&lt;/li&gt;
&lt;li&gt;Store embeddings with document identifiers and access-control metadata.&lt;/li&gt;
&lt;li&gt;Retrieve only the chunks relevant to the current query.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Chunk size should be treated as an application parameter, not a universal constant. Large chunks preserve context but consume more tokens. Small chunks improve retrieval precision but can remove relationships between sentences.&lt;/p&gt;

&lt;p&gt;Metadata is equally important. A vector record might contain:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"policy-2026-17"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"department"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"finance"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"access_level"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"internal"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"source_page"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;12&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The API can then apply authorization filters before retrieval. This prevents a technically valid semantic match from becoming an information-security problem.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Generate With Explicit Context
&lt;/h3&gt;

&lt;p&gt;The second step is to make the LLM operate on retrieved evidence rather than unrestricted assumptions.&lt;/p&gt;

&lt;p&gt;A minimal FastAPI implementation can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/ask&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;ask&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;Query&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Why: retrieve evidence before asking the model to generate.
&lt;/span&gt;    &lt;span class="n"&gt;documents&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;retrieve_documents&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;top_k&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;context&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;join&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;
    Answer using only the supplied context.
    If the context does not contain the answer, say so.

    Context:
    &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;context&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;

    Question:
    &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;query&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;question&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;
    &lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;

    &lt;span class="c1"&gt;# Why: keeping generation behind one service makes model replacement easier.
&lt;/span&gt;    &lt;span class="n"&gt;answer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;generate_with_llm&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;answer&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sources&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;document&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;source&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;document&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;documents&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important design decision is not the framework. It is the boundary between retrieval and generation.&lt;/p&gt;

&lt;p&gt;A model provider can change later without forcing the application to redesign its document store, authorization layer, or API contracts. This is especially useful when comparing hosted models, self-hosted models, or smaller domain-specific models.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Add Guardrails and Observability
&lt;/h3&gt;

&lt;p&gt;The third step is to treat model output as untrusted application data.&lt;/p&gt;

&lt;p&gt;A production pipeline should monitor at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Retrieval latency&lt;/li&gt;
&lt;li&gt;LLM latency&lt;/li&gt;
&lt;li&gt;Token usage&lt;/li&gt;
&lt;li&gt;Retrieval hit quality&lt;/li&gt;
&lt;li&gt;Validation failures&lt;/li&gt;
&lt;li&gt;Model/API errors&lt;/li&gt;
&lt;li&gt;User feedback&lt;/li&gt;
&lt;li&gt;Prompt and model versions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For sensitive applications, add output schemas. For example, a support workflow should return structured fields such as &lt;code&gt;intent&lt;/code&gt;, &lt;code&gt;answer&lt;/code&gt;, &lt;code&gt;confidence&lt;/code&gt;, and &lt;code&gt;source_ids&lt;/code&gt; rather than an uncontrolled string.&lt;/p&gt;

&lt;p&gt;Caching can also reduce repeated retrieval and generation work, but it should be applied carefully. Cache keys should include relevant tenant, user-permission, model, and prompt-version information. Otherwise, a response generated for one security context can become visible in another.&lt;/p&gt;

&lt;p&gt;This is one reason production Generative AI Development Services should be treated as software architecture rather than only prompt engineering.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our Oodles projects, we worked on OmniDimension, a conversational ordering system for restaurants. The architecture combined Twilio for voice communication, Google Speech-to-Text for speech recognition, LangChain with ChatGPT for conversational processing, and Stripe for payment handling.&lt;/p&gt;

&lt;p&gt;The system had to interpret spoken orders, work with menu information, calculate totals, and provide payment links during a phone interaction. Oodles reports that content chunking and prompt engineering were used to improve performance, with a response time of about 2 seconds.&lt;/p&gt;

&lt;p&gt;The architecture illustrates an important production pattern: performance improvements came from changes around the model, not simply from selecting a larger model. Content preparation, prompt construction, speech processing, API calls, and payment operations all contribute to the end-to-end latency budget.&lt;/p&gt;

&lt;p&gt;You can explore more engineering work from &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Separate retrieval from generation. The vector store should provide evidence; the LLM should transform that evidence into a useful response.&lt;/li&gt;
&lt;li&gt;Attach authorization metadata to embeddings. Semantic similarity should never bypass application-level access control.&lt;/li&gt;
&lt;li&gt;Measure the complete request path. Model latency is only one part of an AI application's response time.&lt;/li&gt;
&lt;li&gt;Keep the model behind an application interface. This makes model providers and versions replaceable without rewriting the complete system.&lt;/li&gt;
&lt;li&gt;Treat output as untrusted data. Schema validation, source tracking, logging, and monitoring belong in the production architecture.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Production AI applications require more than an LLM endpoint and a prompt. A RAG pipeline provides a controlled way to connect proprietary information with generative models, while API boundaries, authorization filters, structured outputs, and observability turn the prototype into an application that engineers can operate.&lt;/p&gt;

&lt;p&gt;For developers building Generative AI Development Services, the key architectural decision is to make the model one component of the system rather than the system itself.&lt;/p&gt;

&lt;p&gt;Have you implemented RAG, model routing, vector search, or LLM observability in production? Share your architecture or performance bottleneck in the comments.&lt;/p&gt;

&lt;p&gt;For a technical discussion about Generative AI Development Services, connect with the Oodles engineering team through the &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Generative AI Development Services contact page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What is RAG in a generative AI application?
&lt;/h3&gt;

&lt;p&gt;Retrieval-Augmented Generation combines semantic search with an LLM. The application first retrieves relevant information from a controlled knowledge source, places that information into the model context, and then generates an answer. This allows responses to use current or private data without retraining the model.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. When should developers use RAG instead of fine-tuning?
&lt;/h3&gt;

&lt;p&gt;Use RAG when the model needs access to changing, private, or frequently updated information. Fine-tuning is more appropriate when you need to change model behavior, formatting, or domain-specific response patterns. Many systems can use both, but retrieval should usually handle dynamic knowledge.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. How can Generative AI Development Services control hallucinations?
&lt;/h3&gt;

&lt;p&gt;Generative AI Development Services can reduce hallucination risk by retrieving authoritative sources, limiting prompts to retrieved context, requiring structured outputs, returning source references, validating responses, and recording user feedback. These controls reduce unsupported generation but cannot guarantee that every model response will be correct.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. What should be monitored in a production RAG system?
&lt;/h3&gt;

&lt;p&gt;Monitor retrieval latency, generation latency, token consumption, retrieval relevance, model errors, timeout rates, validation failures, source usage, and user feedback. Tracking these metrics separately helps engineers determine whether a performance problem originates in search, application code, network calls, or model generation.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Which technology stack works for a RAG backend?
&lt;/h3&gt;

&lt;p&gt;A Python backend using FastAPI works well for many RAG applications because it integrates easily with embedding libraries, vector databases, and LLM SDKs. Docker can package the service consistently, while PostgreSQL with vector capabilities or a dedicated vector database can store searchable embeddings.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Build Reliable Zoho Integration Services with Node.js and AWS</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Thu, 17 Sep 2026 06:13:55 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-build-reliable-zoho-integration-services-with-nodejs-and-aws-32pn</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-build-reliable-zoho-integration-services-with-nodejs-and-aws-32pn</guid>
      <description>&lt;p&gt;A common integration failure starts with a simple assumption: if a Zoho API call succeeds, the business operation succeeded. In production, that assumption breaks when requests are retried, webhooks arrive twice, access tokens expire, or a downstream database update fails after Zoho has already accepted the request.&lt;/p&gt;

&lt;p&gt;This is where Zoho Integration services need an application layer rather than direct point-to-point API calls. A Node.js service can isolate Zoho APIs from business logic, while AWS services handle queues, retries, secrets, and observability. For teams evaluating an implementation approach, Oodles provides &lt;a href="https://www.oodles.com/zoho?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_10" rel="noopener noreferrer"&gt;Zoho integration services&lt;/a&gt; around these integration patterns.&lt;/p&gt;

&lt;p&gt;The goal is not simply to connect applications. The goal is to make synchronization predictable when the network, APIs, or application instances fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The recommended architecture separates the integration into four layers: API ingestion, business transformation, asynchronous processing, and persistence.&lt;/p&gt;

&lt;p&gt;A typical flow looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Zoho
  |
  | Webhook / REST API
  v
API Gateway
  |
  v
Node.js Lambda
  |
  v
Amazon SQS
  |
  v
Worker Lambda
  |
  +----&amp;gt; PostgreSQL
  |
  +----&amp;gt; Zoho API
  |
  v
CloudWatch
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation matters because Zoho API latency or temporary failures should not block the application that receives an event.&lt;/p&gt;

&lt;p&gt;The 2024 Stack Overflow Developer Survey reported that 62.3% of respondents had used JavaScript during the previous year, making JavaScript a widely represented technology for web and integration development.&lt;/p&gt;

&lt;p&gt;For AWS-based integrations, another important consideration is duplicate processing. AWS explicitly recommends idempotent Lambda functions because event-driven systems can deliver the same event more than once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing Zoho Integration Services Around Idempotency
&lt;/h2&gt;

&lt;p&gt;The key design principle is to make every externally triggered operation safe to repeat. A webhook can be delivered twice, a worker can retry after a timeout, or a network connection can fail after Zoho has processed a request.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Create an integration boundary
&lt;/h3&gt;

&lt;p&gt;Do not allow application code to call Zoho APIs throughout the codebase.&lt;/p&gt;

&lt;p&gt;Instead, create a dedicated client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// zohoClient.js&lt;/span&gt;
&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;createContact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;accessToken&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://www.zohoapis.com/crm/v6/Contacts&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Zoho-oauthtoken &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;accessToken&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
      &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
        &lt;span class="na"&gt;data&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;contact&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
      &lt;span class="p"&gt;})&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;ok&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`Zoho API failed: &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This boundary gives the rest of the system one place to handle authentication, request formatting, response parsing, and API-specific errors.&lt;/p&gt;

&lt;p&gt;Store OAuth credentials in AWS Secrets Manager rather than source code or container environment files committed to Git.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Put retries behind a queue
&lt;/h3&gt;

&lt;p&gt;A synchronous request should not repeatedly call an external API while the user waits.&lt;/p&gt;

&lt;p&gt;Instead:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Validate the incoming event.&lt;/li&gt;
&lt;li&gt;Generate or extract an idempotency key.&lt;/li&gt;
&lt;li&gt;Store the event in SQS.&lt;/li&gt;
&lt;li&gt;Return an acknowledgement.&lt;/li&gt;
&lt;li&gt;Let a worker process the event.&lt;/li&gt;
&lt;li&gt;Retry transient failures.&lt;/li&gt;
&lt;li&gt;Send permanently failed messages to a dead-letter queue.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;AWS recommends explicit retry strategies, exponential backoff, and idempotent processing for Lambda workloads.&lt;/p&gt;

&lt;p&gt;A worker can implement the idempotency check before performing a write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;processContact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;idempotencyKey&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: prevents the same webhook from creating duplicate records.&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;alreadyProcessed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ignored&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;duplicate&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createContact&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;getZohoAccessToken&lt;/span&gt;&lt;span class="p"&gt;(),&lt;/span&gt;
    &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;contact&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: record completion only after the external operation succeeds.&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;markProcessed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;processed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important detail is the ordering. Marking an event as processed before the external operation succeeds can permanently hide a failed transaction.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Separate transient and permanent errors
&lt;/h3&gt;

&lt;p&gt;Not every error deserves a retry.&lt;/p&gt;

&lt;p&gt;A &lt;code&gt;429&lt;/code&gt;, timeout, or temporary &lt;code&gt;5xx&lt;/code&gt; response can generally enter a retry path. A malformed payload or invalid business identifier should normally fail fast.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;shouldRetry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;statusCode&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Why: validation errors should not consume retry capacity.&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;statusCode&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;400&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;statusCode&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;statusCode&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: rate limits and server errors may recover later.&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;statusCode&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;statusCode&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This distinction also improves operational visibility. A dashboard showing 500 retryable failures is more useful when validation failures are not mixed into the same queue.&lt;/p&gt;

&lt;p&gt;AWS documents that asynchronous Lambda invocations can be retried automatically, and recommends handling duplicate events because the same event may be received more than once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;At Oodles, integration projects are typically structured around an application-owned integration layer rather than embedding third-party API calls directly into business services.&lt;/p&gt;

&lt;p&gt;For a representative CRM synchronization architecture, the system can use Node.js for API orchestration, AWS Lambda for execution, SQS for asynchronous processing, PostgreSQL for synchronization state, and CloudWatch for operational monitoring.&lt;/p&gt;

&lt;p&gt;The measurable engineering targets should be established from the project's actual baseline rather than copied from a generic benchmark. Useful measurements include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;P50 and P95 integration latency&lt;/li&gt;
&lt;li&gt;Zoho API error rate&lt;/li&gt;
&lt;li&gt;duplicate-event rate&lt;/li&gt;
&lt;li&gt;retry count per successful transaction&lt;/li&gt;
&lt;li&gt;queue age during traffic spikes&lt;/li&gt;
&lt;li&gt;failed-message recovery time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;AWS specifically recommends load testing Lambda functions to determine suitable timeout values and to identify dependency bottlenecks.&lt;/p&gt;

&lt;p&gt;For production implementations, &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; can apply these measurements to the actual workload instead of treating a generic response-time number as a guaranteed outcome.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Zoho Integration services should have an application boundary that isolates third-party API behavior from core business logic.&lt;/li&gt;
&lt;li&gt;Idempotency is essential because retries and duplicate events are normal characteristics of distributed systems.&lt;/li&gt;
&lt;li&gt;SQS decouples ingestion from processing, allowing temporary Zoho failures without blocking the caller.&lt;/li&gt;
&lt;li&gt;Retry policies should distinguish transient failures from permanent validation errors.&lt;/li&gt;
&lt;li&gt;Performance should be measured against the actual workload, using latency, queue age, error rate, retries, and recovery time rather than generic benchmarks.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If your team is designing a CRM, ERP, finance, or workflow integration and needs to reason through API boundaries, retries, authentication, data synchronization, or AWS architecture, share your technical scenario in the comments.&lt;/p&gt;

&lt;p&gt;For architecture reviews, implementation planning, or integration engineering discussions, contact Oodles through &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Zoho Integration services&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What are Zoho Integration services?
&lt;/h3&gt;

&lt;p&gt;Zoho Integration services connect Zoho applications with external systems such as CRMs, ERPs, databases, payment platforms, and custom applications. A production implementation commonly includes API authentication, data mapping, retries, webhook processing, error handling, logging, and synchronization controls.&lt;/p&gt;

&lt;h3&gt;
  
  
  How should Node.js handle Zoho API failures?
&lt;/h3&gt;

&lt;p&gt;Node.js should classify failures before retrying. Rate-limit responses, network timeouts, and temporary server errors can use bounded retries with backoff. Invalid requests should generally fail without repeated attempts. The worker should also record the operation state so retries do not create duplicate business records.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why is idempotency important for Zoho integrations?
&lt;/h3&gt;

&lt;p&gt;Idempotency ensures that processing the same event more than once produces the same business result. This is important because distributed systems can retry events or deliver duplicates. An idempotency key stored with processing state allows the integration worker to recognize an operation that has already completed.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Zoho Integration services use AWS Lambda?
&lt;/h3&gt;

&lt;p&gt;Yes. Zoho Integration services can use AWS Lambda for webhook handlers and asynchronous workers, with API Gateway for HTTP ingestion, SQS for buffering, Secrets Manager for credentials, and CloudWatch for monitoring. AWS recommends designing Lambda functions to be idempotent when duplicate events are possible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should Zoho API calls be synchronous or asynchronous?
&lt;/h3&gt;

&lt;p&gt;Use synchronous calls when the caller needs an immediate response and the operation is short and predictable. Use asynchronous processing for synchronization jobs, bulk updates, retries, or workflows involving multiple external systems. Queues provide isolation when downstream APIs experience latency, throttling, or temporary failures.&lt;/p&gt;

</description>
      <category>zoho</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>How to Build Inventory Management Services That Stay Consistent Under Concurrency</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Wed, 16 Sep 2026 07:07:13 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-build-inventory-management-services-that-stay-consistent-under-concurrency-3176</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-build-inventory-management-services-that-stay-consistent-under-concurrency-3176</guid>
      <description>&lt;p&gt;An inventory API can return the wrong stock count even when every individual query looks correct. The problem usually appears when multiple orders reserve the same SKU at nearly the same time, while warehouse updates, cancellations, and payment retries are also modifying inventory.&lt;/p&gt;

&lt;p&gt;This is where Inventory Management Services need more than CRUD endpoints. The service must make stock changes atomic, make retries safe, and provide a clear audit trail for every movement.&lt;/p&gt;

&lt;p&gt;In this guide, we will build that design around Node.js, PostgreSQL, and Docker, with Redis and AWS as optional infrastructure components. For teams evaluating an implementation, Oodles provides &lt;a href="https://www.oodles.com/inventory-warehouse-management-?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_09" rel="noopener noreferrer"&gt;inventory and warehouse management solutions&lt;/a&gt; for systems that require inventory synchronization across operational workflows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The core architecture is a transactional inventory service sitting between sales channels and warehouse operations.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Web / Mobile / Marketplace
          |
       API Gateway
          |
    Inventory Service
      |          |
 PostgreSQL     Redis
      |
 Inventory Ledger
      |
 Warehouse / ERP / Shipping
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important design decision is that PostgreSQL remains the source of truth for stock mutations. Redis can accelerate reads or absorb temporary traffic, but it should not independently decide whether a reservation is valid.&lt;/p&gt;

&lt;p&gt;For a useful performance reference, AWS documented a PostgreSQL &lt;code&gt;pgbench&lt;/code&gt; test using Amazon EBS configurations. Its reported transaction throughput ranged from 5,686 TPS on gp2 to 6,956 TPS on io2 for that particular workload and environment. These figures are benchmark-specific, not guarantees for an inventory workload.&lt;/p&gt;

&lt;p&gt;That distinction matters. Inventory workloads should be benchmarked using their own transaction shape, concurrency, indexes, and contention patterns.&lt;/p&gt;

&lt;h2&gt;
  
  
  Inventory Management Services: Transaction-Safe Stock Reservation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Model stock as a transactional resource
&lt;/h3&gt;

&lt;p&gt;Start with a table that represents the current available quantity and another table that records movements.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight sql"&gt;&lt;code&gt;&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;inventory&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;sku_id&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;available&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt; &lt;span class="k"&gt;CHECK&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;available&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;CREATE&lt;/span&gt; &lt;span class="k"&gt;TABLE&lt;/span&gt; &lt;span class="n"&gt;inventory_movements&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="n"&gt;id&lt;/span&gt; &lt;span class="n"&gt;BIGSERIAL&lt;/span&gt; &lt;span class="k"&gt;PRIMARY&lt;/span&gt; &lt;span class="k"&gt;KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;sku_id&lt;/span&gt; &lt;span class="nb"&gt;BIGINT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;quantity&lt;/span&gt; &lt;span class="nb"&gt;INTEGER&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;reason&lt;/span&gt; &lt;span class="nb"&gt;TEXT&lt;/span&gt; &lt;span class="k"&gt;NOT&lt;/span&gt; &lt;span class="k"&gt;NULL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="n"&gt;created_at&lt;/span&gt; &lt;span class="n"&gt;TIMESTAMPTZ&lt;/span&gt; &lt;span class="k"&gt;DEFAULT&lt;/span&gt; &lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Why two tables?&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;inventory&lt;/code&gt; row provides a fast current-state lookup. The movement table provides historical context for receiving, reservation, release, adjustment, and shipment events.&lt;/p&gt;

&lt;p&gt;This also makes reconciliation possible. If the current quantity differs from the expected quantity calculated from movements, the discrepancy can be investigated instead of silently overwriting it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Lock the SKU during reservation
&lt;/h3&gt;

&lt;p&gt;The reservation operation should read and modify the same row inside one database transaction.&lt;/p&gt;

&lt;p&gt;PostgreSQL's &lt;code&gt;FOR UPDATE&lt;/code&gt; locks selected rows against conflicting updates until the transaction ends. PostgreSQL also supports &lt;code&gt;SKIP LOCKED&lt;/code&gt;, which can be useful for queue-like workloads where workers should avoid waiting on already-locked rows.&lt;/p&gt;

&lt;p&gt;A simplified Node.js implementation using &lt;code&gt;pg&lt;/code&gt; looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;reserveStock&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;skuId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;BEGIN&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Why: row lock prevents concurrent reservations from reading stale stock.&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`SELECT available
       FROM inventory
       WHERE sku_id = $1
       FOR UPDATE`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;skuId&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rowCount&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;available&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="nx"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Insufficient inventory&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Why: update happens inside the same transaction as the locked read.&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`UPDATE inventory
       SET available = available - $1
       WHERE sku_id = $2`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;skuId&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="c1"&gt;// Why: the ledger records why the quantity changed.&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="s2"&gt;`INSERT INTO inventory_movements (sku_id, quantity, reason)
       VALUES ($1, $2, $3)`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;skuId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;reservation&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;COMMIT&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ROLLBACK&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The transaction is deliberately small. Do not call external payment, shipping, or warehouse APIs while holding the database lock.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Make retries idempotent
&lt;/h3&gt;

&lt;p&gt;A network timeout does not tell the client whether the reservation succeeded. Retrying blindly can therefore reserve the same quantity twice.&lt;/p&gt;

&lt;p&gt;Add an idempotency key to reservation requests:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /inventory/reservations
Idempotency-Key: order-8472-item-01
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Store that key with the resulting reservation. A repeated request returns the original result instead of performing another stock mutation.&lt;/p&gt;

&lt;p&gt;AWS describes idempotent APIs as a way to make retries safe by allowing repeated requests to produce no additional side effects.&lt;/p&gt;

&lt;p&gt;The trade-off is additional state and cleanup. Idempotency records need a retention policy, and the database must enforce uniqueness on the key.&lt;/p&gt;

&lt;p&gt;This pattern is preferable to relying exclusively on client-side retry logic because the inventory service owns the business invariant.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In an Oodles implementation of an inventory platform, this architecture can be applied to a system handling SKU reservations across storefront and warehouse workflows. The implementation pattern uses PostgreSQL for transactional stock state, an inventory movement ledger for traceability, Node.js APIs for reservations, and asynchronous workers for downstream warehouse synchronization.&lt;/p&gt;

&lt;p&gt;For production reporting, the meaningful metrics should include reservation p95 latency, database lock wait time, transaction rollback rate, duplicate-request rate, and inventory reconciliation discrepancies.&lt;/p&gt;

&lt;p&gt;Because project-specific production measurements are not included in the supplied brief, production numbers should not be fabricated. Instead, an implementation review should establish a baseline first, then report before-and-after measurements from the actual environment.&lt;/p&gt;

&lt;p&gt;Teams looking at the broader architecture can also review &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; for related engineering and warehouse-management capabilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion / Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Treat stock mutation as a transaction, not as a sequence of independent API operations.&lt;/li&gt;
&lt;li&gt;Lock only the inventory rows that must change so unrelated SKUs can continue processing concurrently.&lt;/li&gt;
&lt;li&gt;Use an inventory ledger to preserve the reason and history behind every quantity change.&lt;/li&gt;
&lt;li&gt;Make reservation APIs idempotent so network retries cannot create duplicate reservations.&lt;/li&gt;
&lt;li&gt;Benchmark the real workload, including lock contention and concurrent reservations, rather than treating generic database TPS as an application guarantee.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Discuss the Architecture
&lt;/h2&gt;

&lt;p&gt;How do you handle concurrent reservations, inventory reconciliation, and retry safety in your systems? Share your approach in the DEV.to comments.&lt;/p&gt;

&lt;p&gt;For architecture discussions or implementation requirements, contact &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Inventory Management Services&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What are Inventory Management Services?
&lt;/h3&gt;

&lt;p&gt;Inventory Management Services are software components that track stock quantities, reservations, warehouse movements, adjustments, and availability. A production implementation normally combines transactional database operations, APIs, audit records, synchronization workflows, and monitoring to keep inventory state consistent across operational systems.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. How do you prevent two orders from reserving the same stock?
&lt;/h3&gt;

&lt;p&gt;Use a database transaction with row-level locking or an equivalent atomic update. In PostgreSQL, &lt;code&gt;SELECT ... FOR UPDATE&lt;/code&gt; can lock the inventory row while the application validates and changes available quantity. This prevents conflicting transactions from simultaneously using the same stock state.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Should Redis be the source of truth for inventory?
&lt;/h3&gt;

&lt;p&gt;Usually, no. Redis can provide fast cached availability or support temporary coordination, but transactional inventory state should generally reside in a durable database. The database should determine whether a reservation succeeds, while caches are invalidated or updated after committed changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Why are idempotency keys important for inventory APIs?
&lt;/h3&gt;

&lt;p&gt;Idempotency keys prevent repeated requests from applying the same business operation multiple times. If a client times out after creating a reservation, it can retry using the same key. The server can then return the existing reservation instead of decrementing inventory again.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How should inventory performance be measured?
&lt;/h3&gt;

&lt;p&gt;Measure application-specific metrics such as p50 and p95 reservation latency, transactions per second, lock wait duration, rollback rate, database CPU, connection-pool saturation, and reconciliation errors. Generic database benchmarks are useful references, but they should not replace workload-specific testing.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How an LMS Development Company Can Design a Scalable Node.js and AWS Learning Platform</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Tue, 15 Sep 2026 04:39:45 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-an-lms-development-company-can-design-a-scalable-nodejs-and-aws-learning-platform-3gl6</link>
      <guid>https://dev.to/naresh_chandralohani/how-an-lms-development-company-can-design-a-scalable-nodejs-and-aws-learning-platform-3gl6</guid>
      <description>&lt;p&gt;An LMS can work perfectly with 100 learners and still fail when thousands of users start watching videos, submitting assessments, and refreshing progress dashboards at the same time. The usual problem is not the course UI. It is the architecture behind authentication, progress tracking, content delivery, background jobs, and database access.&lt;/p&gt;

&lt;p&gt;An LMS Development Company designing for this environment should treat the learning platform as a distributed application rather than a collection of CRUD screens. This article presents a practical Node.js and AWS architecture for handling learner activity without turning the database into the system's bottleneck. If you are evaluating implementation options, see Oodles' &lt;a href="https://www.oodles.com/lms?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_08" rel="noopener noreferrer"&gt;LMS development services&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The correct architecture starts by separating synchronous learner actions from asynchronous workloads.&lt;/p&gt;

&lt;p&gt;A typical LMS contains:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Web or mobile clients for learners, instructors, and administrators.&lt;/li&gt;
&lt;li&gt;Node.js APIs for authentication, courses, assessments, enrolments, and progress.&lt;/li&gt;
&lt;li&gt;PostgreSQL or another relational database for transactional records.&lt;/li&gt;
&lt;li&gt;Amazon S3 for documents, images, and course assets.&lt;/li&gt;
&lt;li&gt;Amazon CloudFront for distributing static and media content.&lt;/li&gt;
&lt;li&gt;Amazon SQS for jobs such as certificate generation, notifications, and analytics processing.&lt;/li&gt;
&lt;li&gt;AWS Lambda or containerized workers for asynchronous processing.&lt;/li&gt;
&lt;li&gt;CloudWatch for application and infrastructure observability.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This separation matters because a learner opening a course should not wait for a certificate-generation task or analytics calculation to finish.&lt;/p&gt;

&lt;p&gt;There is also a useful ecosystem signal for this architecture. The 2024 Stack Overflow Developer Survey collected responses from more than 65,000 developers, and 62.3% reported using JavaScript during the previous year. Node.js also remained the most-used web technology in that survey. [Source: Stack Overflow Developer Survey 2024.]&lt;/p&gt;

&lt;p&gt;For AWS Lambda specifically, AWS recommends reusing execution environments, initializing SDK clients and database connections outside the handler, and writing idempotent functions. [Source: AWS Lambda Best Practices.]&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing the LMS Development Company Architecture
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Separate Transactional and Learning Events
&lt;/h3&gt;

&lt;p&gt;The first design decision is to distinguish between operations that require an immediate response and operations that can run later.&lt;/p&gt;

&lt;p&gt;For example, updating a learner's quiz attempt is transactional. Sending an email about the completed course is not.&lt;/p&gt;

&lt;p&gt;A request can therefore follow this path:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Client → API → Database → SQS → Worker → Notification/Analytics&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The API confirms the important database transaction first. The worker then handles secondary work.&lt;/p&gt;

&lt;p&gt;This prevents slow downstream services from increasing API latency.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Keep Node.js Database Access Concurrency-Safe
&lt;/h3&gt;

&lt;p&gt;An LMS Development Company should pay particular attention to database connections when Node.js services run on AWS Lambda.&lt;/p&gt;

&lt;p&gt;Creating a new database connection for every invocation can exhaust database connection limits as Lambda concurrency increases. AWS recommends connection management strategies such as Amazon RDS Proxy for workloads that create frequent short-lived connections. [Source: AWS Lambda documentation.]&lt;/p&gt;

&lt;p&gt;A simplified Node.js example looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;pg&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pg&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;pool&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nx"&gt;pg&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Pool&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;connectionString&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;max&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;pool&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;connect&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SELECT id, progress FROM learner_progress WHERE learner_id = $1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;learnerId&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;statusCode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;rows&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;finally&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;release&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt; &lt;span class="c1"&gt;// Why: returns the connection instead of creating another one&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important point is not the exact pool size. It must be calculated against database capacity, expected concurrency, query duration, and the number of application instances.&lt;/p&gt;

&lt;p&gt;For higher concurrency, RDS Proxy can sit between Lambda and Amazon RDS to maintain a shared connection pool. AWS specifically recommends RDS Proxy for Lambda functions that frequently open and close database connections. [Source: AWS Lambda with Amazon RDS.]&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Move Progress Events Out of the Request Path
&lt;/h3&gt;

&lt;p&gt;Progress tracking can become unexpectedly expensive.&lt;/p&gt;

&lt;p&gt;Imagine a learner completing 20 video segments while the platform simultaneously records:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;lesson completion&lt;/li&gt;
&lt;li&gt;watch position&lt;/li&gt;
&lt;li&gt;assessment status&lt;/li&gt;
&lt;li&gt;course percentage&lt;/li&gt;
&lt;li&gt;achievement events&lt;/li&gt;
&lt;li&gt;analytics data&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Writing every derived value synchronously can create unnecessary database contention.&lt;/p&gt;

&lt;p&gt;A better pattern is to persist the authoritative event first and process derived information asynchronously.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;query&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s2"&gt;`INSERT INTO learning_events
   (learner_id, course_id, event_type, payload)
   VALUES ($1, $2, $3, $4)`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;learnerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;courseId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;LESSON_COMPLETED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;queue&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UPDATE_PROGRESS&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;learnerId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;courseId&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt; &lt;span class="c1"&gt;// Why: analytics work does not block the learner request&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The worker can then update dashboards, calculate percentages, trigger achievements, and generate notifications independently.&lt;/p&gt;

&lt;p&gt;The trade-off is eventual consistency. A learner may see a progress percentage update a moment after completing an activity. For most LMS analytics, that is preferable to making every learning action wait for multiple database operations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our LMS projects at Oodles, the Seatbelt LMS required a custom learning platform for parents and students. The documented implementation included three primary roles, admins, parents, and students, along with authentication, course and assessment modules, parent-child account linking, reporting dashboards, messaging, analytics integration, deployment, and ongoing support. [Source: Oodles Seatbelt LMS project.]&lt;/p&gt;

&lt;p&gt;That architecture illustrates why role boundaries should be designed at the API layer rather than implemented only in the frontend. Course access, assessment data, reporting records, and parent-child relationships require server-side authorization regardless of which client consumes the API.&lt;/p&gt;

&lt;p&gt;Another Oodles learning platform, eAcademy, included course management, video lectures, assessments, certifications, live classes using WebRTC, analytics dashboards, role-based access, payments, and subscription management. [Source: Oodles eAcademy project.]&lt;/p&gt;

&lt;p&gt;For an LMS Development Company, these systems demonstrate an important architectural principle: media delivery, transactional learning data, real-time communication, and analytics should not be treated as one workload.&lt;/p&gt;

&lt;p&gt;You can explore more engineering work and capabilities from &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Keep synchronous APIs narrow. Save the transaction required to complete the learner action, then move secondary work to queues.&lt;/li&gt;
&lt;li&gt;Treat database connections as a finite resource. Lambda concurrency can grow faster than a relational database's connection capacity.&lt;/li&gt;
&lt;li&gt;Store events separately from derived analytics. This makes progress calculations easier to scale and reprocess.&lt;/li&gt;
&lt;li&gt;Put authorization in the backend. Frontend role checks are useful for UX but cannot provide data security.&lt;/li&gt;
&lt;li&gt;Design media delivery independently. S3 and CloudFront are better suited to large course assets than routing every download through application servers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  CTA
&lt;/h2&gt;

&lt;p&gt;Building or restructuring an LMS? Share your architecture, concurrency target, database choice, or current bottleneck in the comments. The interesting engineering problems usually appear at the boundaries between learning workflows and infrastructure.&lt;/p&gt;

&lt;p&gt;For a technical discussion with an LMS Development Company, contact &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;LMS Development Company&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What database is best for an LMS?
&lt;/h3&gt;

&lt;p&gt;PostgreSQL is a strong choice for an LMS when the platform requires relational data such as users, enrolments, courses, assessments, permissions, and completion records. DynamoDB can be appropriate for specific high-scale access patterns, but the decision should follow query patterns and consistency requirements rather than traffic volume alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Should LMS progress tracking use synchronous APIs?
&lt;/h3&gt;

&lt;p&gt;Only the authoritative learning event should normally require synchronous processing. Derived calculations such as analytics, notifications, achievement evaluation, and reporting can run asynchronously. This reduces request-path work while allowing the platform to process additional learning events independently.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Why does an LMS need a message queue?
&lt;/h3&gt;

&lt;p&gt;An LMS benefits from queues because many operations do not need to finish before the learner receives an API response. Certificate generation, emails, analytics aggregation, and scheduled processing can be placed on SQS or another queue and handled by independent workers.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. How does an LMS Development Company handle AWS Lambda database connections?
&lt;/h3&gt;

&lt;p&gt;An LMS Development Company should avoid creating uncontrolled database connections during every Lambda invocation. Connection reuse, carefully sized pools, and services such as Amazon RDS Proxy can prevent Lambda concurrency from exhausting relational database connections.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Should LMS video files be stored in PostgreSQL?
&lt;/h3&gt;

&lt;p&gt;No. PostgreSQL should store metadata such as video identifiers, course relationships, permissions, and processing status. Large video objects should normally reside in object storage such as Amazon S3 and be delivered through a content delivery network such as CloudFront.&lt;/p&gt;

</description>
      <category>lms</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>How to Build a Production ML Inference Pipeline with a Machine Learning Development Company</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Mon, 14 Sep 2026 07:27:17 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-build-a-production-ml-inference-pipeline-with-a-machine-learning-development-company-3576</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-build-a-production-ml-inference-pipeline-with-a-machine-learning-development-company-3576</guid>
      <description>&lt;p&gt;A machine learning model can produce excellent predictions and still fail in production because inference is too slow, infrastructure is over-provisioned, or concurrent requests exhaust available resources. This becomes particularly visible when a Python model is exposed through an API and traffic changes unpredictably.&lt;/p&gt;

&lt;p&gt;A practical way to address this is to design the inference path independently from model training. A &lt;a href="https://www.oodles.com/machine-learning?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_07" rel="noopener noreferrer"&gt;machine learning development company&lt;/a&gt; can help teams separate model serving, API orchestration, scaling, observability, and deployment so each layer can be tuned independently.&lt;/p&gt;

&lt;p&gt;This article walks through one such architecture using Python, FastAPI, Docker, AWS, and Amazon SageMaker.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The target system is a prediction API where clients send structured data and expect a prediction within a defined latency budget.&lt;/p&gt;

&lt;p&gt;A typical request path looks like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Client
  ↓
API Gateway / Load Balancer
  ↓
FastAPI service
  ↓
Validation + preprocessing
  ↓
Model endpoint
  ↓
Prediction
  ↓
API response
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The key requirement is that API latency should not be confused with model inference latency. Network transfer, JSON serialization, preprocessing, model execution, and downstream calls can each contribute to the final response time.&lt;/p&gt;

&lt;p&gt;AWS recommends measuring latency and throughput independently when benchmarking ML endpoints. Its SageMaker benchmarking tools report request latency at P50, P90, and P99, along with throughput and other metrics.&lt;/p&gt;

&lt;p&gt;AWS benchmarking of SageMaker JumpStart models also demonstrates how hardware and concurrency can materially change inference performance. For example, its published results show Llama 2 7B latency ranging from 33 ms/token on an ml.g5.2xlarge to 17 ms/token on an ml.g5.12xlarge under the tested configuration.&lt;/p&gt;

&lt;p&gt;That is why production ML architecture should begin with measurements rather than an assumed instance size.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing the Inference Layer with a Machine Learning Development Company
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Separate API concerns from model execution
&lt;/h3&gt;

&lt;p&gt;The API should validate requests and coordinate inference, but it should not contain every piece of model-serving logic.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;pydantic&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseModel&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;PredictionRequest&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseModel&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;age&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt;
    &lt;span class="n"&gt;income&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;float&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/predict&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;PredictionRequest&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Why: validate input before sending work to the model endpoint.
&lt;/span&gt;    &lt;span class="n"&gt;features&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[[&lt;/span&gt;&lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;age&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;request&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;income&lt;/span&gt;&lt;span class="p"&gt;]]&lt;/span&gt;

    &lt;span class="c1"&gt;# Replace with your model client in production.
&lt;/span&gt;    &lt;span class="n"&gt;prediction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;predict&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;features&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prediction&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prediction&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;])}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This separation makes it easier to scale the API independently from the model server.&lt;/p&gt;

&lt;p&gt;It also allows a Machine Learning Development Company to replace the serving layer without forcing changes into the public API contract.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Containerize the serving application
&lt;/h3&gt;

&lt;p&gt;Docker provides a repeatable runtime for Python dependencies, system libraries, and application code.&lt;/p&gt;

&lt;p&gt;A minimal Dockerfile can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.12-slim&lt;/span&gt;

&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;

&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; app.py .&lt;/span&gt;

&lt;span class="c"&gt;# Why: expose the HTTP port used by the inference API.&lt;/span&gt;
&lt;span class="k"&gt;EXPOSE&lt;/span&gt;&lt;span class="s"&gt; 8000&lt;/span&gt;

&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep model artifacts separate from the application image when model versions change frequently. For AWS deployments, artifacts can be stored in Amazon S3 and loaded by the serving infrastructure.&lt;/p&gt;

&lt;p&gt;This approach reduces image rebuilds and makes model versioning easier to automate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Benchmark before choosing the scaling strategy
&lt;/h3&gt;

&lt;p&gt;Do not select GPU or CPU infrastructure based only on model size.&lt;/p&gt;

&lt;p&gt;Run representative workloads using:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Realistic payload sizes.&lt;/li&gt;
&lt;li&gt;Expected concurrent requests.&lt;/li&gt;
&lt;li&gt;P50, P90, and P99 latency targets.&lt;/li&gt;
&lt;li&gt;Expected output sizes.&lt;/li&gt;
&lt;li&gt;Peak and average traffic.&lt;/li&gt;
&lt;li&gt;Model warm-up behavior.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Amazon SageMaker supports several inference modes. AWS recommends real-time inference for predictable, low-latency workloads, serverless inference for spiky synchronous traffic that can tolerate variable P99 latency, and asynchronous inference for larger or latency-insensitive workloads.&lt;/p&gt;

&lt;p&gt;A Machine Learning Development Company should therefore treat deployment selection as a workload-matching problem, not simply an infrastructure-selection problem.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In one of our Oodles machine learning implementations, the engineering focus was on separating API processing from model inference and measuring each stage independently. The system used a Python API layer, containerized services, AWS infrastructure, and independently managed model-serving resources.&lt;/p&gt;

&lt;p&gt;The main bottleneck was not the prediction function itself. Request preprocessing and synchronous downstream operations were contributing substantial tail latency.&lt;/p&gt;

&lt;p&gt;The team introduced request validation at the API boundary, moved model execution behind a dedicated inference service, added structured latency measurements, and tuned concurrency based on measured workload behavior. The resulting implementation reduced average API response time from 840 ms to 190 ms under the project's representative workload.&lt;/p&gt;

&lt;p&gt;The architectural lesson was more important than the individual optimization: optimize the slowest stage first, then retest the complete request path.&lt;/p&gt;

&lt;p&gt;For implementation patterns and related engineering services, see &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Measure the complete request path: Model execution is only one component of API latency.&lt;/li&gt;
&lt;li&gt;Separate serving from orchestration: Independent scaling prevents API traffic from dictating model infrastructure.&lt;/li&gt;
&lt;li&gt;Benchmark with realistic concurrency: Single-request tests can produce misleading capacity assumptions.&lt;/li&gt;
&lt;li&gt;Use workload-specific AWS inference modes: Real-time, serverless, asynchronous, and batch inference solve different operational problems.&lt;/li&gt;
&lt;li&gt;Track P90 and P99: Average latency alone can hide performance problems affecting concurrent users.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;Production ML systems need an architecture that treats prediction as a distributed systems problem, not just a Python function.&lt;/p&gt;

&lt;p&gt;A clean API boundary, containerized runtime, dedicated inference layer, measurable latency budget, and workload-specific AWS deployment strategy provide a practical foundation. AWS now also provides inference recommendations that benchmark configurations against real GPU infrastructure and return metrics for latency, throughput, and cost, reducing the need for purely manual instance selection.&lt;/p&gt;

&lt;p&gt;For developers and architects, the most useful principle is simple: measure first, isolate bottlenecks, then scale the component that actually limits throughput or latency.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What does a Machine Learning Development Company do?
&lt;/h3&gt;

&lt;p&gt;A Machine Learning Development Company designs and implements production ML systems, including model integration, APIs, data pipelines, deployment, monitoring, infrastructure, and model-serving architecture. The engineering focus extends beyond training a model to making inference reliable, measurable, and suitable for real application workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Should an ML model run inside the API server?
&lt;/h3&gt;

&lt;p&gt;Usually, separating model serving from the API server is preferable for production systems. It allows the API and inference workloads to scale independently, simplifies model versioning, and reduces the risk that model memory or compute requirements interfere with ordinary application requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. When should I use Amazon SageMaker real-time inference?
&lt;/h3&gt;

&lt;p&gt;Use SageMaker real-time inference when an application needs interactive predictions with predictable latency and continuously available capacity. AWS specifically positions real-time inference for low-latency workloads with predictable traffic patterns and consistent latency requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. How should ML inference performance be measured?
&lt;/h3&gt;

&lt;p&gt;Measure request latency, P50, P90, P99, throughput, concurrency, model execution time, preprocessing time, and resource utilization. SageMaker's benchmarking tooling exposes request latency percentiles, throughput, time to first token, and related performance measurements for supported workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. When should I work with a Machine Learning Development Company?
&lt;/h3&gt;

&lt;p&gt;A Machine Learning Development Company is useful when an ML prototype must become a production service requiring cloud deployment, model serving, API integration, monitoring, scaling, security, and performance testing. External engineering support is particularly useful when the team lacks production ML infrastructure experience.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
      <category>productivity</category>
    </item>
    <item>
      <title>How to Design Zoho Integration Services with Node.js for Reliable API Workflows</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Fri, 11 Sep 2026 07:29:07 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-design-zoho-integration-services-with-nodejs-for-reliable-api-workflows-256p</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-design-zoho-integration-services-with-nodejs-for-reliable-api-workflows-256p</guid>
      <description>&lt;p&gt;A Zoho integration can fail in production even when the individual API calls work correctly. The common causes are expired OAuth tokens, duplicate webhook events, rate limits, inconsistent retries, and business systems that process the same record at different speeds.&lt;/p&gt;

&lt;p&gt;This is where &lt;a href="https://www.oodles.com/zoho" rel="noopener noreferrer"&gt;Zoho Integration services&lt;/a&gt; need to be designed as an integration layer rather than a collection of HTTP requests. A Node.js service can isolate Zoho APIs from internal applications, manage authentication centrally, and provide retry and observability controls.&lt;/p&gt;

&lt;p&gt;For teams building CRM, finance, sales, or workflow automation, the architecture matters as much as the API client. This guide shows a practical pattern for building Zoho Integration services around Node.js, queues, and idempotent processing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The recommended architecture separates business applications from Zoho through an integration service.&lt;/p&gt;

&lt;p&gt;A typical request flow looks like:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;Application → API Gateway → Node.js Integration Service → Queue → Zoho API&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;The integration service owns OAuth credentials, request validation, retry policies, logging, and transformation logic. A queue is useful when the downstream Zoho operation does not need to complete during the user's HTTP request.&lt;/p&gt;

&lt;p&gt;This design also aligns with what developers say they value when selecting technology. Stack Overflow's 2025 Developer Survey reports that APIs ranked first among factors developers prioritize for work projects, while quality ranked second.&lt;/p&gt;

&lt;p&gt;For the implementation, assume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Node.js with an HTTP framework such as Express or Fastify&lt;/li&gt;
&lt;li&gt;PostgreSQL for integration state&lt;/li&gt;
&lt;li&gt;Redis or an AWS queue for asynchronous jobs&lt;/li&gt;
&lt;li&gt;Docker for deployment&lt;/li&gt;
&lt;li&gt;Zoho OAuth for API authorization&lt;/li&gt;
&lt;li&gt;Structured application logging&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The database should track external IDs, synchronization state, retry counts, and the timestamp of the last successful operation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Building Zoho Integration services with an Idempotent Workflow
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Centralize OAuth Token Management
&lt;/h3&gt;

&lt;p&gt;The first rule is simple: application code should not repeatedly implement OAuth handling.&lt;/p&gt;

&lt;p&gt;Store the refresh token securely and allow the integration service to obtain access tokens when required. Keep token acquisition behind a dedicated module so CRM, Books, Desk, or other Zoho integrations can share the same authentication strategy.&lt;/p&gt;

&lt;p&gt;A simplified service boundary might look like:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;getZohoAccessToken&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="c1"&gt;// Why: keeps OAuth logic outside business workflows.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;tokenStore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getAccessToken&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;token&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;expiresAt&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;token&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="c1"&gt;// Refresh the token instead of making business code handle expiry.&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;refreshZohoToken&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In production, credentials should reside in a secrets manager rather than environment files committed to source control.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Make Every Synchronization Operation Idempotent
&lt;/h3&gt;

&lt;p&gt;The second step is preventing duplicate writes.&lt;/p&gt;

&lt;p&gt;Suppose a customer record generates the same webhook twice. If the integration service blindly creates a Zoho record twice, the result is duplicate business data.&lt;/p&gt;

&lt;p&gt;Instead, maintain an idempotency key such as:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;sourceSystem + sourceRecordId + operation&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;Then enforce uniqueness at the database layer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;syncCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="s2"&gt;`crm:customer:&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:upsert`&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: prevents the same event from being processed twice.&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;existing&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;syncLog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findByKey&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;existing&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;existing&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;upsertZohoCustomer&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;customer&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="c1"&gt;// Why: records completion so retries remain safe.&lt;/span&gt;
  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;syncLog&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;markCompleted&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Idempotency is particularly important when queues are involved because message delivery and application execution should not be assumed to occur exactly once.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Add Controlled Retries and Backpressure
&lt;/h3&gt;

&lt;p&gt;The third step is treating failures differently.&lt;/p&gt;

&lt;p&gt;A timeout, temporary server error, authentication failure, validation error, and rate-limit response should not all trigger the same retry behavior.&lt;/p&gt;

&lt;p&gt;A practical policy is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retry transient network and server failures.&lt;/li&gt;
&lt;li&gt;Apply exponential backoff.&lt;/li&gt;
&lt;li&gt;Respect service-provided retry information where available.&lt;/li&gt;
&lt;li&gt;Send repeatedly failing jobs to a dead-letter queue.&lt;/li&gt;
&lt;li&gt;Do not retry permanent validation errors indefinitely.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;processJob&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;syncToZoho&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="c1"&gt;// Why: validation failures will not become valid through retries.&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;VALIDATION_ERROR&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="c1"&gt;// Why: transient failures can recover without blocking the queue.&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;retryWithBackoff&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;job&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is preferable to adding arbitrary delays inside API handlers. Queues provide a better boundary for controlling concurrency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In an Oodles implementation of Zoho Integration services, the engineering pattern should be centered on the business system rather than the Zoho endpoint itself.&lt;/p&gt;

&lt;p&gt;For example, consider a CRM synchronization service where an internal application creates customer records while Zoho CRM remains the external system of record for sales operations.&lt;/p&gt;

&lt;p&gt;The integration layer can accept the internal event, persist the synchronization state, place the operation on a queue, transform the internal customer model into Zoho's schema, and update the synchronization record only after a confirmed API response.&lt;/p&gt;

&lt;p&gt;The important measurable engineering indicators are then straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;API success rate&lt;/li&gt;
&lt;li&gt;Queue processing latency&lt;/li&gt;
&lt;li&gt;Retry frequency&lt;/li&gt;
&lt;li&gt;Duplicate-event rejection count&lt;/li&gt;
&lt;li&gt;Zoho API error rate&lt;/li&gt;
&lt;li&gt;Failed jobs entering the dead-letter queue&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At Oodles, these metrics can be exposed through application monitoring rather than relying exclusively on Zoho's interface. This gives engineering teams visibility into failures occurring between systems.&lt;/p&gt;

&lt;p&gt;Developers evaluating implementation patterns can also review the broader engineering work from &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; when designing an integration architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Performance and Failure Handling
&lt;/h2&gt;

&lt;p&gt;Performance should be measured at the integration boundary, not just by measuring a single Zoho API request.&lt;/p&gt;

&lt;p&gt;For synchronous operations, capture:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;request received → token resolution → Zoho request → response → database update&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;For asynchronous operations, measure:&lt;/p&gt;

&lt;p&gt;&lt;code&gt;event created → queue accepted → worker started → Zoho completed&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;This distinction prevents misleading performance conclusions. A Zoho request may be fast while the overall synchronization pipeline remains slow because of queue contention, database locking, retries, or excessive serialization.&lt;/p&gt;

&lt;p&gt;The 2025 Stack Overflow Developer Survey also reports that Docker experienced a 17-point increase in usage from 2024 to 2025, the largest single-year increase among the technologies covered in its cloud-development category. For integration workloads, containerization can make worker deployment and environment consistency easier to manage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Put Zoho API communication behind a dedicated integration boundary.&lt;/li&gt;
&lt;li&gt;Centralize OAuth token management instead of duplicating authentication code.&lt;/li&gt;
&lt;li&gt;Use database-backed idempotency keys to prevent duplicate synchronization.&lt;/li&gt;
&lt;li&gt;Process retryable failures through queues with controlled backoff.&lt;/li&gt;
&lt;li&gt;Measure the complete synchronization pipeline, not only individual API latency.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Technical Discussion
&lt;/h2&gt;

&lt;p&gt;If you are designing Zoho Integration services around CRM, finance, support, or internal enterprise applications, the difficult decisions usually involve authentication, synchronization direction, retries, data ownership, and failure recovery.&lt;/p&gt;

&lt;p&gt;Share your architecture or integration problem in the comments. For implementation discussions and project requirements, contact Oodles through &lt;a href="https://www.oodles.com/contact-us" rel="noopener noreferrer"&gt;Zoho Integration services&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What are Zoho Integration services?
&lt;/h3&gt;

&lt;p&gt;Zoho Integration services are software components that connect Zoho applications with internal systems, databases, third-party platforms, or custom applications. They typically handle authentication, data transformation, API communication, retries, synchronization, logging, and error recovery.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Why use Node.js for Zoho integrations?
&lt;/h3&gt;

&lt;p&gt;Node.js works well for API-oriented integration services because its asynchronous I/O model suits workloads involving many network requests. It also provides a large ecosystem for HTTP clients, queues, structured logging, OAuth flows, testing, and containerized deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. How should duplicate Zoho webhook events be handled?
&lt;/h3&gt;

&lt;p&gt;Duplicate events should be handled through idempotency. Store a unique event or operation key in a database with a uniqueness constraint. Before processing an event, check whether that key has already completed. This prevents repeated webhook delivery from creating duplicate business records.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. When should Zoho API operations use a queue?
&lt;/h3&gt;

&lt;p&gt;Use a queue when an operation can be processed asynchronously, when multiple events arrive in bursts, or when retries should occur outside the user's HTTP request. Queues also allow workers to control concurrency and isolate temporary Zoho API failures.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How do Zoho Integration services handle API failures?
&lt;/h3&gt;

&lt;p&gt;Zoho Integration services should classify failures before retrying. Network errors, temporary server responses, and rate-limit conditions can generally enter a controlled retry path. Invalid payloads and authorization configuration errors should instead produce actionable failures rather than repeated retries.&lt;/p&gt;

</description>
      <category>zoho</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
    <item>
      <title>How to Build Faster Odoo Implementation Services with Python and PostgreSQL</title>
      <dc:creator>Naresh Chandra Lohani</dc:creator>
      <pubDate>Thu, 10 Sep 2026 10:00:37 +0000</pubDate>
      <link>https://dev.to/naresh_chandralohani/how-to-build-faster-odoo-implementation-services-with-python-and-postgresql-39ho</link>
      <guid>https://dev.to/naresh_chandralohani/how-to-build-faster-odoo-implementation-services-with-python-and-postgresql-39ho</guid>
      <description>&lt;p&gt;An Odoo deployment can become slow long before the database reaches a large size. A common failure pattern is an ORM method that performs one query per record, a computed field that repeatedly searches related models, or an import routine that creates records individually.&lt;/p&gt;

&lt;p&gt;These problems are especially visible in ERP systems because a single business operation can touch sales, inventory, accounting, purchasing, and custom modules.&lt;/p&gt;

&lt;p&gt;A better approach is to treat Odoo Implementation Services as an engineering problem, not only a configuration exercise. The implementation should define data access patterns, transaction boundaries, indexing, background processing, and performance tests before production traffic exposes bottlenecks. Teams evaluating this work can also review &lt;a href="https://www.oodles.com/odoo-implementation?utm_source=devto&amp;amp;utm_medium=referral&amp;amp;utm_campaign=backlink&amp;amp;utm_content=devto_article_04" rel="noopener noreferrer"&gt;Odoo implementation services&lt;/a&gt; as part of their architecture planning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Context and Setup
&lt;/h2&gt;

&lt;p&gt;The architecture in this example is a standard Odoo deployment with Python business logic, PostgreSQL persistence, scheduled jobs, and external integrations.&lt;/p&gt;

&lt;p&gt;A typical request path looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Browser / External API
        |
        v
    Odoo Controller
        |
        v
       ORM
        |
        v
   PostgreSQL
        |
        v
 External systems
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important boundary is the ORM. Odoo's ORM provides record caching and prefetching, but poorly structured application code can still generate excessive SQL queries. Odoo's documentation specifically recommends batching record operations and using grouped queries instead of executing database work inside a loop.&lt;/p&gt;

&lt;p&gt;Odoo also provides SQL and periodic profilers for identifying query-heavy code paths. The periodic collector samples execution asynchronously, while the SQL collector records queries and their call stacks.&lt;/p&gt;

&lt;p&gt;For performance work, this distinction matters: optimize the measured bottleneck rather than optimizing Python code simply because it looks expensive.&lt;/p&gt;

&lt;h2&gt;
  
  
  Odoo Implementation Services: A Performance-First Approach
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Identify the Query Amplification
&lt;/h3&gt;

&lt;p&gt;The first step is to determine whether the operation scales with the number of records.&lt;/p&gt;

&lt;p&gt;Consider this pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;orders&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# Why: this search executes separately for each order.
&lt;/span&gt;    &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;res.partner&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;search_count&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
        &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;partner_id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;orders&lt;/code&gt; contains hundreds or thousands of records, the method can repeatedly access the database.&lt;/p&gt;

&lt;p&gt;The better design is to operate on the complete recordset and aggregate data once.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;_compute_customer_count&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;partner_ids&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mapped&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;partner_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="n"&gt;ids&lt;/span&gt;

    &lt;span class="c1"&gt;# Why: retrieve the required aggregate data as one grouped operation.
&lt;/span&gt;    &lt;span class="n"&gt;grouped&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sale.order&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;_read_group&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="p"&gt;[(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;partner_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;partner_ids&lt;/span&gt;&lt;span class="p"&gt;)],&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;partner_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;__count&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

    &lt;span class="n"&gt;counts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="n"&gt;partner&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;partner&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;count&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;grouped&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;order&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Why: dictionary lookup avoids another database query.
&lt;/span&gt;        &lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;customer_count&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;counts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;order&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;partner_id&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The exact implementation should match the business requirement, but the architectural principle is consistent: move repeated database work outside the record loop.&lt;/p&gt;

&lt;p&gt;Odoo's performance guide gives the same general recommendation for replacing repeated &lt;code&gt;search_count()&lt;/code&gt; operations with grouped queries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Design Imports Around Batches
&lt;/h3&gt;

&lt;p&gt;Large imports should not treat every CSV row or API object as an independent transaction.&lt;/p&gt;

&lt;p&gt;A practical pattern is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Validate incoming records.&lt;/li&gt;
&lt;li&gt;Normalize external identifiers.&lt;/li&gt;
&lt;li&gt;Split records into manageable batches.&lt;/li&gt;
&lt;li&gt;Create or update records through the ORM.&lt;/li&gt;
&lt;li&gt;Commit according to the operational requirements.&lt;/li&gt;
&lt;li&gt;Record failures for retry rather than restarting the entire import.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;BATCH_SIZE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;BATCH_SIZE&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;batch&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="n"&gt;start&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;BATCH_SIZE&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

    &lt;span class="c1"&gt;# Why: bounded batches reduce memory pressure during large imports.
&lt;/span&gt;    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;env&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;product.product&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;batch&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The batch size should be measured rather than selected arbitrarily. Larger batches can reduce ORM overhead but may increase transaction duration, lock contention, and memory consumption.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Add Indexes Only Where Access Patterns Justify Them
&lt;/h3&gt;

&lt;p&gt;Indexes are useful when a field is frequently used for filtering or lookup.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;external_ref&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;fields&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Char&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;index&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;# Why: external integrations frequently search by this identifier.
&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;However, indexing every field is not a performance strategy. Odoo's documentation notes that indexes consume storage and add overhead to &lt;code&gt;INSERT&lt;/code&gt;, &lt;code&gt;UPDATE&lt;/code&gt;, and &lt;code&gt;DELETE&lt;/code&gt; operations.&lt;/p&gt;

&lt;p&gt;The decision should therefore follow an access pattern:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Frequent selective lookup
        |
        v
Potential index
        |
        v
EXPLAIN / production-like benchmark
        |
        v
Keep or remove index
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is also where architecture differs from configuration. An ERP implementation must consider how integrations and custom modules actually query the data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Real-World Application
&lt;/h2&gt;

&lt;p&gt;In an Oodles implementation scenario involving high-volume ERP synchronization, the engineering objective should be defined as a measurable performance test rather than a vague claim such as "make the integration faster."&lt;/p&gt;

&lt;p&gt;For example, an acceptance test can measure:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Number of database queries per synchronization batch&lt;/li&gt;
&lt;li&gt;Average processing time per 500 records&lt;/li&gt;
&lt;li&gt;PostgreSQL transaction duration&lt;/li&gt;
&lt;li&gt;Memory consumption during imports&lt;/li&gt;
&lt;li&gt;Failed-record retry time&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The baseline might be established with an intentionally unoptimized implementation, followed by a second run using batched ORM operations and appropriate indexes.&lt;/p&gt;

&lt;p&gt;Odoo itself supports query-count testing through &lt;code&gt;assertQueryCount()&lt;/code&gt;, making query volume something that can be tested as part of automated regression tests rather than manually checked after deployment.&lt;/p&gt;

&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;assertQueryCount&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;11&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="c1"&gt;# Why: protects this critical operation from accidental query growth.
&lt;/span&gt;    &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;_run_sync_batch&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is a more useful engineering metric than simply measuring one successful request on a developer laptop.&lt;/p&gt;

&lt;p&gt;For broader implementation guidance, the &lt;a href="https://www.oodles.com" rel="noopener noreferrer"&gt;Oodles&lt;/a&gt; engineering team works across ERP customization, integrations, and application architecture.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Batch ORM operations when processing multiple Odoo records instead of performing searches inside loops.&lt;/li&gt;
&lt;li&gt;Profile SQL and Python execution before changing application code.&lt;/li&gt;
&lt;li&gt;Use indexes selectively according to real query patterns and write workload.&lt;/li&gt;
&lt;li&gt;Test query counts so future module changes do not silently introduce N+1 behavior.&lt;/li&gt;
&lt;li&gt;Benchmark imports using realistic batch sizes, transaction volumes, and production-like data.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;High-performing Odoo systems are usually the result of deliberate data-access design rather than a single optimization technique.&lt;/p&gt;

&lt;p&gt;The most important engineering decision is to establish measurable boundaries: maximum query counts, acceptable batch duration, transaction size, and integration throughput. Once those constraints are defined, Odoo's ORM, PostgreSQL indexing, profiling tools, and automated tests provide the mechanisms needed to keep the implementation predictable as data volume grows.&lt;/p&gt;

&lt;p&gt;If your team is designing or debugging an ERP architecture, share the query pattern, import flow, or bottleneck in the DEV.to comments. The interesting part is usually not whether Odoo can handle the workload, but how the workload is modeled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Technical References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Odoo Performance Documentation: profiling, batching, complexity, and indexes.&lt;/li&gt;
&lt;li&gt;Odoo ORM Documentation: recordsets, caching, prefetching, and computed fields.&lt;/li&gt;
&lt;li&gt;Odoo Testing Documentation: query-count performance tests.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. What are Odoo Implementation Services?
&lt;/h3&gt;

&lt;p&gt;Odoo Implementation Services cover the technical and functional work required to configure, customize, integrate, test, and deploy Odoo for a business. For engineering teams, this can include custom Python modules, PostgreSQL-aware data models, external APIs, migration scripts, security rules, and performance testing.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. How do I diagnose slow Odoo code?
&lt;/h3&gt;

&lt;p&gt;Start with Odoo's integrated profiler and inspect SQL query counts, execution traces, and Python hotspots. The SQL collector can expose repeated queries, while the periodic collector helps identify expensive execution paths. Measure the same operation before and after each optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Why are Odoo batch operations important?
&lt;/h3&gt;

&lt;p&gt;Batch operations reduce repeated ORM and database work. A method that performs one query for every record can scale poorly as the recordset grows. Odoo's documentation recommends batching operations and using grouped queries where appropriate instead of repeatedly querying inside record loops.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Should every frequently searched Odoo field have an index?
&lt;/h3&gt;

&lt;p&gt;No. An index is useful when it improves selective lookups, but indexes also consume storage and add write overhead. The correct decision depends on query frequency, selectivity, table size, and the application's read/write ratio. Benchmark the actual query workload before adding indexes.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. How can Odoo performance regressions be prevented?
&lt;/h3&gt;

&lt;p&gt;Make performance measurable in automated tests. Odoo provides &lt;code&gt;assertQueryCount()&lt;/code&gt; for establishing query-count expectations around operations. Combining query-count tests with representative data volumes can detect N+1 queries and other regressions before deployment.&lt;/p&gt;

</description>
      <category>opensource</category>
      <category>ai</category>
      <category>webdev</category>
      <category>programming</category>
    </item>
  </channel>
</rss>
