<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Convobrains</title>
    <description>The latest articles on DEV Community by Convobrains (@convobrains).</description>
    <link>https://dev.to/convobrains</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4158456%2F9248b274-1378-4b88-a035-94fde0391f4c.png</url>
      <title>DEV Community: Convobrains</title>
      <link>https://dev.to/convobrains</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/convobrains"/>
    <language>en</language>
    <item>
      <title>Our Dashboard Looked Fine. The Database Was Out of Connections.</title>
      <dc:creator>Convobrains</dc:creator>
      <pubDate>Fri, 02 Oct 2026 20:51:59 +0000</pubDate>
      <link>https://dev.to/convobrains/our-dashboard-looked-fine-the-database-was-out-of-connections-36lf</link>
      <guid>https://dev.to/convobrains/our-dashboard-looked-fine-the-database-was-out-of-connections-36lf</guid>
      <description>&lt;p&gt;At 50 conversations a week, one database feels like enough for everyone. At thousands, every service that touches it becomes a tenant fighting for the same front door.&lt;/p&gt;

&lt;p&gt;We learned this the hard way. The dashboards broke first. The CPU graph looked fine. That mismatch sent us down the wrong path for longer than it should have.&lt;/p&gt;

&lt;h2&gt;
  
  
  How the system behaved at low volume
&lt;/h2&gt;

&lt;p&gt;We run a conversation intelligence stack. Audio comes in. It gets transcribed. It gets audited against a scorecard. Results land in dashboards for team leads and QA managers.&lt;/p&gt;

&lt;p&gt;For a long time, this worked on a shared Postgres instance. Multiple services connected to the same database. Each one opened a connection pool so it would be ready when traffic arrived.&lt;/p&gt;

&lt;p&gt;At modest volume, this felt efficient. Pools sat mostly idle. Queries returned fast. Dashboards loaded. Ingest kept up. CloudWatch showed CPU in single digits. We had headroom, or so we thought.&lt;/p&gt;

&lt;p&gt;Nobody was doing anything obviously wrong. This is a normal pattern for a growing B2B product. One database, many services, everyone holding a few connections just in case.&lt;/p&gt;

&lt;h2&gt;
  
  
  What volume did to it
&lt;/h2&gt;

&lt;p&gt;Then conversation volume grew. The preprocessing pipeline scaled out to handle more audio. More worker tasks came online. Each task brought its own connection pool.&lt;/p&gt;

&lt;p&gt;The audit pipeline ran in parallel. The web app served more dashboard users. Staging and dev environments shared the same database instance as production.&lt;/p&gt;

&lt;p&gt;The database hit a hard connection ceiling. Not slow queries. Not timeouts on heavy reports. New connections were refused entirely.&lt;/p&gt;

&lt;p&gt;Auditor dashboards started throwing server errors. Transcript fetches failed. Ingest jobs retried and piled up. Our monitoring tools could not query the database either.&lt;/p&gt;

&lt;p&gt;The login page still loaded. That made the failure feel localized, like a bug in one feature. It was not.&lt;/p&gt;

&lt;h2&gt;
  
  
  How long before we knew something was wrong
&lt;/h2&gt;

&lt;p&gt;Longer than I want to admit.&lt;/p&gt;

&lt;p&gt;The first signal was dashboard crashes on auditor and team-lead pages. We investigated application errors. Server component failures. Internal server errors on transcript fetch.&lt;/p&gt;

&lt;p&gt;The database was pegged at its connection limit for hours. CPU stayed low. Memory was tight but not the failure mode.&lt;/p&gt;

&lt;p&gt;The click for me was when read-only monitoring queries failed with the same connection error. This was not a dashboard code problem. Every service sharing that database was competing for the same finite slot budget.&lt;/p&gt;

&lt;p&gt;We had scaled compute to handle more conversations. We had not scaled the connection math.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we learned about building systems that survive growth
&lt;/h2&gt;

&lt;p&gt;The problem was pool stacking.&lt;/p&gt;

&lt;p&gt;Every service reserves connections whether it uses them or not. Preprocessing workers hold pools open while waiting for audio. Web containers hold pools open between requests. Audit workers hold pools open between jobs. Staging holds pools open while nobody is testing.&lt;/p&gt;

&lt;p&gt;Add more workers to handle volume, and you add more idle reservations. The database becomes a parking lot where every spot stays occupied even when nobody is using it.&lt;/p&gt;

&lt;p&gt;A bigger database instance does not fix this by itself. Idle pools expand to fill a larger limit. You pay more and hit the same wall later.&lt;/p&gt;

&lt;p&gt;What actually helps:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Connection budgets per service.&lt;/strong&gt; New components now ship with explicit small pool limits. Not "use whatever the driver defaults to." A hard cap negotiated against the shared budget.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Treat connections as a shared resource, not a per-service perk.&lt;/strong&gt; Before you scale workers horizontally, multiply pool size by task count. If the product exceeds what the database allows, you will fail before CPU does.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Separate production from non-production data stores.&lt;/strong&gt; Dev and staging pools competing with production ingest is an expensive habit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A pooling layer in front of Postgres.&lt;/strong&gt; Transaction pooling lets many application connections share fewer database slots. We had been deferring this. Volume stopped letting us defer it.&lt;/p&gt;

&lt;p&gt;The lesson is not "Postgres breaks at scale." The lesson is that connection slots are a hidden capacity limit that only shows up when compute scales faster than connection planning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why every founder scaling a data-heavy product will face a version of this
&lt;/h2&gt;

&lt;p&gt;If your product touches customer conversations, you are data-heavy by definition. Ingest, storage, analysis, dashboards. Multiple services. Multiple pools. One database.&lt;/p&gt;

&lt;p&gt;This pattern shows up everywhere we look. A BPO platform running QA on thousands of agent calls. A CRM adding conversation intelligence for every deal. An EdTech company scoring counseling calls at batch scale. A workforce management tool surfacing quality metrics to supervisors.&lt;/p&gt;

&lt;p&gt;The failure mode looks the same. Dashboards break. Ingest stalls. The team debugs application code while the real bottleneck sits one layer down.&lt;/p&gt;

&lt;p&gt;You do not need to be running millions of calls to hit this. You need enough parallel workers that their idle pools exceed what your database was sized to hold.&lt;/p&gt;




&lt;p&gt;List every service that connects to your production database and its max pool size. Multiply pool size by running instances across prod, staging, and workers. Does the total exceed your database connection limit? If you have never done that math, your next traffic spike will do it for you.&lt;/p&gt;

&lt;p&gt;At ConvoBrains we run conversation intelligence infrastructure at scale for B2B companies. The technical failures we hit, we hit in production with real customer data on the line. The lessons above came from those moments. Understanding how systems break at volume is not optional when your product is built on top of that volume.&lt;/p&gt;

&lt;p&gt;If you are processing customer conversations at scale and want to know what is actually inside them, we will run a free conversation audit. Send us a sample. We will show you what your current process is missing. &lt;a href="https://www.convobrains.com/onboarding?utm_source=devto&amp;amp;utm_medium=organic&amp;amp;utm_campaign=editorial&amp;amp;utm_content=dashboard-database-connection-exhaustion-scale" rel="noopener noreferrer"&gt;Start here&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;If you are building in sales enablement, contact center, CRM, EdTech, QSR, or workforce management and your customers generate audio at scale, we should talk about what a conversation intelligence layer could do for your product. &lt;a href="https://www.convobrains.com/contact?utm_source=devto&amp;amp;utm_medium=organic&amp;amp;utm_campaign=editorial&amp;amp;utm_content=dashboard-database-connection-exhaustion-scale-partner" rel="noopener noreferrer"&gt;Get in touch&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Discussion question:&lt;/strong&gt; If you doubled your conversation volume tomorrow, which part of your stack would break first, and do you have a graph that would tell you before your users do?&lt;/p&gt;

</description>
      <category>architecture</category>
      <category>database</category>
      <category>monitoring</category>
      <category>postgres</category>
    </item>
  </channel>
</rss>
