<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Eiji</title>
    <description>The latest articles on DEV Community by Eiji (@eiji-kudo).</description>
    <link>https://dev.to/eiji-kudo</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F1232413%2F4322d407-df42-4fb2-ba89-597b7f7e240f.jpeg</url>
      <title>DEV Community: Eiji</title>
      <link>https://dev.to/eiji-kudo</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/eiji-kudo"/>
    <language>en</language>
    <item>
      <title>How I Led Capacity Planning for an Approximately 1,000x E-Commerce Traffic Spike</title>
      <dc:creator>Eiji</dc:creator>
      <pubDate>Sat, 03 Oct 2026 16:37:54 +0000</pubDate>
      <link>https://dev.to/eiji-kudo/how-i-led-capacity-planning-for-an-approximately-1000x-e-commerce-traffic-spike-32g5</link>
      <guid>https://dev.to/eiji-kudo/how-i-led-capacity-planning-for-an-approximately-1000x-e-commerce-traffic-spike-32g5</guid>
      <description>&lt;p&gt;I am the tech lead for an influencer e-commerce team. We were preparing for a traffic event expected to receive approximately 1,000× normal traffic. A previous event had saturated the primary database, so we could not treat this as a routine scaling exercise.&lt;/p&gt;

&lt;p&gt;The difficult part was not setting up k6 or listing bottlenecks. It was deciding what to test locally and in staging, estimating the infrastructure cost before scaling the environment, and turning the findings into a remediation plan that multiple teams could execute. The work also required tracing a request from the frontend through the backend and into the database and infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Split the tests by cost and fidelity
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cjw0g15ksq3igov5m9g.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F3cjw0g15ksq3igov5m9g.png" alt="Three-stage capacity testing strategy: local, staging, and traffic-event validation" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Scaling staging close to production capacity costs money for every hour it runs. I first used a LocalStack-based local environment to evaluate queries, indexes, transaction boundaries, and single-task throughput.&lt;/p&gt;

&lt;p&gt;We ran the actual application image with MySQL and replaced external APIs with deterministic stubs. Instead of copying production data, we generated synthetic data with similar table sizes and index cardinality. For single-task tests, Docker CPU and memory limits matched production.&lt;/p&gt;

&lt;p&gt;I separated single-endpoint tests from mixed-flow tests based on traffic patterns observed during the previous incident. A simplified mixed-flow scenario looked like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;options&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;scenarios&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;flows&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;executor&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;constant-arrival-rate&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;rate&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;RATE&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
      &lt;span class="na"&gt;timeUnit&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;1s&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;DURATION&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="k"&gt;default&lt;/span&gt; &lt;span class="nf"&gt;function &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;pickByWeight&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;READ_WEIGHT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;readEndpoint&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;WRITE_WEIGHT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;writeEndpoint&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;CALLBACK_WEIGHT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;asyncCallback&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;__ENV&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;AUTH_WEIGHT&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;authEndpoint&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="p"&gt;])();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;MySQL ran with &lt;code&gt;long_query_time=0&lt;/code&gt;, and the slow query log was cleared before each run. I compared query counts, &lt;code&gt;rows examined&lt;/code&gt;, execution plans, and connection hold time—not only response latency.&lt;/p&gt;

&lt;p&gt;This exposed missing indexes, duplicate queries, unnecessary data fetching, and long transactions. In one case, the expensive query was not the SQL under investigation but a full scan triggered by an ORM relationship.&lt;/p&gt;

&lt;h2&gt;
  
  
  Agree on the staging cost before running the test
&lt;/h2&gt;

&lt;p&gt;The staging test required temporarily scaling the database, container workloads, cache, load generators, and an external-service stub. I estimated the hourly cost, expected duration, and spending cap before provisioning them, then agreed on the plan with the engineers responsible for the environment.&lt;/p&gt;

&lt;p&gt;Because the test occupied staging, I also coordinated a window that would not block development or QA. Longer runs were scheduled outside normal working hours, with the start time, affected features, and rollback plan shared in advance.&lt;/p&gt;

&lt;p&gt;This preparation mattered as much as the test script. It made the cost explicit and allowed the team to repeat tests when a fix needed verification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reproduce the complete request path
&lt;/h2&gt;

&lt;p&gt;Each k6 virtual user represented one isolated client with its own token and test data. We modelled two traffic groups:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authenticated clients repeatedly exercising the read path before the traffic event&lt;/li&gt;
&lt;li&gt;New clients arriving at the start, authenticating, and entering the write path&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The simplified flow was:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;sustainedTraffic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;pollReadPathUntilStart&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;executeWritePath&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;newSessionTraffic&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;authenticate&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;executeWritePath&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;executeWritePath&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;createResource&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;enterAdmissionControl&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;pollUntilAdmitted&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="nf"&gt;submitWrite&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The scenario covered read-heavy requests, session creation, admission control, state-changing requests, and asynchronous callbacks. The external-service stub also emitted callbacks with realistic retry intervals.&lt;/p&gt;

&lt;p&gt;An HTTP 200 was not enough to call a run successful. We reconciled k6 application errors, callback state, and persisted records, including domain-invariant checks.&lt;/p&gt;

&lt;p&gt;k6 ran on dedicated EC2 instances in the staging environment. The instances and related IAM and network resources were provisioned with Terraform. We selected their size from the target concurrency and iteration duration, adjusted operating-system limits, and saved the dashboard and summary from every run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fix the request path across frontend and backend
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr1he5njf2jku8pct148d.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fr1he5njf2jku8pct148d.png" alt="Before-and-after connection-pool flow with the primary database on the left and reader database on the right" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The largest issue was in the write-path transaction. It held a primary connection while waiting for a reader connection, exhausting the database connection pool even when the database itself still had capacity. Adding more application tasks gave little improvement. Moving the read before the transaction removed most of the errors under the same load.&lt;/p&gt;

&lt;p&gt;Other improvements crossed the frontend/backend boundary. We traced each client action through browser requests, API calls, SQL, and infrastructure metrics. We removed redundant requests from the read-heavy client path and reduced duplicate backend calls.&lt;/p&gt;

&lt;p&gt;We also avoided direct database reads where short-lived data did not require them. Admission-control state was cached briefly in Valkey, while signed URLs were reused from an in-memory cache inside each ECS task.&lt;/p&gt;

&lt;p&gt;Finally, I increased ECS tasks and EKS pods and nodes step by step to measure the throughput of each configuration. From the database connection limit, I created a connection budget that included rolling deployments and scheduled jobs. This established the safe horizontal-scaling ceiling and the prescale configuration for the event.&lt;/p&gt;

&lt;h2&gt;
  
  
  Turn findings into team-owned work
&lt;/h2&gt;

&lt;p&gt;I worked as an individual contributor on the transaction, query, cache, and request-flow changes. But the scope was larger than one engineer or one codebase.&lt;/p&gt;

&lt;p&gt;I grouped the remaining findings by request path and failure mode, then created remediation issues containing the test evidence, proposed change, and expected impact. I shared the plan with the frontend, backend, and infrastructure owners and led prioritisation across the request path.&lt;/p&gt;

&lt;p&gt;During each test, I also showed the engineers how client behaviour translated into request volume, admission-control progress, and error conditions. This let us discuss application behaviour and technical limits using the same evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Result
&lt;/h2&gt;

&lt;p&gt;The traffic event completed with a low error rate and stable request latency, without the database saturation seen previously.&lt;/p&gt;

&lt;p&gt;The most useful output was not a single maximum request number. It was a shared plan that separated problems requiring code changes from capacity that should be reserved through prescaling—and assigned each change to an owner.&lt;/p&gt;

&lt;p&gt;Delivering that plan required working from frontend behaviour through backend transactions to infrastructure capacity, implementing critical fixes myself, and leading the remaining work across teams.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open to opportunities
&lt;/h2&gt;

&lt;p&gt;I am currently looking for software engineering opportunities in Canada and the UK. I am authorised to work in both Canada and the UK.&lt;/p&gt;

&lt;p&gt;If you are hiring for full-stack engineering, cloud infrastructure, SRE, performance engineering, or developer productivity, please get in touch through &lt;a href="https://eiji-kudo.github.io/portfolio/?lang=en" rel="noopener noreferrer"&gt;my portfolio&lt;/a&gt; or &lt;a href="https://x.com/tech_eassy" rel="noopener noreferrer"&gt;X&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>performance</category>
      <category>devops</category>
      <category>architecture</category>
      <category>aws</category>
    </item>
  </channel>
</rss>
