<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mitya Zima</title>
    <description>The latest articles on DEV Community by Mitya Zima (@mitya_zima_6b84bbc9a16bfc).</description>
    <link>https://dev.to/mitya_zima_6b84bbc9a16bfc</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4147614%2F3ec1d997-5a6d-41e9-a4ba-8622b2b95213.jpg</url>
      <title>DEV Community: Mitya Zima</title>
      <link>https://dev.to/mitya_zima_6b84bbc9a16bfc</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/mitya_zima_6b84bbc9a16bfc"/>
    <language>en</language>
    <item>
      <title>Why your "zero-downtime" Docker Compose deploy still drops requests</title>
      <dc:creator>Mitya Zima</dc:creator>
      <pubDate>Mon, 28 Sep 2026 16:47:35 +0000</pubDate>
      <link>https://dev.to/mitya_zima_6b84bbc9a16bfc/why-your-zero-downtime-docker-compose-deploy-still-drops-requests-35k4</link>
      <guid>https://dev.to/mitya_zima_6b84bbc9a16bfc/why-your-zero-downtime-docker-compose-deploy-still-drops-requests-35k4</guid>
      <description>&lt;p&gt;For a while I was pretty happy with my deploy setup. Push to main, GitHub Actions builds an image, SSHes into the VPS, pulls, restarts. Simple, no Kubernetes, nothing to babysit. Then I actually measured what happened to in-flight requests during a deploy instead of just eyeballing it, and it wasn't zero.&lt;/p&gt;

&lt;p&gt;The stack is nothing exotic. Docker Compose, Traefik in front, one VPS, no orchestrator, because for a handful of services on one box an orchestrator is more infrastructure than the problem justifies. What I wanted was for &lt;code&gt;docker compose up -d&lt;/code&gt; to stop yanking the old container out from under active connections.&lt;/p&gt;

&lt;p&gt;docker-rollout (&lt;a href="https://github.com/wowu/docker-rollout" rel="noopener noreferrer"&gt;https://github.com/wowu/docker-rollout&lt;/a&gt;) does the clever part: starts the new container next to the old one, waits for it to report healthy, points the proxy at it, only then removes the old one. Good piece of engineering, and the reason I didn't have to write that logic myself.&lt;/p&gt;

&lt;p&gt;Wiring it into CI is where it got less clean. Upload compose files over SSH, log into a registry without leaving creds lying around, run migrations before the switch, roll out several services in the right order. None of it is individually hard, it's just the kind of glue you copy between repos and half-remember how it works six months later.&lt;/p&gt;

&lt;p&gt;Here's the part that actually cost me an afternoon. A healthcheck tells docker-rollout the new container is ready. It says nothing about the old one. Traefik keeps routing to it right up until it's removed, and if that removal lands mid-request, that request just dies with a 502. I only caught it because I put a deploy under constant load and counted failures instead of curling the site twice and calling it done, which is roughly how I'd been "testing" this before. Every single deploy was dropping a handful of requests, consistently, invisibly.&lt;/p&gt;

&lt;p&gt;The fix is connection draining. Make the healthcheck also check for a marker file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;healthcheck&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;test&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;test ! -f /tmp/drain &amp;amp;&amp;amp; curl -fsS http://localhost:8000/health&lt;/span&gt;
  &lt;span class="na"&gt;interval&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;5s&lt;/span&gt;
  &lt;span class="na"&gt;retries&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;1&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and before stopping the old container, touch that file and wait:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pre-stop-hook: &lt;span class="nb"&gt;touch&lt;/span&gt; /tmp/drain &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nb"&gt;sleep &lt;/span&gt;10
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Touching /tmp/drain makes the healthcheck fail on purpose. Traefik notices within interval × retries and quietly stops sending new traffic there. The sleep needs to cover that window plus whatever your slowest in-flight request takes. By the time docker-rollout actually stops the container, nothing's running on it anymore to interrupt. Ran the same load test with draining added: zero dropped requests.&lt;/p&gt;

&lt;p&gt;I packaged the whole thing, file upload, registry login, pull, a pre-deploy hook for migrations, rollout with draining, automatic rollback if a release never goes healthy, into a GitHub Action (&lt;a href="https://github.com/mtizima/docker-rollout-action):" rel="noopener noreferrer"&gt;https://github.com/mtizima/docker-rollout-action):&lt;/a&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;uses&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mtizima/docker-rollout-action@v1&lt;/span&gt;
  &lt;span class="na"&gt;with&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;host&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.SSH_HOST }}&lt;/span&gt;
    &lt;span class="na"&gt;user&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;deploy&lt;/span&gt;
    &lt;span class="na"&gt;ssh-key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.SSH_KEY }}&lt;/span&gt;
    &lt;span class="na"&gt;known-hosts&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.SSH_KNOWN_HOSTS }}&lt;/span&gt;
    &lt;span class="na"&gt;project-dir&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;/opt/myapp&lt;/span&gt;
    &lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;web&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Everything runs in one SSH session, plain bash, nothing to install on the runner beyond ssh itself. Inputs go over stdin instead of the command line so nothing shows up in ps, and registry credentials sit in a Docker config that's deleted the second the deploy finishes.&lt;/p&gt;

&lt;p&gt;Couple of real constraints worth knowing before you touch this. Services you roll out this way can't have &lt;code&gt;container_name&lt;/code&gt; or published ports, since two containers of the same service run side by side during the swap, so you need a proxy routing to them instead. And they need an actual Docker healthcheck. Without one, docker-rollout just waits a fixed number of seconds and hopes for the best, which is exactly the kind of thing that gets you a repeat of this whole post.&lt;/p&gt;

&lt;p&gt;If you're running Kubernetes already, none of this is for you. If you're on one or two boxes with Traefik or nginx-proxy in front and you've never actually measured whether your deploys drop requests, it might be worth five minutes to check. Mine didn't, until it did.&lt;/p&gt;

</description>
      <category>docker</category>
      <category>devops</category>
      <category>githubactions</category>
      <category>showdev</category>
    </item>
  </channel>
</rss>
