<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ionut-Robert Sandu</title>
    <description>The latest articles on DEV Community by Ionut-Robert Sandu (@riskbitsdotnet).</description>
    <link>https://dev.to/riskbitsdotnet</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4122206%2Fbbbcac33-36d4-4160-bb69-5ff839e2485a.png</url>
      <title>DEV Community: Ionut-Robert Sandu</title>
      <link>https://dev.to/riskbitsdotnet</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/riskbitsdotnet"/>
    <language>en</language>
    <item>
      <title>Your Health Check Is Green and Your Service Is Down</title>
      <dc:creator>Ionut-Robert Sandu</dc:creator>
      <pubDate>Fri, 25 Sep 2026 17:11:37 +0000</pubDate>
      <link>https://dev.to/riskbitsdotnet/your-health-check-is-green-and-your-service-is-down-5eec</link>
      <guid>https://dev.to/riskbitsdotnet/your-health-check-is-green-and-your-service-is-down-5eec</guid>
      <description>&lt;p&gt;Monitoring has two ways to be wrong, and we treat them as if they were&lt;br&gt;
symmetrical. They are not, and the asymmetry should drive how you build checks.&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;false alarm&lt;/strong&gt; costs you an interruption. Someone looks, finds nothing, mutters,&lt;br&gt;
goes back to what they were doing. Annoying, self-correcting, and — this is the&lt;br&gt;
important part — &lt;strong&gt;it tells you the monitoring is alive.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A &lt;strong&gt;false all-clear&lt;/strong&gt; costs you the entire outage. Nobody looks, because nothing&lt;br&gt;
asked them to. The first signal is a customer, and by then you have lost both the&lt;br&gt;
time and the claim that you knew before they did.&lt;/p&gt;

&lt;p&gt;We spend most of our tuning effort on the first failure mode. Almost all of the&lt;br&gt;
damage comes from the second.&lt;/p&gt;
&lt;h2&gt;
  
  
  Three ways a check reports green on a broken service
&lt;/h2&gt;
&lt;h3&gt;
  
  
  1. The check is narrower than the thing it claims to cover
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;GET /health&lt;/code&gt; returns 200. The endpoint returns a literal &lt;code&gt;{"status":"ok"}&lt;/code&gt; from a&lt;br&gt;
handler that touches nothing. It confirms that the process is running and the HTTP&lt;br&gt;
listener is bound. It says nothing about the database connection pool, the expired&lt;br&gt;
credential to the payment provider, or the disk that filled up twenty minutes ago.&lt;/p&gt;

&lt;p&gt;This is not laziness; it is drift. The check was written when the service had two&lt;br&gt;
dependencies and was honest about both. Four years and eleven dependencies later,&lt;br&gt;
the check still passes because it still does exactly what it did in 2022.&lt;/p&gt;

&lt;p&gt;The useful question is not "is the check passing" but &lt;strong&gt;"what would have to break&lt;br&gt;
for this check to fail?"&lt;/strong&gt; If the answer is "the process would have to be dead",&lt;br&gt;
you have a liveness probe wearing a health check's name tag. Both are legitimate,&lt;br&gt;
but only one of them is allowed on the dashboard the on-call engineer looks at.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. The check runs somewhere that cannot fail the way users do
&lt;/h3&gt;

&lt;p&gt;Checks that run inside the cluster share fate with the thing they measure. They use&lt;br&gt;
internal DNS, skip the load balancer, bypass the CDN, sit inside the network policy,&lt;br&gt;
and often talk to the pod directly rather than through the service. Every one of&lt;br&gt;
those is a component that can break for users while leaving the check perfectly&lt;br&gt;
happy.&lt;/p&gt;

&lt;p&gt;Certificate monitoring makes this concrete. Probe public HTTPS endpoints from inside&lt;br&gt;
an environment whose egress&lt;br&gt;
&lt;a href="https://dev.to/riskbitsdotnet/your-app-works-everywhere-except-the-corporate-network-2697"&gt;terminates and re-signs TLS&lt;/a&gt;&lt;br&gt;
and every certificate that comes back is issued by the gateway, not by the site you&lt;br&gt;
think you are watching. Build a certificate monitor there, point it at production,&lt;br&gt;
and it reports healthy forever — because it is measuring the appliance one hop away&lt;br&gt;
rather than the service on the other side of the internet. The check is fine. The&lt;br&gt;
vantage point makes it meaningless.&lt;/p&gt;

&lt;p&gt;The rule that falls out: &lt;strong&gt;at least one check has to traverse the same path a user&lt;br&gt;
does&lt;/strong&gt;, including the parts you do not own. Internal checks tell you which component&lt;br&gt;
broke. Only an external check tells you whether anyone is actually affected.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. The check stopped running
&lt;/h3&gt;

&lt;p&gt;This is the one that produces the longest outages, and it is structurally&lt;br&gt;
invisible. A cron job that no longer fires, a worker that died at 03:00, a pipeline&lt;br&gt;
whose credential silently expired — none of these emit anything. And the absence of&lt;br&gt;
an alert is indistinguishable from the absence of a problem.&lt;/p&gt;

&lt;p&gt;Think about what your dashboard actually renders when a checker is dead. Most of&lt;br&gt;
them show the last known state, which was green, with a timestamp nobody reads. The&lt;br&gt;
system is not lying to you. You just asked it a question it has no way to answer.&lt;/p&gt;
&lt;h2&gt;
  
  
  The heartbeat inverts the logic
&lt;/h2&gt;

&lt;p&gt;The fix for the third case is to stop asking "did something report a failure" and&lt;br&gt;
start asking "did something report at all".&lt;/p&gt;

&lt;p&gt;A dead man's switch is an external service expecting a ping on a schedule. Your&lt;br&gt;
checker pings it after every successful run. If the ping stops, the external&lt;br&gt;
service alerts. The alert now fires on &lt;em&gt;silence&lt;/em&gt;, which means the failure of your&lt;br&gt;
monitoring is itself a monitored event.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/sh&lt;/span&gt;
&lt;span class="nb"&gt;set&lt;/span&gt; &lt;span class="nt"&gt;-e&lt;/span&gt;

run_all_checks            &lt;span class="c"&gt;# exits non-zero if anything is wrong&lt;/span&gt;

&lt;span class="c"&gt;# only reached when the checks ran AND passed&lt;/span&gt;
curl &lt;span class="nt"&gt;-fsS&lt;/span&gt; &lt;span class="nt"&gt;-m&lt;/span&gt; 10 &lt;span class="nt"&gt;--retry&lt;/span&gt; 3 &lt;span class="s2"&gt;"https://hc-ping.example/&lt;/span&gt;&lt;span class="nv"&gt;$CHECK_UUID&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; /dev/null
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three things about that snippet are load-bearing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;set -e&lt;/code&gt; and the ordering.&lt;/strong&gt; The ping is last. If the checks fail, or crash, or&lt;br&gt;
the box runs out of memory halfway through, the ping never happens, and you get an&lt;br&gt;
alert from the outside. Ping first and you have built a system that reports success&lt;br&gt;
before knowing whether there is any.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The pinged service must be external.&lt;/strong&gt; A heartbeat monitored by the same&lt;br&gt;
infrastructure that runs the checker shares fate with it. Hosted dead man's switch&lt;br&gt;
services exist and are cheap; the point is the independence, not the vendor.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The expected period should be roughly twice the run interval.&lt;/strong&gt; One missed run on&lt;br&gt;
a five-minute job is usually a blip. Two is a pattern. Alert on the pattern, or you&lt;br&gt;
will train yourself to ignore the alert — which returns you to a false all-clear by&lt;br&gt;
a slower route.&lt;/p&gt;

&lt;h2&gt;
  
  
  Designing checks that fail loudly
&lt;/h2&gt;

&lt;p&gt;A few habits that follow from all of this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make "unknown" a distinct state from "healthy".&lt;/strong&gt; If your checker cannot reach a&lt;br&gt;
target, that is not a pass and it is not necessarily a fail — it is stale data, and&lt;br&gt;
it should render differently. Most dashboards have two colours where they need&lt;br&gt;
three. A check that has not reported in an hour should look visibly wrong even if&lt;br&gt;
its last result was green.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test the failure path, not just the success path.&lt;/strong&gt; A check nobody has ever seen&lt;br&gt;
fail is a check nobody has verified. Break the dependency deliberately in a staging&lt;br&gt;
environment and confirm the alert actually arrives, at the actual destination, in&lt;br&gt;
the actual channel. Alert routing rots quietly: people leave, channels get archived,&lt;br&gt;
webhooks expire. The routing is part of the system and needs the same treatment as&lt;br&gt;
the code.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Write down what each check does not cover.&lt;/strong&gt; One line next to the definition.&lt;br&gt;
"Does not verify database connectivity." It costs nothing to write and it is the&lt;br&gt;
sentence that saves you during the post-mortem, because it converts an unknown&lt;br&gt;
unknown into a known gap that somebody can choose to close or accept.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Prefer a check that occasionally cries wolf over one that never speaks.&lt;/strong&gt; Given&lt;br&gt;
the asymmetry at the top of this post, a slightly noisy check is a trade you should&lt;br&gt;
be willing to make on purpose rather than one you drift into.&lt;/p&gt;

&lt;h2&gt;
  
  
  The uncomfortable summary
&lt;/h2&gt;

&lt;p&gt;Green means one of two things: the system is healthy, or the system is not being&lt;br&gt;
measured. Most monitoring setups cannot distinguish between the two, and the&lt;br&gt;
dashboard renders both identically.&lt;/p&gt;

&lt;p&gt;The work is not making the checks more sensitive. It is making sure that when the&lt;br&gt;
measurement itself fails, something notices — because that is the failure mode that&lt;br&gt;
costs you the whole outage, and it is the only one that gets quieter the worse it&lt;br&gt;
gets.&lt;/p&gt;

</description>
      <category>sre</category>
      <category>devops</category>
      <category>monitoring</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Your Certificate Monitoring Only Checks the Leaf. That Is Not the Same Thing as Your Chain Being Valid.</title>
      <dc:creator>Ionut-Robert Sandu</dc:creator>
      <pubDate>Fri, 18 Sep 2026 08:10:41 +0000</pubDate>
      <link>https://dev.to/riskbitsdotnet/your-certificate-monitoring-only-checks-the-leaf-that-is-not-the-same-thing-as-your-chain-being-1k00</link>
      <guid>https://dev.to/riskbitsdotnet/your-certificate-monitoring-only-checks-the-leaf-that-is-not-the-same-thing-as-your-chain-being-1k00</guid>
      <description>&lt;p&gt;Here is a monitoring result. It is from a server built in a lab a few minutes earlier,&lt;br&gt;
and every number in it is real.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$ &lt;/span&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; | openssl s_client &lt;span class="nt"&gt;-connect&lt;/span&gt; app.lab.example:443 &lt;span class="nt"&gt;-servername&lt;/span&gt; app.lab.example &lt;span class="se"&gt;\&lt;/span&gt;
  2&amp;gt;/dev/null | openssl x509 &lt;span class="nt"&gt;-noout&lt;/span&gt; &lt;span class="nt"&gt;-enddate&lt;/span&gt; &lt;span class="nt"&gt;-checkend&lt;/span&gt; 2592000

&lt;span class="nv"&gt;notAfter&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;Dec 11 14:12:02 2026 GMT
Certificate will not expire
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ninety days of runway, comfortably past the thirty-day threshold. Green.&lt;/p&gt;

&lt;p&gt;The site stops working in twenty days.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the check missed
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;openssl s_client&lt;/code&gt; hands back the leaf certificate by default. That is the&lt;br&gt;
certificate for the hostname, the one with your domain in it, the one everybody&lt;br&gt;
means when they say "the certificate". It is also only the first entry in a list.&lt;br&gt;
A browser does not trust a leaf because the leaf says a date. It trusts it&lt;br&gt;
because it can build a path from that leaf to a root it already trusts, and&lt;br&gt;
&lt;strong&gt;every certificate on that path has to be valid at the same moment.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Here is the same server, with every certificate it sends checked instead of only&lt;br&gt;
the first one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;CN = app.lab.example               Dec 11 14:12:02 2026 GMT   ok
CN = Lab Intermediate CA (short)   Oct  2 14:12:02 2026 GMT   EXPIRES WITHIN 30d
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The leaf outlives the intermediate that signed it. On October 2nd, the chain&lt;br&gt;
stops validating, and the leaf's own December expiry becomes irrelevant. Clients&lt;br&gt;
will report an expired certificate. Your dashboard will report ninety days&lt;br&gt;
remaining, because the thing your dashboard is looking at does, in fact, have&lt;br&gt;
ninety days remaining.&lt;/p&gt;

&lt;p&gt;Nothing here is exotic. A CA can perfectly well issue you a certificate that&lt;br&gt;
outlives its own issuing intermediate, and the certificate is not malformed when&lt;br&gt;
it does. The constraint lives in path validation at the client, not in the&lt;br&gt;
issuance.&lt;/p&gt;
&lt;h2&gt;
  
  
  Build it yourself in two minutes
&lt;/h2&gt;

&lt;p&gt;Do not take my word for the shape of this. The whole thing is four &lt;code&gt;openssl&lt;/code&gt;&lt;br&gt;
invocations.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Root CA, long lived&lt;/span&gt;
openssl req &lt;span class="nt"&gt;-x509&lt;/span&gt; &lt;span class="nt"&gt;-newkey&lt;/span&gt; rsa:2048 &lt;span class="nt"&gt;-nodes&lt;/span&gt; &lt;span class="nt"&gt;-keyout&lt;/span&gt; root.key &lt;span class="nt"&gt;-out&lt;/span&gt; root.crt &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-days&lt;/span&gt; 3650 &lt;span class="nt"&gt;-subj&lt;/span&gt; &lt;span class="s2"&gt;"/CN=Lab Root CA"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-addext&lt;/span&gt; &lt;span class="s2"&gt;"basicConstraints=critical,CA:TRUE"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-addext&lt;/span&gt; &lt;span class="s2"&gt;"keyUsage=critical,keyCertSign,cRLSign"&lt;/span&gt;

&lt;span class="c"&gt;# Intermediate CA that expires in 20 days&lt;/span&gt;
openssl req &lt;span class="nt"&gt;-newkey&lt;/span&gt; rsa:2048 &lt;span class="nt"&gt;-nodes&lt;/span&gt; &lt;span class="nt"&gt;-keyout&lt;/span&gt; inter.key &lt;span class="nt"&gt;-out&lt;/span&gt; inter.csr &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-subj&lt;/span&gt; &lt;span class="s2"&gt;"/CN=Lab Intermediate CA (short)"&lt;/span&gt;
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'basicConstraints=critical,CA:TRUE,pathlen:0\nkeyUsage=critical,keyCertSign,cRLSign\n'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; inter.ext
openssl x509 &lt;span class="nt"&gt;-req&lt;/span&gt; &lt;span class="nt"&gt;-in&lt;/span&gt; inter.csr &lt;span class="nt"&gt;-CA&lt;/span&gt; root.crt &lt;span class="nt"&gt;-CAkey&lt;/span&gt; root.key &lt;span class="nt"&gt;-CAcreateserial&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-out&lt;/span&gt; inter.crt &lt;span class="nt"&gt;-days&lt;/span&gt; 20 &lt;span class="nt"&gt;-extfile&lt;/span&gt; inter.ext

&lt;span class="c"&gt;# Leaf valid for 90 days, signed by that intermediate&lt;/span&gt;
openssl req &lt;span class="nt"&gt;-newkey&lt;/span&gt; rsa:2048 &lt;span class="nt"&gt;-nodes&lt;/span&gt; &lt;span class="nt"&gt;-keyout&lt;/span&gt; leaf.key &lt;span class="nt"&gt;-out&lt;/span&gt; leaf.csr &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-subj&lt;/span&gt; &lt;span class="s2"&gt;"/CN=app.lab.example"&lt;/span&gt;
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'basicConstraints=CA:FALSE\nkeyUsage=critical,digitalSignature,keyEncipherment\nextendedKeyUsage=serverAuth\nsubjectAltName=DNS:app.lab.example\n'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; leaf.ext
openssl x509 &lt;span class="nt"&gt;-req&lt;/span&gt; &lt;span class="nt"&gt;-in&lt;/span&gt; leaf.csr &lt;span class="nt"&gt;-CA&lt;/span&gt; inter.crt &lt;span class="nt"&gt;-CAkey&lt;/span&gt; inter.key &lt;span class="nt"&gt;-CAcreateserial&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-out&lt;/span&gt; leaf.crt &lt;span class="nt"&gt;-days&lt;/span&gt; 90 &lt;span class="nt"&gt;-extfile&lt;/span&gt; leaf.ext

&lt;span class="nb"&gt;cat &lt;/span&gt;leaf.crt inter.crt &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; fullchain.crt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Serve &lt;code&gt;fullchain.crt&lt;/code&gt; from anything, point your existing certificate monitor at&lt;br&gt;
it, and watch it tell you everything is fine.&lt;/p&gt;
&lt;h2&gt;
  
  
  Checking the whole chain
&lt;/h2&gt;

&lt;p&gt;The fix is not complicated, which is part of why the gap is annoying. Pull every&lt;br&gt;
certificate the server actually sends, and run the expiry check against each one.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;#!/bin/sh&lt;/span&gt;
&lt;span class="c"&gt;# Usage: tls-chain-check.sh host:port [sni] [days]&lt;/span&gt;
&lt;span class="nv"&gt;HOST&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$1&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;SNI&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;2&lt;/span&gt;&lt;span class="k"&gt;:-${&lt;/span&gt;&lt;span class="nv"&gt;1&lt;/span&gt;&lt;span class="p"&gt;%%&lt;/span&gt;:&lt;span class="p"&gt;*&lt;/span&gt;&lt;span class="k"&gt;}}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;DAYS&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;3&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;30&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;TMP&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;mktemp&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;trap&lt;/span&gt; &lt;span class="s1"&gt;'rm -rf "$TMP"'&lt;/span&gt; EXIT

&lt;span class="nb"&gt;echo&lt;/span&gt; | openssl s_client &lt;span class="nt"&gt;-connect&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOST&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-servername&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$SNI&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-showcerts&lt;/span&gt; 2&amp;gt;/dev/null &lt;span class="se"&gt;\&lt;/span&gt;
| &lt;span class="nb"&gt;awk&lt;/span&gt; &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nv"&gt;d&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TMP&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s1"&gt;'
    /-----BEGIN CERTIFICATE-----/ { n++; p=1 }
    p                             { print &amp;gt;&amp;gt; (d "/c" n ".pem") }
    /-----END CERTIFICATE-----/   { p=0 }'&lt;/span&gt;

&lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="nt"&gt;-f&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TMP&lt;/span&gt;&lt;span class="s2"&gt;/c1.pem"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"no certificates received from &lt;/span&gt;&lt;span class="nv"&gt;$HOST&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;2&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;

&lt;span class="nv"&gt;rc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0
&lt;span class="k"&gt;for &lt;/span&gt;f &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$TMP&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;/c&lt;span class="k"&gt;*&lt;/span&gt;.pem&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
    &lt;/span&gt;&lt;span class="nv"&gt;subj&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;openssl x509 &lt;span class="nt"&gt;-in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-noout&lt;/span&gt; &lt;span class="nt"&gt;-subject&lt;/span&gt; | &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s/^subject=[ ]*//'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
    &lt;span class="nv"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;openssl x509 &lt;span class="nt"&gt;-in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-noout&lt;/span&gt; &lt;span class="nt"&gt;-enddate&lt;/span&gt; | &lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="s1"&gt;'s/^notAfter=//'&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;openssl x509 &lt;span class="nt"&gt;-in&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$f&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-noout&lt;/span&gt; &lt;span class="nt"&gt;-checkend&lt;/span&gt; &lt;span class="k"&gt;$((&lt;/span&gt;DAYS &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="m"&gt;86400&lt;/span&gt;&lt;span class="k"&gt;))&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt;/dev/null 2&amp;gt;&amp;amp;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
        &lt;/span&gt;&lt;span class="nv"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"ok"&lt;/span&gt;
    &lt;span class="k"&gt;else
        &lt;/span&gt;&lt;span class="nv"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"EXPIRES WITHIN &lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;DAYS&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;d"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nv"&gt;rc&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;1
    &lt;span class="k"&gt;fi
    &lt;/span&gt;&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'%-34s %-26s %s\n'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$subj&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$end&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$state&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;done
&lt;/span&gt;&lt;span class="nb"&gt;exit&lt;/span&gt; &lt;span class="nv"&gt;$rc&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Non-zero exit when anything on the path is inside the window, which is what you&lt;br&gt;
want for a cron job or a CI gate.&lt;/p&gt;

&lt;p&gt;Two details in there are deliberate and worth stealing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;-showcerts&lt;/code&gt; before anything else.&lt;/strong&gt; Without it you get the leaf and nothing&lt;br&gt;
else, which is how you ended up here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;-servername&lt;/code&gt; separate from the connect address.&lt;/strong&gt; On a host serving multiple&lt;br&gt;
sites, the certificate you get depends on the SNI you send, not on the IP you&lt;br&gt;
connected to. If you monitor by IP and omit SNI, you are checking whatever the&lt;br&gt;
default virtual host happens to present, which may be a completely different&lt;br&gt;
certificate from the one your users receive. This is a good way to monitor&lt;br&gt;
something real and useless for eighteen months.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three more places the same mistake hides
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The chain the server sends is not always the chain the client builds.&lt;/strong&gt;&lt;br&gt;
A client with a cached intermediate, or one that follows the AIA extension to&lt;br&gt;
fetch a missing issuer, can successfully validate a chain your server sent&lt;br&gt;
incompletely. Your desktop browser says fine; a stripped-down container with no&lt;br&gt;
cached intermediates and no outbound access to the CA's AIA endpoint says&lt;br&gt;
handshake failure. Checking from a machine with a rich trust state hides the&lt;br&gt;
problem from you specifically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Your checker's trust store is not your users' trust store.&lt;/strong&gt; If the checker&lt;br&gt;
runs &lt;code&gt;openssl verify&lt;/code&gt; against the OS bundle on a long-lived VM, you are&lt;br&gt;
validating against whatever roots that VM had at its last update. Mobile clients,&lt;br&gt;
older Java runtimes, and embedded devices all carry different sets. "Valid" is&lt;br&gt;
not a property of a certificate. It is a relationship between a certificate and&lt;br&gt;
a particular trust store at a particular time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The interception layer has its own copy.&lt;/strong&gt; If traffic crosses a proxy that&lt;br&gt;
&lt;a href="https://dev.to/riskbitsdotnet/your-app-works-everywhere-except-the-corporate-network-2697"&gt;terminates and re-signs TLS&lt;/a&gt;,&lt;br&gt;
there are now two certificates on the path to your user, issued by two different&lt;br&gt;
authorities, renewed by two different processes.&lt;br&gt;
Probe a public site from inside an environment with intercepted egress and every&lt;br&gt;
certificate comes back signed by the gateway's own CA — the leaf you would be&lt;br&gt;
"monitoring" was manufactured about a second earlier. If you run monitoring from&lt;br&gt;
inside a corporate network and alert on the certificate you receive, you may be&lt;br&gt;
monitoring the health of your own proxy.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to actually alert on
&lt;/h2&gt;

&lt;p&gt;Alert on the &lt;strong&gt;minimum remaining lifetime across the full path&lt;/strong&gt;, not on the&lt;br&gt;
leaf. It is one number, it is the one that determines when things break, and it&lt;br&gt;
is strictly more conservative than what most tools give you today.&lt;/p&gt;

&lt;p&gt;Then make the thresholds reflect who fixes what. A leaf inside thirty days is&lt;br&gt;
your problem and your automation should have handled it. An intermediate inside&lt;br&gt;
sixty days is usually your CA's problem, and the remediation is to re-issue and&lt;br&gt;
redeploy — a different task, on a different timeline, often owned by a different&lt;br&gt;
team. Folding both into a single "certificate expiring" alert guarantees that at&lt;br&gt;
least one of them gets handled by someone who cannot fix it.&lt;/p&gt;

&lt;p&gt;This matters more every year, not less. Public certificate lifetimes are already&lt;br&gt;
down from 398 days to 200, drop to 100 in March 2027, and reach 47 in March 2029.&lt;br&gt;
Renewal stops being a calendar event and becomes a pipeline. Pipelines fail&lt;br&gt;
quietly, which means the monitoring has to be right about what it is looking at.&lt;/p&gt;

&lt;p&gt;Checking the leaf was a reasonable approximation when certificates lasted a year.&lt;br&gt;
It is not one anymore.&lt;/p&gt;

</description>
      <category>security</category>
      <category>tls</category>
      <category>devops</category>
      <category>sre</category>
    </item>
    <item>
      <title>Your App Works Everywhere Except the Corporate Network</title>
      <dc:creator>Ionut-Robert Sandu</dc:creator>
      <pubDate>Sat, 12 Sep 2026 13:57:36 +0000</pubDate>
      <link>https://dev.to/riskbitsdotnet/your-app-works-everywhere-except-the-corporate-network-2697</link>
      <guid>https://dev.to/riskbitsdotnet/your-app-works-everywhere-except-the-corporate-network-2697</guid>
      <description>&lt;p&gt;You ship something. It works on your machine, in CI, in staging, and for every user who tries it. Then one customer opens a ticket: it doesn't work for them. Same version, same config, same everything. It just hangs, or throws a certificate error, or fails in some way your error handling never anticipated.&lt;/p&gt;

&lt;p&gt;They're on a corporate network. Somewhere between their machine and your server, a device is opening every TLS connection, reading it, and building a new one.&lt;/p&gt;

&lt;p&gt;This is normal. Large organisations are often legally required to inspect traffic leaving their network, and they've been doing it for twenty years. The problem isn't that it happens. The problem is that from where you're standing, it's nearly invisible — and the failures it causes look like bugs in your code.&lt;/p&gt;

&lt;p&gt;Here's what's actually happening, why five different things break in five different ways, and how to work out which one you're looking at without access to the customer's network team.&lt;/p&gt;

&lt;h2&gt;
  
  
  What interception actually does
&lt;/h2&gt;

&lt;p&gt;A normal TLS connection is between your client and your server. The server presents a certificate, the client checks it chains to a certificate authority it trusts, and the two negotiate keys that nobody in the middle can derive. That's the entire point.&lt;/p&gt;

&lt;p&gt;An inspection proxy breaks this into two connections. It terminates the client's TLS session itself, then opens a separate one to the real server. In the middle, it holds plaintext.&lt;/p&gt;

&lt;p&gt;For the client to accept this, the proxy has to present a certificate for the site the client asked for. It generates one on the fly, signed by a CA whose root certificate the organisation has installed on every managed machine. On a managed laptop, this works transparently. The browser sees a valid chain to a root it trusts, and shows a padlock.&lt;/p&gt;

&lt;p&gt;One consequence is worth stating plainly, because a lot of developers get it backwards: &lt;strong&gt;this isn't a vulnerability being exploited.&lt;/strong&gt; It works because an administrator deliberately installed a root certificate. Browsers make a specific exception for locally-installed roots — Certificate Transparency requirements and most key-pinning checks are enforced for publicly-trusted CAs but relaxed for locally-added ones. If they weren't, corporate inspection would break the web for a large fraction of enterprise users, so the exception is deliberate.&lt;/p&gt;

&lt;p&gt;That exception is also the reason the failures are so confusing. Interception works fine for the browser, which is why the customer insists their internet is "working" — and breaks for everything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five mechanisms, five different failure signatures
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Your language runtime doesn't use the system trust store
&lt;/h3&gt;

&lt;p&gt;This is the most common one and the least known.&lt;/p&gt;

&lt;p&gt;The corporate root CA gets installed into the operating system trust store. The browser picks it up. Your application might not, because several runtimes ship their own bundled list of certificate authorities and ignore the OS entirely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Node.js&lt;/strong&gt; bundles its own root store. It does not read the system store by default. The escape hatch is &lt;code&gt;NODE_EXTRA_CA_CERTS&lt;/code&gt;, pointing at a PEM file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Python&lt;/strong&gt;, when using &lt;code&gt;requests&lt;/code&gt; or anything else built on &lt;code&gt;certifi&lt;/code&gt;, uses the &lt;code&gt;certifi&lt;/code&gt; bundle rather than the system store. &lt;code&gt;REQUESTS_CA_BUNDLE&lt;/code&gt; and &lt;code&gt;SSL_CERT_FILE&lt;/code&gt; override it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Java&lt;/strong&gt; uses its own &lt;code&gt;cacerts&lt;/code&gt; keystore, managed with &lt;code&gt;keytool&lt;/code&gt;, entirely separate from the OS.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Go&lt;/strong&gt; does read system roots on Linux and macOS, but any code that sets a custom &lt;code&gt;tls.Config&lt;/code&gt; with its own &lt;code&gt;RootCAs&lt;/code&gt; pool has opted out.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Signature:&lt;/strong&gt; the site loads in the browser on the same machine, but your application throws a certificate verification error. That specific combination — browser fine, code broken — almost always means a trust store mismatch rather than anything wrong with the network.&lt;/p&gt;

&lt;p&gt;Firefox is worth a special mention. It maintains its own trust store independently of the OS, so an environment where Chrome works and Firefox doesn't is the same problem wearing a different hat.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Certificate pinning
&lt;/h3&gt;

&lt;p&gt;If your application pins a specific certificate or public key — common in mobile apps and in anything handling payments — you've explicitly said you will only accept one identity, regardless of what the trust store says. An inspection proxy cannot satisfy that. It doesn't have the private key.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signature:&lt;/strong&gt; fails on every corporate network, works everywhere else, and no amount of installing certificates fixes it. Unlike the trust store problem, this one is unfixable from the client side. That's the point of pinning.&lt;/p&gt;

&lt;p&gt;If you pin, you need a documented way for administrators to allow-list your domain from inspection, and your error message should say so. Most apps that pin fail with a generic network error, which turns a five-minute fix into a week of support tickets.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Mutual TLS
&lt;/h3&gt;

&lt;p&gt;If your server requires a client certificate, the proxy has to present one on the client's behalf. It doesn't have the client's private key, so it can't. Some proxies detect this and pass the connection through untouched; others don't, and the handshake dies.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signature:&lt;/strong&gt; the handshake fails after the server requests a certificate, often with an unhelpful message about a missing or bad certificate. The tell is that it fails at a specific handshake stage, not at verification.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Protocol upgrades and long-lived connections
&lt;/h3&gt;

&lt;p&gt;WebSockets, server-sent events, gRPC streams and anything else that holds a connection open or upgrades it mid-flight depend on the middlebox handling that correctly. Many do. Some strip the &lt;code&gt;Upgrade&lt;/code&gt; header, some apply an idle timeout far shorter than your keepalive interval, and some buffer responses that you intended to stream.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signature:&lt;/strong&gt; the initial request succeeds and then the connection dies, or the upgrade silently doesn't happen and you fall back to polling. Reconnect loops with no error are typical.&lt;/p&gt;

&lt;p&gt;This one is nasty because it is intermittent and timing-dependent, which makes it look like a bug in your reconnection logic. If your connections die at a suspiciously round interval — sixty seconds, five minutes — that's an idle timeout somewhere in the path, not your code.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. QUIC and HTTP/3
&lt;/h3&gt;

&lt;p&gt;QUIC encrypts most of its transport metadata, which defeats the inspection techniques built for TLS over TCP. The usual enterprise response is to block UDP/443 outright and force clients back to TCP.&lt;/p&gt;

&lt;p&gt;Browsers handle this gracefully — they try QUIC, fail, and fall back. Custom clients often don't.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Signature:&lt;/strong&gt; a delay of several seconds before every connection, or a hang, on one network only. If your HTTP client prefers HTTP/3, test it with QUIC disabled before blaming anything else.&lt;/p&gt;

&lt;h2&gt;
  
  
  Triage from the client side
&lt;/h2&gt;

&lt;p&gt;You usually can't get access to the customer's network. You can get them to run two commands.&lt;/p&gt;

&lt;p&gt;The first question is always: &lt;strong&gt;is the connection being intercepted at all?&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openssl s_client &lt;span class="nt"&gt;-connect&lt;/span&gt; example.com:443 &lt;span class="nt"&gt;-servername&lt;/span&gt; example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look at the issuer line. On an unintercepted connection you'll see a public CA. On an intercepted one you'll see something internal — the organisation's name, or a firewall vendor's. That single line answers the question.&lt;/p&gt;

&lt;p&gt;To see the whole chain, including what the proxy is signing with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;openssl s_client &lt;span class="nt"&gt;-connect&lt;/span&gt; example.com:443 &lt;span class="nt"&gt;-servername&lt;/span&gt; example.com &lt;span class="nt"&gt;-showcerts&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then compare what your application sees against what the system sees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-v&lt;/span&gt; https://example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;curl&lt;/code&gt; succeeds and your application fails on the same machine, you're in mechanism 1 — a trust store mismatch. If &lt;code&gt;curl&lt;/code&gt; fails too, the problem is at the network layer and affects everything.&lt;/p&gt;

&lt;p&gt;To confirm a certificate issue specifically rather than a network one, try without verification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-vk&lt;/span&gt; https://example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If it works with &lt;code&gt;-k&lt;/code&gt; and fails without, it's trust. If it fails both ways, it isn't. Never leave &lt;code&gt;-k&lt;/code&gt; in anything but a diagnostic.&lt;/p&gt;

&lt;p&gt;For a suspected HTTP/3 problem, there is a trap worth knowing about. &lt;code&gt;curl&lt;/code&gt; does &lt;strong&gt;not&lt;/strong&gt; negotiate HTTP/3 by default — that needs a build with QUIC support and an explicit flag. So check what you actually got before drawing conclusions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-sS&lt;/span&gt; &lt;span class="nt"&gt;-o&lt;/span&gt; /dev/null &lt;span class="nt"&gt;-w&lt;/span&gt; &lt;span class="s2"&gt;"%{http_version}&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; https://example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that prints &lt;code&gt;2&lt;/code&gt;, curl never attempted QUIC, and comparing it against &lt;code&gt;--http1.1&lt;/code&gt; proves nothing at all. Force the attempt instead:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-v&lt;/span&gt; &lt;span class="nt"&gt;--http3&lt;/span&gt; https://example.com
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the explicit HTTP/3 attempt hangs or fails while the ordinary request succeeds, UDP/443 is being blocked somewhere in the path. Browsers conceal this, because they try QUIC, fail, and fall back silently — which is precisely why testing in a browser will tell you everything is fine.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to build so this is less painful
&lt;/h2&gt;

&lt;p&gt;You will not stop corporate networks from inspecting traffic. What you can do is stop your application from being the hardest part of the diagnosis.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Say what actually failed.&lt;/strong&gt; "Network error" is worthless. "Certificate verification failed: issuer not trusted" points the user at their administrator immediately. Surface the issuer name if you have it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Document the domains you need.&lt;/strong&gt; Every serious enterprise vendor publishes a list of hostnames to allow-list, and states plainly if a domain must be exempt from inspection. If you pin certificates, this documentation is not optional.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Make the trust store configurable.&lt;/strong&gt; Respect &lt;code&gt;NODE_EXTRA_CA_CERTS&lt;/code&gt;, &lt;code&gt;SSL_CERT_FILE&lt;/code&gt;, or whatever your runtime's equivalent is, and say so in your docs. Administrators know what to do with that.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Never disable verification as a workaround.&lt;/strong&gt; The number of production systems running with verification switched off because someone hit this once and needed it working by Friday is genuinely alarming. It converts a support ticket into a permanent vulnerability, and nobody ever turns it back on.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Test against an interception proxy.&lt;/strong&gt; mitmproxy in front of your integration tests, with its CA installed, reproduces most of this in an afternoon. It's a better use of time than debugging it live against a customer who can't share their network configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  The short version
&lt;/h2&gt;

&lt;p&gt;If a customer reports that your application fails only on their network, the question isn't whether your code is wrong. The question is which of five things a middlebox is doing to your connection — and the first &lt;code&gt;openssl s_client&lt;/code&gt; tells you most of it.&lt;/p&gt;

</description>
      <category>security</category>
      <category>webdev</category>
      <category>devops</category>
      <category>networking</category>
    </item>
  </channel>
</rss>
