<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: vivek Itp</title>
    <description>The latest articles on DEV Community by vivek Itp (@vivek_itp).</description>
    <link>https://dev.to/vivek_itp</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4118048%2Fc6bc377d-ae65-4385-89f6-c3d3fe39cfbf.png</url>
      <title>DEV Community: vivek Itp</title>
      <link>https://dev.to/vivek_itp</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/vivek_itp"/>
    <language>en</language>
    <item>
      <title>What an artifact signature proves, and what it does not</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Wed, 30 Sep 2026 04:02:21 +0000</pubDate>
      <link>https://dev.to/vivek_itp/what-an-artifact-signature-proves-and-what-it-does-not-43da</link>
      <guid>https://dev.to/vivek_itp/what-an-artifact-signature-proves-and-what-it-does-not-43da</guid>
      <description>&lt;p&gt;I wrote earlier about verifying commit signatures in the pipeline, and mentioned in passing that the build signs its own output. That sentence deserves a post of its own, because signing an artifact is easy and the claims people make about it are usually larger than what it actually supports.&lt;/p&gt;

&lt;p&gt;What we have works. It is also narrower than the phrase "supply chain security" suggests, and the honest version is more useful to anyone building the same thing than a diagram with green ticks on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Sign the JAR and the image, because only one of them gets deployed
&lt;/h2&gt;

&lt;p&gt;We sign both the JAR and the container image. For a while it was only the JAR, and that was the wrong place to stop.&lt;/p&gt;

&lt;p&gt;The JAR is what the build produces. The image is what actually runs. Sign only the JAR and you have proof about an input to a later step and nothing about the thing that reaches the cluster, so anyone who can influence how the image is assembled sits in a gap the signature does not cover.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What you sign&lt;/th&gt;
&lt;th&gt;What that lets you check&lt;/th&gt;
&lt;th&gt;What is still open&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The JAR only&lt;/td&gt;
&lt;td&gt;The library came from a trusted build&lt;/td&gt;
&lt;td&gt;Nothing about the image that wraps it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The image only&lt;/td&gt;
&lt;td&gt;What runs was built by a trusted pipeline&lt;/td&gt;
&lt;td&gt;Nothing about the code inside it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Both&lt;/td&gt;
&lt;td&gt;Each handoff is checkable on its own&lt;/td&gt;
&lt;td&gt;Whether the build itself did what it claimed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the second signature got added&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first version signed the JAR at publish time and stopped there. On paper the chain looked finished: signed commits, signed artifact, verification before deploy.&lt;/p&gt;

&lt;p&gt;The question that broke it came from outside the platform team: &lt;em&gt;the thing running in production is an image, so what have you proved about the image?&lt;/em&gt; The answer was nothing. The image was built in a separate job, from a base image we pulled and a JAR we trusted, and no signature covered the result. Signing the image closed that step — not the whole chain, as the rest of this post gets into, but it moved the boundary somewhere defensible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Only the build can sign, and that is the whole control
&lt;/h2&gt;

&lt;p&gt;Access to the signing key is restricted to the runner's role. No developer has it, no local build can produce a signed artifact, and there is no break-glass path that puts the key on a laptop.&lt;/p&gt;

&lt;p&gt;That restriction is the entire security property. Everything else — the verification step, the metadata, the pipeline configuration — is plumbing that only means something because the key is unreachable from anywhere except the build. The moment a person can sign, a signature stops saying "this came from the pipeline" and starts saying "this came from the pipeline, or from someone who was in a hurry".&lt;/p&gt;

&lt;p&gt;So the interesting question is not &lt;em&gt;how do we sign&lt;/em&gt; but &lt;em&gt;who can schedule a job on a runner that holds the signing role&lt;/em&gt;. That is a CI access question wearing cryptographic clothes, and it is where the real risk lives.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The request I keep having to turn down&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The recurring ask is for a way to sign locally, and it is always reasonable in the moment: a release is blocked, the pipeline is red for an unrelated reason, the artifact is fine — can someone just sign it and push it through?&lt;/p&gt;

&lt;p&gt;Saying no to that is the job. If a human can sign, every signature in the system becomes "probably from the build", and a signature that means &lt;em&gt;probably&lt;/em&gt; is not worth the pipeline minutes it costs. What we do instead is keep the pipeline path fast enough that the question is about waiting rather than being blocked — and when that stops being true, the pressure to build a bypass comes straight back.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Verify at the point of use, and fail the deploy
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The specifics in this section are the pattern rather than a story I am retelling exactly.&lt;/em&gt; Verification happens in the deploy job, before the artifact is pulled down and installed. If the signature does not verify, the deploy fails. There is no warning mode and no override flag.&lt;/p&gt;

&lt;p&gt;Two details matter more than where the check sits. &lt;strong&gt;Fail closed:&lt;/strong&gt; a step that logs a warning and continues is documentation, not a control. And &lt;strong&gt;say why in the first line:&lt;/strong&gt; "no signature found" and "signed with an unknown key" are different problems with different fixes, and nobody should need a security runbook to tell them apart.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The failure message worth writing properly&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most common verification failure is not an attack. It is an artifact that predates the control, or one promoted between environments by a path that did not carry the signature with it. With a generic message both produce the same outcome: a red deploy, a developer who did nothing wrong, and a support request.&lt;/p&gt;

&lt;p&gt;Both are obvious if the message names the artifact, says whether a signature was absent or unrecognised, and links to the one page explaining what to do next. The check is the easy part; the message decides whether people trust it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. A signature is not provenance
&lt;/h2&gt;

&lt;p&gt;Here is the part I would want to read if someone else had written this post.&lt;/p&gt;

&lt;p&gt;What we record alongside the artifact is the commit and the signature. That is it. There is no signed attestation describing the build, no record of which runner produced it, no statement of what the pipeline configuration was at the time.&lt;/p&gt;

&lt;p&gt;So the chain supports this claim: &lt;em&gt;this artifact was signed by our build, and it says it came from this commit.&lt;/em&gt; It does not support the claim people tend to hear, which is: &lt;em&gt;this artifact was built from this commit, by this pipeline, with nothing else added.&lt;/em&gt; The commit reference is metadata the build wrote about itself. A signature over self-reported metadata proves the signer, not the statement.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxrczcuh2ke3kx3w78rqx.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxrczcuh2ke3kx3w78rqx.png" alt="Diagram" width="800" height="1389"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That is not a reason to skip signing. It raises the cost of tampering after the build, which is the most likely place for it to happen. It is a reason to be careful about the sentence you put in a compliance document. "Traceable" is true. "Provable" is not, yet.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The distinction that matters in an audit conversation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The useful move in review is to say plainly which claim the setup supports, rather than showing the diagram and letting people read the stronger one into it.&lt;/p&gt;

&lt;p&gt;Traceability is genuinely valuable: given an artifact we can get to a commit, and given a commit to a verified author, and every link is checkable. What we cannot do is prove the build did only what the pipeline file said. Closing that means attesting the build itself, which is a larger piece of work than signing ever was.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Key rotation is the failure you will actually hit
&lt;/h2&gt;

&lt;p&gt;Not an attacker. A lost key.&lt;/p&gt;

&lt;p&gt;When a developer loses their key — a replaced laptop, usually — they generate a new one, and everything signed with the old key is suddenly signed by something the verifier does not recognise. The signature is still valid; the key it points to is no longer registered.&lt;/p&gt;

&lt;p&gt;The fix is to keep old public keys rather than replace them. The parameter store holds retired keys alongside the current one, so historic signatures keep verifying while new work is signed with the new key, and revocation becomes a deliberate act for a key you believe is compromised rather than a side effect of someone getting a new machine.&lt;/p&gt;

&lt;p&gt;This is the single most useful thing I would tell anyone rolling out signing of any kind. &lt;strong&gt;Design for rotation before you design for enforcement.&lt;/strong&gt; Enforcement is a day of work. Rotation is what determines whether the control survives contact with normal life.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What it looks like when you have not planned for it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The pattern is always the same. Someone gets a new laptop, generates a fresh key, and their next few pieces of work fail verification. From their side the control is broken: they did nothing unusual and the pipeline is refusing their change. Keeping retired keys in the store fixes it, and it costs nothing except deciding once that key history is state you keep rather than state you overwrite.&lt;/p&gt;

&lt;p&gt;What still needs a human is deciding whether a key was &lt;em&gt;lost&lt;/em&gt; or &lt;em&gt;taken&lt;/em&gt;. Those need opposite responses — keep the old key so history verifies, or revoke it so it stops — and no automation tells you which one you are looking at.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. The part that is still unsolved: knowing which commits are new
&lt;/h2&gt;

&lt;p&gt;The gap I have not closed is not about signing at all. It is about deciding which commits a pipeline run is supposed to check.&lt;/p&gt;

&lt;p&gt;When a branch is pushed for the first time, the "before" reference the hook receives is all zeros — there is no previous state to compare against. You cannot tell from that alone where the new work starts. If the branch was taken from another branch rather than from the mainline, walking back from the tip picks up commits that belong to whoever wrote them, not to the person pushing now.&lt;/p&gt;

&lt;p&gt;The consequences run both ways, and both are bad:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Too wide:&lt;/strong&gt; the check covers commits from before the branch point, and a developer is blocked by history they did not write.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Too narrow:&lt;/strong&gt; the range misses commits, and unverified work reaches a protected branch while the pipeline reports green.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Comparing against the mainline merge base is the obvious answer and it is not reliable either — a branch taken from a branch has a merge base that is not where the developer thinks it is, and a rebase moves it after the fact.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;How we live with it today&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We compare against the mainline and accept that a branch taken from another branch is a case we get wrong. When it happens the developer sees a failure they cannot explain, asks in the channel, and we look at the range by hand. Calling that a solution would be generous — it is a known rough edge with a manual escape, which is the honest description of most controls at this stage.&lt;/p&gt;

&lt;p&gt;What I think the real fix looks like is recording the verified range rather than recomputing it: storing what has already been checked, so the pipeline asks &lt;em&gt;what is new since the last verification&lt;/em&gt; instead of inferring the boundary from the push. I have not built it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Sign every artifact that gets handed on&lt;/strong&gt;, not only the first one. The JAR and the image are both handoffs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the signing key reachable only from the build.&lt;/strong&gt; That restriction is the security property; everything else is plumbing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fail the deploy on a bad signature&lt;/strong&gt;, and make the message distinguish "no signature" from "unknown key".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be precise about what you have.&lt;/strong&gt; A signature over self-reported metadata gives you traceability, not attested provenance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan key rotation first.&lt;/strong&gt; Keep retired public keys so history still verifies, and treat revocation as a deliberate decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Work out how you will identify new commits&lt;/strong&gt; before you enforce anything on them. It is harder than the signing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The signature part of this was a week of work. The part I am still circling is knowing exactly which commits a given run is responsible for, and until that is solved the enforcement around it is only as trustworthy as the range it was handed.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/30/what-an-artifact-signature-proves/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/30/what-an-artifact-signature-proves/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>supplychain</category>
      <category>signing</category>
      <category>ci</category>
      <category>security</category>
    </item>
    <item>
      <title>Why our Java builds run on Fargate and our Docker builds do not</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Sun, 27 Sep 2026 13:20:34 +0000</pubDate>
      <link>https://dev.to/vivek_itp/why-our-java-builds-run-on-fargate-and-our-docker-builds-do-not-336a</link>
      <guid>https://dev.to/vivek_itp/why-our-java-builds-run-on-fargate-and-our-docker-builds-do-not-336a</guid>
      <description>&lt;p&gt;Most discussions about self-hosted CI runners start with cost. Managed minutes are expensive, the argument goes, so run your own and save money. That was not why we did it, and cost is usually the weakest reason on the list.&lt;/p&gt;

&lt;p&gt;We run our own GitLab runners because of a requirement we could not meet any other way: builds have to sign the artifacts they produce, with a key the build host is allowed to use and nobody else is. Everything below follows from that one constraint, and from the fact that not every build wants the same kind of machine.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The reason to self-host is usually a requirement, not a saving
&lt;/h2&gt;

&lt;p&gt;If your builds are ordinary, managed runners are the right answer. Somebody else patches them, scales them, and gets paged when they break. You should need a reason to give that up.&lt;/p&gt;

&lt;p&gt;Ours was artifact signing. The publish step signs the JAR with a KMS key before it reaches the artifact repository, and the deploy step refuses anything that does not verify. That only works if the machine doing the signing is one we control: our account, our IAM role, our network path to the key, and no shared tenancy with builds that should never reach it.&lt;/p&gt;

&lt;p&gt;Once you own the compute for that reason, the other arguments start to matter — private network access, predictable machine sizes, caches that survive between jobs. But they are consequences, not the case.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What actually forced the decision&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had already moved to signed commits, verified in the pipeline. The next question from the security side was the obvious one: you have proved who wrote the code, so can you prove which build produced the binary?&lt;/p&gt;

&lt;p&gt;Answering that meant the runner had to hold an identity: not a secret pasted into a CI variable, but a role with permission to use one specific key, scoped so only jobs from projects meant to sign could assume it. You cannot do that on shared managed compute in any way I was comfortable defending.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Split the compute by what the job builds, not by who owns the job
&lt;/h2&gt;

&lt;p&gt;We run everything from one centralised AWS account, with two kinds of executor behind it: ECS Fargate and EC2. The split is not by team or by environment. It is by what the job actually does.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Job type&lt;/th&gt;
&lt;th&gt;Runs on&lt;/th&gt;
&lt;th&gt;Why&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Java builds, tests, artifact publish&lt;/td&gt;
&lt;td&gt;ECS Fargate&lt;/td&gt;
&lt;td&gt;Ordinary processes, no privileged access needed, scales to zero&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Container image builds&lt;/td&gt;
&lt;td&gt;EC2&lt;/td&gt;
&lt;td&gt;Needs a real Docker daemon and a disk that behaves like a disk&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Fargate is the better default. No host to patch, capacity appears when a job appears, and nothing is running means nothing is billed. For a Maven or Gradle build that is all upside.&lt;/p&gt;

&lt;p&gt;Container builds are where it stops working. Building an image wants a daemon, layer storage that behaves like a local disk, and occasionally privileges a serverless container platform is specifically designed not to give you. You can fight that with alternative builders, and we looked at it, but EC2 hosts do this without argument and Fargate does it with a fight.&lt;/p&gt;

&lt;p&gt;So Java builds go to Fargate, Docker builds go to EC2, and the developer writes neither of those words. They pick a runner tag and the tag decides.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The rule that removed the question&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Before we split it, the recurring question in the platform channel was &lt;em&gt;"which runner should my job use?"&lt;/em&gt;, and the answers people gave each other were folklore copied from whichever project they had worked on last.&lt;/p&gt;

&lt;p&gt;One sentence fixed most of it: &lt;strong&gt;if the job builds an image it runs on the EC2 tag, and everything else runs on the Fargate tag.&lt;/strong&gt; That fits in a template comment, and nobody has to understand executor internals to get it right. The jobs that still get it wrong are the ones that do both, which is usually a sign the job is doing too much anyway.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Isolation is a decision you make once, early, and live with
&lt;/h2&gt;

&lt;p&gt;Each executor type runs as two stacks, and each stack has two sets of runners separated by a restricted flag. Restricted runners can reach the signing key and the production-facing artifact paths. Unrestricted ones cannot, and that is enforced by the role attached to the runner, not by anything in the project's pipeline file.&lt;/p&gt;

&lt;p&gt;The reason for two sets rather than one is simple: if every runner can assume the signing role, every project in the instance can sign, including the one someone created this morning to try something out. Permission on a runner is permission for whoever can schedule a job onto it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fboaz9wzlmgxk8q6taovp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fboaz9wzlmgxk8q6taovp.png" alt="Diagram" width="720" height="2708"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The thing I would tell anyone starting this: decide the isolation boundary before you have twenty projects on the runners, not after. Adding a project to the restricted set is a five-minute change. Taking access away from projects that have quietly come to depend on it is a quarter of conversations.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the flag sits on the runner and not in the pipeline&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first version had the signing step guarded by a check inside the pipeline template: if the project was on the approved list, the step ran. That is a check a project can edit its way around, because the pipeline file lives in the project's own repository. The list was in the right place for convenience and the wrong place for security.&lt;/p&gt;

&lt;p&gt;Moving the boundary onto the runner's IAM role changed the failure mode. A project that is not supposed to sign does not fail a policy check. It gets an access-denied from AWS, because the machine it runs on genuinely cannot use the key.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. For critical projects, the pipeline configuration is not the team's to change
&lt;/h2&gt;

&lt;p&gt;Alongside the tagging, we run what we call guarded CI. For a short list of critical projects, the pipeline configuration is hardcoded: the runner tags, the signing step, and the verification step are fixed, and the project cannot override them from its own file.&lt;/p&gt;

&lt;p&gt;This is a deliberate reduction in flexibility and it is unpopular in exactly the way you would expect. The argument for it is that the controls on a critical application should not be removable by editing a YAML file in that application's own repository. If a team can turn off the check that protects them, the check is a suggestion.&lt;/p&gt;

&lt;p&gt;The cost is real: guarded projects wait on us for changes other teams make themselves. We keep the list short so that stays survivable.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where guarding earns its keep&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The moment that justified it was mundane. A build was failing on the verification step, and the quickest path to green — the one anybody under pressure reaches for — was to comment the step out and come back to it later.&lt;/p&gt;

&lt;p&gt;On a normal project that works, and sometimes nobody notices for a month. On a guarded project it does nothing, because the step does not come from the project's file. The team pinged us instead, and the cause was a stale key reference that took minutes to fix. Being unable to bypass the check is what got the right people looking at it the same day.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Caching is where self-hosted runners give the time back
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;The specifics here are the pattern I would recommend rather than a story I am retelling exactly.&lt;/em&gt; A Java build that resolves dependencies from scratch every time spends a meaningful chunk of every run redoing work. Self-hosted runners let you fix that, and it is the clearest developer-visible win available.&lt;/p&gt;

&lt;p&gt;Two things matter more than the rest:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A shared dependency cache in S3&lt;/strong&gt;, keyed on the lockfile, so a build only pays full resolution cost when dependencies actually change.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Warm layer storage on the EC2 hosts&lt;/strong&gt; for image builds, so a rebuild of an application whose base image has not changed reuses the layers it already has.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The trap is treating the cache as free. A cache that is never invalidated eventually serves something stale and costs you an afternoon, and a cache keyed too loosely is a correctness problem wearing a performance costume. Key it on the lockfile, expire it on a schedule, and make cache hits and misses obvious in the job log.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The shape of the win&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The pattern to expect is that the first build after a dependency change pays full price and every build after it takes noticeably less. What makes it worth doing is the common case rather than the average: a small code change on a project whose dependencies have not moved in weeks gets fast, and that is where developers spend almost all their time.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. The maintenance is the part nobody puts in the estimate
&lt;/h2&gt;

&lt;p&gt;&lt;em&gt;Also a recommendation rather than a specific recollection.&lt;/em&gt; Building the runner setup gets estimated. The ongoing cost usually does not, and it is not small:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Patching the EC2 hosts, and rotating them often enough that patching is routine rather than an event.&lt;/li&gt;
&lt;li&gt;Keeping the runner version in step with the GitLab instance, because drift produces failures that look like project problems.&lt;/li&gt;
&lt;li&gt;Watching the autoscaling behaviour, because a capacity problem arrives at the worst moment and looks like a slow pipeline.&lt;/li&gt;
&lt;li&gt;Telling platform failures apart from project failures quickly enough that the distinction is useful.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Budget it as a standing share of somebody's time, not as a project with an end date. If nobody owns it, the runners keep working right up until the day they very loudly do not.&lt;/p&gt;

&lt;h2&gt;
  
  
  7. What developers actually notice
&lt;/h2&gt;

&lt;p&gt;Not the architecture, and not the isolation model. They notice that a build sometimes takes longer to start than it should, and that the reason is invisible from where they sit.&lt;/p&gt;

&lt;p&gt;That is a queueing effect and we have not solved it. Concurrency limits are an awkward trade: set them high and a busy afternoon has several heavy builds fighting over the same hosts, so everything slows at once. Set them low and jobs queue while capacity sits idle, because the limit rather than the machine is the constraint. We are still tuning it.&lt;/p&gt;

&lt;p&gt;The mitigation that helped most was not a capacity change. It was making the wait visible: if a job is queued rather than running, a developer should see that without asking, because &lt;em&gt;queued&lt;/em&gt; and &lt;em&gt;slow&lt;/em&gt; are different problems and only one of them is theirs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-host for a requirement you cannot meet otherwise&lt;/strong&gt;, not for a saving. Ours was signing artifacts with a key only our build hosts can use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Split executors by what the job builds.&lt;/strong&gt; Fargate for ordinary builds that scale to zero, EC2 for image builds that want a real daemon and a real disk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the isolation boundary on the runner's role&lt;/strong&gt;, not in a pipeline file the project can edit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guard the configuration for critical projects&lt;/strong&gt;, keep that list short, and accept that it costs those teams some autonomy.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache deliberately&lt;/strong&gt;, key on the lockfile, and make hits and misses visible in the log.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget maintenance as standing time.&lt;/strong&gt; Patching, version drift and scaling behaviour do not stop arriving.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I have not worked out is the concurrency question. Every setting we have tried is a choice about who waits and when, and the right answer probably changes with the time of day in a way a single number cannot express.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/27/why-java-builds-run-on-fargate/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/27/why-java-builds-run-on-fargate/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ci</category>
      <category>gitlab</category>
      <category>aws</category>
      <category>runners</category>
    </item>
    <item>
      <title>Measuring developer experience without a survey</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Wed, 23 Sep 2026 22:55:37 +0000</pubDate>
      <link>https://dev.to/vivek_itp/measuring-developer-experience-without-a-survey-5hj4</link>
      <guid>https://dev.to/vivek_itp/measuring-developer-experience-without-a-survey-5hj4</guid>
      <description>&lt;p&gt;The usual first move when someone decides to take developer experience seriously is to send a survey. It is the obvious thing to do and it is rarely the most useful thing available, because the data you actually need is already sitting in systems you run.&lt;/p&gt;

&lt;p&gt;I am not against surveys. They capture things instrumentation cannot: whether people feel trusted, whether they think anyone is listening, whether the platform is something they use or something that happens to them. But as a starting point they have three problems, and all three are avoidable.&lt;/p&gt;

&lt;p&gt;They are answered by the people who already engage. They measure sentiment at a moment rather than friction over time. And they take weeks to run, which is long enough for the thing that annoyed everyone to be forgotten or fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers are already there
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What to measure&lt;/th&gt;
&lt;th&gt;Where it already lives&lt;/th&gt;
&lt;th&gt;What a bad number means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Time from push to a merge signal&lt;/td&gt;
&lt;td&gt;Your CI system&lt;/td&gt;
&lt;td&gt;People context-switch while waiting&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Re-run rate on checks&lt;/td&gt;
&lt;td&gt;Your CI system&lt;/td&gt;
&lt;td&gt;Nobody trusts red any more&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time from repository creation to first production deploy&lt;/td&gt;
&lt;td&gt;Repo events and deploy logs&lt;/td&gt;
&lt;td&gt;Onboarding a service costs days&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Support requests per team per month&lt;/td&gt;
&lt;td&gt;Your support channel&lt;/td&gt;
&lt;td&gt;The platform cannot answer for itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time from merge to production&lt;/td&gt;
&lt;td&gt;Deploy logs&lt;/td&gt;
&lt;td&gt;Shipping is a decision, not a default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rollback frequency and time to roll back&lt;/td&gt;
&lt;td&gt;Deploy logs&lt;/td&gt;
&lt;td&gt;Either quality, or fear of deploying&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;None of those requires a new vendor. They require someone to go and count, and then to keep counting.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Pick numbers that describe a developer's day
&lt;/h2&gt;

&lt;p&gt;The test for a developer experience metric is whether a developer would recognise it as describing their own week. Deployment frequency across the organisation does not. Waiting eleven minutes for a pipeline does.&lt;/p&gt;

&lt;p&gt;That rules out most of what gets reported upward. Aggregate throughput numbers are useful to someone, but they average away exactly the thing you are trying to see: that one team waits forty minutes and another waits four, and the average of twenty-two describes nobody.&lt;/p&gt;

&lt;p&gt;Keep the numbers per team, or per service, and look at the distribution rather than the mean. The interesting signal is almost always in the tail.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first number I tracked properly was time from push to a merge signal, per repository rather than overall. The organisation-wide figure had looked acceptable and had told me nothing.&lt;/p&gt;

&lt;p&gt;Broken out per repository, the picture was completely different. A small number of repositories were far slower than the rest, and they belonged to the teams who complained most. The complaint and the data had been saying the same thing for a while; only one of them had been legible.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Re-run rate is the most honest number you have
&lt;/h2&gt;

&lt;p&gt;If I could keep only one, it would be this: how often does someone re-run a check without changing anything.&lt;/p&gt;

&lt;p&gt;Nobody re-runs a check they believe. A high re-run rate means the pipeline has stopped carrying information, and that the team has learned to treat red as noise. It is the clearest possible measurement of trust, and it is free, because your CI system already records it.&lt;/p&gt;

&lt;p&gt;It is also a leading indicator. Trust erodes before anything visible breaks. By the time a real failure gets re-run three times and merged, the erosion happened months earlier.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What the re-run rate exposed&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Counting re-runs turned a vague sense that the pipeline was flaky into a short, ordered list of the specific jobs people did not believe. The same few appeared over and over.&lt;/p&gt;

&lt;p&gt;That list was more actionable than any survey response about pipeline confidence would have been, because it named the jobs rather than the feeling.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Measure the whole path, not your part of it
&lt;/h2&gt;

&lt;p&gt;A platform team naturally measures the parts it owns, which is exactly how the slow parts stay invisible. The queue you control looks fast because it is fast. The wait that hurts is somewhere between two systems, or in another team's inbox.&lt;/p&gt;

&lt;p&gt;Measure end to end, from the developer's first action to the outcome they wanted, and treat every handoff as part of the number. If getting a namespace takes two days in someone else's queue, that belongs in time to first deploy whether or not you own the queue.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frpxx67l0y5vn4heqjv6l.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Frpxx67l0y5vn4heqjv6l.png" alt="Diagram" width="800" height="100"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The dotted spans are the measurements. Both cross boundaries the platform team does not own, and that is the point: they describe the developer's experience rather than the platform's.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The wait that was not in our numbers&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I walked the new-service path myself and timed it, almost all of the elapsed time was waiting in queues belonging to other teams. Every system I owned was fast, and the developer's experience was still days long.&lt;/p&gt;

&lt;p&gt;Our dashboards had been honest about our part and silent about the total, which made them reassuring and useless.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Count the questions
&lt;/h2&gt;

&lt;p&gt;Support volume is a developer experience metric and almost nobody treats it as one. Every question is a moment where the platform failed to answer for itself.&lt;/p&gt;

&lt;p&gt;You do not need tooling for this. A tally of what gets asked, bucketed roughly and counted, will show you a distribution steep enough to act on within a few weeks. The top few buckets are the platform's real backlog.&lt;/p&gt;

&lt;p&gt;Watch the trend rather than the total. A channel that gets quieter each quarter is a platform that is learning. A channel with a stable volume and a fast response time is a platform that has got good at answering the same questions forever.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Trends, not benchmarks
&lt;/h2&gt;

&lt;p&gt;The question worth asking is whether a number is better than it was last quarter. The question that wastes time is whether it is better than some other company's number.&lt;/p&gt;

&lt;p&gt;Published benchmarks come from organisations with different constraints, different regulatory positions, and different definitions of the same word. Comparing against them produces either false comfort or a target nobody can reach for reasons that have nothing to do with effort.&lt;/p&gt;

&lt;p&gt;Your own history is the only honest comparison. It shares every confounding factor, which is exactly what makes it useful.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Do not turn them into targets
&lt;/h2&gt;

&lt;p&gt;The moment a number becomes a target, it stops measuring what it measured. This is not a new observation and platform teams fall into it anyway, usually with the best intentions and a quarterly goal attached.&lt;/p&gt;

&lt;p&gt;If re-run rate becomes a target, the tempting fix is to remove the flaky tests rather than fix them. If time to first deploy becomes a target, the fix is to redefine when the clock starts. Both improve the number and neither improves anyone's day.&lt;/p&gt;

&lt;p&gt;Keep them as instruments. They tell you where to look and whether something you changed helped. They are not a scoreboard, and the moment anyone's performance review touches them, they are finished as measurements.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where I would put the numbers&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Putting the measurements where developers could see them, rather than only in a report going upward, changed how they were received. Teams recognised their own figures and occasionally disputed them, which was useful, because a disputed number gets examined.&lt;/p&gt;

&lt;p&gt;The times these numbers were least useful were when they appeared in a summary for people who did not work with the platform daily. At that altitude they stop being diagnostic and start being a score.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Then run the survey
&lt;/h2&gt;

&lt;p&gt;Everything above measures friction. It cannot tell you whether people feel they can get things changed, whether they trust the platform team, or whether the documentation is any good.&lt;/p&gt;

&lt;p&gt;So run a survey, but run it second, when you already know where the friction is. The questions get sharper, the responses are easier to interpret, and you can ask about the specific things your numbers flagged rather than asking people to rate the developer experience out of ten.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Start with what you already record.&lt;/strong&gt; Your CI system and deploy logs hold most of it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per team, not aggregate,&lt;/strong&gt; and look at the distribution. The average describes nobody.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run rate is the single most honest number&lt;/strong&gt; and it costs nothing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure end to end,&lt;/strong&gt; including the queues you do not own.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count support questions&lt;/strong&gt; and watch the trend, not the total.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compare against your own history,&lt;/strong&gt; never a published benchmark.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Never make them targets.&lt;/strong&gt; They are instruments.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run the survey second,&lt;/strong&gt; to ask about the things instrumentation cannot see.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still do not know how to measure is the work that never gets attempted. Some fraction of good ideas die because someone correctly judges that the platform would make them painful, and none of those decisions leave a trace in any system I run. Everything above measures the friction people pushed through, not the friction that stopped them.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/23/measuring-devex-without-a-survey/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/23/measuring-devex-without-a-survey/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devex</category>
      <category>metrics</category>
      <category>platformengineering</category>
    </item>
    <item>
      <title>What migrating logging from Datadog to Dynatrace actually cost us</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Mon, 21 Sep 2026 23:38:56 +0000</pubDate>
      <link>https://dev.to/vivek_itp/what-migrating-logging-from-datadog-to-dynatrace-actually-cost-us-4lcg</link>
      <guid>https://dev.to/vivek_itp/what-migrating-logging-from-datadog-to-dynatrace-actually-cost-us-4lcg</guid>
      <description>&lt;p&gt;Every platform team eventually has to take something away. The technical part is usually straightforward. What goes wrong is everything between announcing it and switching it off, and the cost almost never lands where you expect.&lt;/p&gt;

&lt;p&gt;The clearest example I have is moving our logging from Datadog to Dynatrace. The driver was ordinary: cost, plus consolidating logs and metrics onto one platform instead of paying two vendors to hold two halves of the same picture.&lt;/p&gt;

&lt;p&gt;Before anything else: this is not a comparison of the two products, and nothing here is a complaint about either. Both do the job. The decision was commercial, the way most tooling decisions are, and what follows is about the migration rather than the tools. I have kept the internal specifics deliberately light.&lt;/p&gt;

&lt;p&gt;I expected the hard part to be the applications. It was not. The applications barely noticed. What actually cost us was the monitoring built on top of the logs, and the fact that everyone who used logs daily had to relearn how to ask a question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I expected to be hard, and what was
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Expected to be hard&lt;/th&gt;
&lt;th&gt;Actually hard&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Changing every application&lt;/td&gt;
&lt;td&gt;The forwarder changed; apps mostly did not&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Getting teams to prioritise the work&lt;/td&gt;
&lt;td&gt;It rode along with their normal deployments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Log data itself&lt;/td&gt;
&lt;td&gt;Custom masking rules tied to the old vendor&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nothing in particular&lt;/td&gt;
&lt;td&gt;Every monitor and alert built on the old logs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nothing in particular&lt;/td&gt;
&lt;td&gt;Teams learning to query in the new tool&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The two bottom rows are the whole story. Neither was on my list at the start.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Find the thing that is not in the obvious place
&lt;/h2&gt;

&lt;p&gt;On paper the change was small. Logs went to a forwarder, and the forwarder needed to point somewhere new. No application code, no library swap, no redeploy required to change a destination.&lt;/p&gt;

&lt;p&gt;What was not on paper was everything that had accumulated inside that forwarder. Over the years it had picked up custom masking rules, written against the old vendor's configuration format, quietly redacting sensitive values before logs left our systems. They were not documented as a dependency on the vendor. They were just part of how logging worked.&lt;/p&gt;

&lt;p&gt;That is the shape of the hidden dependency in most deprecations. Not the integration everyone knows about, but the small accretions around it that nobody wrote down because they were never a decision, only a series of fixes.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The forwarder swap itself was an afternoon of work. The masking rules were not, because they had to be reproduced exactly in a different configuration language, and "exactly" is the operative word when the rules exist to keep sensitive values out of a log store.&lt;/p&gt;

&lt;p&gt;Going through them one at a time was tedious and it was also the only responsible option. That was the point where the migration stopped being a config change and became a piece of work with a real review attached.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Run both, and validate, before removing anything
&lt;/h2&gt;

&lt;p&gt;The decision that made everything else possible was refusing to cut over. For a couple of months, logs flowed to both platforms at once. Nothing was removed while we compared.&lt;/p&gt;

&lt;p&gt;Dual shipping costs money, briefly, and it buys the only thing that makes removal safe: evidence. Not a belief that the new pipeline works, but a side-by-side you can point at. Same volumes, same fields, same masking applied, same events present in both.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zmrgcqcixlno732rhbd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8zmrgcqcixlno732rhbd.png" alt="Diagram" width="798" height="172"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The loop in the middle is what a validation period is for. Without it, "are we ready to switch off" is a judgement call made under time pressure. With it, the answer is a comparison anybody can check.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What the overlap actually caught&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The overlap did not produce a dramatic discovery, and I count that as the period doing its job rather than being unnecessary. It let us confirm that masking was being applied the same way on both sides, which was the thing I was least willing to guess about.&lt;/p&gt;

&lt;p&gt;Had we cut over directly and been wrong about that, we would have found out from the wrong direction entirely.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Let it ride on work teams are already doing
&lt;/h2&gt;

&lt;p&gt;We did not run this as a migration project with a deadline and a chase list. Teams picked up the change as part of their normal feature deployments. Whatever they were shipping next carried the new configuration with it.&lt;/p&gt;

&lt;p&gt;That choice is worth being honest about, because it trades one problem for another.&lt;/p&gt;

&lt;p&gt;What you gain is real. There is no separate piece of work competing with a team's roadmap, no deadline anyone has to defend, and no incentive to rush. The change arrives with something the team wanted to ship anyway.&lt;/p&gt;

&lt;p&gt;What you give up is predictability. A team that does not deploy for six weeks has not migrated for six weeks, and the tail is as long as your slowest-moving service. You cannot put a date on the board and be confident about it.&lt;/p&gt;

&lt;p&gt;I think it was the right call here, specifically because dual shipping made a long tail harmless. Nothing was breaking while we waited. Take away the overlap and the same approach becomes a slow-motion outage waiting for the least active team.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;No chase list&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The thing I noticed most was the absence of the usual conversation. Nobody had to be persuaded, because nobody was being asked to stop what they were doing.&lt;/p&gt;

&lt;p&gt;The trade shows up at the other end. The final stretch is a small number of services that simply had not deployed recently, and those needed individual conversations rather than a broadcast.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. The applications were fine. The monitoring was not
&lt;/h2&gt;

&lt;p&gt;Here is the part I would tell anyone planning a similar move: the migration is not the logs, it is everything built on top of them.&lt;/p&gt;

&lt;p&gt;Monitors, alerts and dashboards are written against a specific query language, specific field names, and a specific way the platform structures data. None of that survives a vendor change automatically. Every one of them has to be rebuilt, tested, and trusted again.&lt;/p&gt;

&lt;p&gt;That work is invisible in the plan, because nobody thinks of an alert as an integration. It is also the work where getting it wrong is worst: an alert that silently stops firing is far more dangerous than a log line that fails to arrive, and it fails quietly by definition.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Rebuilding the things nobody counted&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The monitors were the bulk of the real effort and were not what I had scoped at the start. Each one had to be re-expressed in the new platform and then actually verified, because the only way to trust an alert is to make it fire.&lt;/p&gt;

&lt;p&gt;If I ran this again, I would treat monitor migration as the main workstream from day one, and the forwarder change as the small task it turned out to be.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. The real cost was people relearning how to ask
&lt;/h2&gt;

&lt;p&gt;The last thing, and the one I keep thinking about, is that the hardest part of this was not technical at all.&lt;/p&gt;

&lt;p&gt;People had years of muscle memory in the old tool. They knew how to get from an alert to the relevant log line in three moves, because they had done it hundreds of times. The new platform can answer all the same questions and it answers them differently, and that difference lands during incidents, when nobody has patience for learning.&lt;/p&gt;

&lt;p&gt;That is a genuine cost of any tooling migration and it usually goes unnamed, because it does not appear as a ticket or an outage. It appears as things taking longer for a while, and as a quiet preference for the old tool for as long as the old tool still exists.&lt;/p&gt;

&lt;p&gt;What would have helped is treating query fluency as part of the migration rather than an afterthought: a short page of the ten queries people actually run, translated, and a channel where asking how to express something was normal rather than an admission.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Leave a tombstone
&lt;/h2&gt;

&lt;p&gt;When the old destination finally went away, what mattered was that anyone who went looking found an explanation rather than silence. What was removed, when, where the data lives now, and how to ask the question they were trying to ask.&lt;/p&gt;

&lt;p&gt;Keep it up longer than feels necessary. The person who needs it is the one who only looks at logs when something is badly wrong, and they will arrive months later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Look for the accretions,&lt;/strong&gt; not the integration. The forwarder was easy; the masking rules written for the old vendor were not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run both and compare&lt;/strong&gt; before removing anything. The overlap buys evidence rather than confidence.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Riding on normal deployments avoids the chase list&lt;/strong&gt; and gives up predictability. It is only safe if nothing breaks while you wait.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scope the monitors as the main work,&lt;/strong&gt; not the pipeline. An alert that stops firing fails silently.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the relearning cost.&lt;/strong&gt; Query fluency is part of the migration, and it is paid during incidents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Leave an explanation behind&lt;/strong&gt; for the person who shows up months later.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still have no good answer for is the deprecation with no replacement, where the honest message is that a capability is going away and teams will have to do without it. Everything above works because there was somewhere to go. When there is not, the conversation is about priorities rather than migration, and I have never found a way to make that one go smoothly.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/21/migrating-logging-datadog-to-dynatrace/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/21/migrating-logging-datadog-to-dynatrace/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>platformengineering</category>
      <category>deprecation</category>
      <category>observability</category>
      <category>migration</category>
    </item>
    <item>
      <title>Local development environments that match production closely enough</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Mon, 21 Sep 2026 23:38:51 +0000</pubDate>
      <link>https://dev.to/vivek_itp/local-development-environments-that-match-production-closely-enough-2pnh</link>
      <guid>https://dev.to/vivek_itp/local-development-environments-that-match-production-closely-enough-2pnh</guid>
      <description>&lt;p&gt;Two failure modes, and I have caused both. The first is a local environment so far from production that "it works on my machine" stops being a joke and becomes a daily fact. The second is chasing full parity until starting the thing locally takes fifteen minutes, eats most of a laptop, and developers quietly stop using it.&lt;/p&gt;

&lt;p&gt;Neither is a discipline problem. They are both the result of not deciding, explicitly, which differences between a laptop and production are allowed to exist.&lt;/p&gt;

&lt;p&gt;I have written before about local and CI needing to run the same commands. This is the neighbouring question and a harder one: how close does a developer's running system need to be to the real one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Some differences matter and most do not
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;th&gt;Does it change behaviour?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Database major version&lt;/td&gt;
&lt;td&gt;Yes. Syntax, query planning, and migrations all move&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runtime version&lt;/td&gt;
&lt;td&gt;Yes. This is where most "works locally" bugs live&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One replica instead of forty&lt;/td&gt;
&lt;td&gt;No, until you are debugging concurrency&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Real payment provider vs a fake&lt;/td&gt;
&lt;td&gt;No, if the fake matches the contract&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No TLS locally&lt;/td&gt;
&lt;td&gt;Usually not, until a client library behaves differently&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ten rows instead of ten million&lt;/td&gt;
&lt;td&gt;Not for correctness. Absolutely for performance work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Different operating system&lt;/td&gt;
&lt;td&gt;Sometimes. Path handling, file watching, case sensitivity&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The right-hand column is the whole exercise. Parity is not a slider you push toward one hundred percent. It is a set of individual decisions, each of which should have a reason.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Match the contract, not the infrastructure
&lt;/h2&gt;

&lt;p&gt;The thing a service needs from its dependencies is the contract: the same interface, the same error shapes, the same semantics. It almost never needs the same topology, the same scale, or the same managed service.&lt;/p&gt;

&lt;p&gt;A local queue that delivers messages at least once, out of order, and occasionally redelivers is a good stand-in for the production one, whatever runs underneath. A local queue that delivers exactly once and in order is worse than useless, because it lets people write code that cannot work in production and gives them confidence while they do it.&lt;/p&gt;

&lt;p&gt;This is the test I use for a fake: does it fail in the same ways as the real thing? If not, it is teaching the wrong lessons.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had a local stand-in for a message broker that was much friendlier than the real one. It delivered in order, never redelivered, and never dropped anything.&lt;/p&gt;

&lt;p&gt;Code written against it worked locally and then behaved oddly in production, where duplicates and reordering were normal. The stand-in had been chosen because it was easy to run, and the ways in which it was easier were exactly the ways that mattered. Replacing it with something that redelivered and reordered made local development slightly more annoying and caught a category of bug before it shipped.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Pin the things that change behaviour
&lt;/h2&gt;

&lt;p&gt;Runtime version and database major version are the two that have caused me the most confusion, and both are cheap to fix. Pin them in the repository, read the same pin everywhere, and the whole class of problem disappears.&lt;/p&gt;

&lt;p&gt;The rule I use: if a difference can change the result of running the code, it goes in a file in the repository, and both the laptop and the pipeline read that file. If it cannot, it does not need to match.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# one place, read by the local tooling and by CI&lt;/span&gt;
&lt;span class="na"&gt;runtime&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3.12.4"&lt;/span&gt;
&lt;span class="na"&gt;database&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;postgres:16.3"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What does not work is documenting the required versions in a README. Nobody re-reads a README after their first week, and versions drift silently from the moment they are written down.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The version that drifted&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The most time-consuming local problem I dealt with was not dramatic: a developer's database was a major version behind what production ran, because they had installed it a year earlier and nothing had ever told them to change it.&lt;/p&gt;

&lt;p&gt;A migration worked on their machine and failed in the pipeline with an error that pointed at the migration rather than the version. Pinning the database image and having the local tooling refuse to start on a mismatch removed that whole category, and the refusal message did more good than any documentation would have.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Decide what you are deliberately not matching
&lt;/h2&gt;

&lt;p&gt;The parts that are not worth matching should be an explicit, written decision rather than an accident. Scale, multi-region behaviour, real third-party accounts, production data. Each of those costs a great deal to reproduce and buys very little for the work a developer does on a normal day.&lt;/p&gt;

&lt;p&gt;Writing them down matters more than it sounds. When a developer hits a bug that only appears at scale, the useful response is "yes, local does not cover that, here is the environment that does" rather than an argument about whether local should have caught it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Funm2x41spou3lypaavha.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Funm2x41spou3lypaavha.png" alt="Diagram" width="800" height="894"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The branch on the right is where most local setups go wrong. They fake something, the fake is friendlier than reality, and nobody notices until production disagrees.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Start-up time is a parity feature
&lt;/h2&gt;

&lt;p&gt;An environment that takes fifteen minutes to start is not a high-parity environment. It is an environment nobody runs, which means its parity is irrelevant.&lt;/p&gt;

&lt;p&gt;Every service added to the local stack has a cost paid by every developer, every day, forever. That cost is rarely weighed against the benefit, because the benefit is visible and the cost is spread thin.&lt;/p&gt;

&lt;p&gt;The questions worth asking before adding anything to the local stack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Does a developer need it running to do their normal work, or only occasionally?&lt;/li&gt;
&lt;li&gt;Can it start lazily, when something actually calls it?&lt;/li&gt;
&lt;li&gt;Can it be a fake that matches the contract rather than the real service?&lt;/li&gt;
&lt;li&gt;What does it add to start-up time, and is that worth it every single day?&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Cutting the stack down&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our local environment had grown to include everything a service might talk to, because each addition had been individually reasonable. Starting it had become something you did while making coffee, and a lot of people had stopped doing it at all.&lt;/p&gt;

&lt;p&gt;Splitting it into a small default set that starts quickly, plus optional extras you opt into when you need them, was not technically interesting and made more difference than anything else on this list. Most developers, most days, needed a fraction of what we had been starting for them.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Make the difference visible, not surprising
&lt;/h2&gt;

&lt;p&gt;When local and production do differ, the developer should be able to see it rather than discover it through a failure.&lt;/p&gt;

&lt;p&gt;A short banner at start-up listing what is real and what is faked costs nothing and prevents a lot of confusion. "Payments: fake. Queue: local, redelivers. Database: postgres 16.3, schema current. Search: not running." A developer who sees that line does not spend an hour wondering why a payment did not arrive.&lt;/p&gt;

&lt;p&gt;The same applies to data. If the local database has a small generated dataset, say so, ideally with a note about what it does not contain. Silence gets interpreted as "this is like production", and it never is.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Test the environment itself
&lt;/h2&gt;

&lt;p&gt;A local environment is a piece of software that the whole team depends on, and it usually has no tests at all. It rots quietly: a fake drifts from the real service's contract, a pin goes stale, a start-up script depends on something that was removed.&lt;/p&gt;

&lt;p&gt;A small scheduled job that starts the local environment from scratch on a clean machine and runs the smoke tests against it catches this early. When it breaks, it breaks for one person who can fix it, rather than for whoever next tries to onboard.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The setup that only worked if you already had it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our setup instructions worked perfectly for everyone who had run them before, and reliably failed for new joiners. The scripts had come to depend on tools and state that existing machines happened to have.&lt;/p&gt;

&lt;p&gt;Nobody had noticed because nobody started from scratch. Running the whole setup on a clean machine on a schedule surfaced it immediately, and onboarding stopped being a two-day exercise in asking colleagues what else they had installed.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Parity is a set of decisions, not a percentage.&lt;/strong&gt; For each difference, ask whether it can change the result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Match the contract, not the infrastructure.&lt;/strong&gt; A fake that is friendlier than reality teaches the wrong lessons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pin what changes behaviour&lt;/strong&gt; in the repository, read by both the laptop and the pipeline. Never in a README.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write down what you deliberately do not match,&lt;/strong&gt; so the gap is a known limit rather than an argument.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Guard start-up time.&lt;/strong&gt; An environment nobody runs has no parity at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Print what is real and what is faked&lt;/strong&gt; at start-up.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the environment on a clean machine&lt;/strong&gt; on a schedule, or it rots.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The one I keep going back and forth on is production data. Realistic data catches a whole class of bug that generated data never will, and every approach I have used to get it locally, whether subsetting, masking, or anonymising, has either leaked something it should not have or been so lossy that it stopped being realistic. I do not have a good answer there.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/21/local-environments-close-enough-to-production/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/21/local-environments-close-enough-to-production/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>localdevelopment</category>
      <category>devex</category>
      <category>tooling</category>
    </item>
    <item>
      <title>The support channel is your best product research</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:44:38 +0000</pubDate>
      <link>https://dev.to/vivek_itp/the-support-channel-is-your-best-product-research-4gfo</link>
      <guid>https://dev.to/vivek_itp/the-support-channel-is-your-best-product-research-4gfo</guid>
      <description>&lt;p&gt;Platform teams are usually bad at user research. We do not run interviews, we do not have a researcher, and the survey we send once a year gets answered by the people who already like us. Meanwhile the same team runs a support channel that receives dozens of real, unprompted, specific accounts of the platform failing its users, every single day.&lt;/p&gt;

&lt;p&gt;We just do not read it that way. We answer the question, mark it done, and move on. The answer is the least valuable thing in the exchange. The question is the data.&lt;/p&gt;

&lt;p&gt;I changed how I worked when I started treating the channel as a research feed instead of a queue. This post is what that looked like.&lt;/p&gt;

&lt;h2&gt;
  
  
  The exchange has two halves and we keep the wrong one
&lt;/h2&gt;

&lt;p&gt;A support request has a visible half and an invisible half.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Visible&lt;/th&gt;
&lt;th&gt;Invisible&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;"How do I get logs for my staging pod?"&lt;/td&gt;
&lt;td&gt;They looked for the answer and could not find it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Can someone re-run my pipeline?"&lt;/td&gt;
&lt;td&gt;They do not have permission, or do not know they do&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Is the registry down?"&lt;/td&gt;
&lt;td&gt;The error they saw did not say what was wrong&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"What is the right base image?"&lt;/td&gt;
&lt;td&gt;There is no obvious default, or there are three&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;"Who owns the deploy for service X?"&lt;/td&gt;
&lt;td&gt;Ownership is not discoverable from the tooling&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The visible half gets answered in two minutes and helps one person. The invisible half, if you collect it, tells you what to build. Every single question is a place where the platform failed to answer for itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Count the questions before you answer them
&lt;/h2&gt;

&lt;p&gt;The cheapest useful thing a platform team can do is keep a tally. Not a ticketing system, not a taxonomy. A tally.&lt;/p&gt;

&lt;p&gt;Every time a question comes in, put it in a bucket. Reuse buckets aggressively; if it is roughly the same question, it is the same bucket. After a few weeks the distribution will be extremely uneven, and the top few buckets will account for most of the volume.&lt;/p&gt;

&lt;p&gt;That list is your backlog, ordered by how much pain each item causes, measured in the only currency that matters: how often a human had to ask another human.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I started with a text file and a habit. Every question I answered got a line. No categories decided up front, just a short description, and when something repeated I merged it into the existing line and added a tick.&lt;/p&gt;

&lt;p&gt;After about a month there were maybe thirty lines, and the top five accounted for most of the ticks. None of the top five was on our roadmap. All five were small. That file changed what we worked on for the next quarter more than any planning session did.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. The repeat question is the bug
&lt;/h2&gt;

&lt;p&gt;A question asked once is a person having a bad day. The same question asked twenty times is a defect in the platform, and it should be treated exactly like one: reproduced, root-caused, fixed, and verified.&lt;/p&gt;

&lt;p&gt;The trap is that answering is so cheap. Two minutes, a link, done. Twenty times two minutes is not the real cost, though. The real cost is the twenty people who were blocked until someone was available, plus the ones who did not ask and worked around it instead.&lt;/p&gt;

&lt;p&gt;When a question repeats, the fix is almost never documentation. Documentation is where answers go to be not found. The fix is usually to make the question impossible: a better default, a clearer error, a link in the place they were already looking.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The question I answered too many times&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The single most repeated question I dealt with was some version of "how do I see the logs for my service in staging". I had a good answer. I gave it dozens of times. I linked the documentation page, which existed and was correct.&lt;/p&gt;

&lt;p&gt;What finally stopped it was not a better page. It was putting a direct link to that service's logs into the deploy notification, so the answer arrived before the question. The documentation page did not change at all. The questions stopped almost entirely.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Watch for the questions that are not questions
&lt;/h2&gt;

&lt;p&gt;Some of the most valuable signals in a support channel are not phrased as requests. They are asides, complaints, and jokes.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Ha, the pipeline is red again, business as usual."&lt;/li&gt;
&lt;li&gt;"I just re-run it twice, works eventually."&lt;/li&gt;
&lt;li&gt;"We stopped using that, it was easier to do it ourselves."&lt;/li&gt;
&lt;li&gt;"Do not worry, everyone knows you have to do X first."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of those is asking for help, so none of them gets logged anywhere. Each one describes a workaround that has become normal. A workaround that has become normal is a feature the platform failed to provide, and nobody will ever file a request for it because they have stopped expecting it to change.&lt;/p&gt;

&lt;p&gt;I started keeping these in the same tally as the questions. They were often more useful, because they described problems people had already given up on.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The joke that was a roadmap item&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Someone joked in the channel that the correct way to deploy was to "press the button twice and go get a coffee". It got a few laughs, including mine, and then I realised everyone in the thread knew exactly what it meant. The first deploy attempt failed often enough that re-running had become the accepted procedure.&lt;/p&gt;

&lt;p&gt;Nobody had ever reported it, because it was not broken in a way you could report. It just did not work the first time, reliably enough that people had built a ritual around it. That went into the tally and turned into real work.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Answer in public, always
&lt;/h2&gt;

&lt;p&gt;A support question answered in a direct message helps one person and produces no data. The same question answered in the channel helps everyone watching, gets found by search later, and stays in your tally.&lt;/p&gt;

&lt;p&gt;Push everything into the open channel, politely and consistently. When someone asks privately, answer in the channel and link them. Not to shame anyone, but because the private answer is a small gift to one person and a loss to everyone else, including you.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg18x6wr6x8evzjr6a5e9.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg18x6wr6x8evzjr6a5e9.png" alt="Diagram" width="800" height="98"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The loop only closes if step three actually happens. A channel where everything is answered and nothing is counted stays exactly as busy next year as it is today.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Moving everything into the open&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;A lot of our support traffic arrived as direct messages, usually to whoever the person had spoken to last. That felt friendly and it was quietly destructive: the same question was being answered in four private conversations, and none of us knew it was common.&lt;/p&gt;

&lt;p&gt;Moving to public-by-default took a few weeks of gently redirecting people. The channel got noisier, which looked like a regression. It was not. The noise had always existed, spread across private messages where it could not be measured.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Close the loop out loud
&lt;/h2&gt;

&lt;p&gt;When you fix something that came from the channel, say so in the channel, and name the question it came from.&lt;/p&gt;

&lt;p&gt;This is not self-promotion. It is the only way people learn that reporting friction leads to it being fixed. Teams that never see that connection stop reporting, and once they stop, the research feed dries up and you are back to guessing.&lt;/p&gt;

&lt;p&gt;The message is short: this thing you all kept hitting, here is what changed, here is what you do now. It costs nothing and it is the difference between a channel that gets quieter every quarter and one that gets more useful every quarter.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What changed when we said it out loud&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had been fixing things from the channel for a while without announcing it, because each fix felt too small to mention. The result was that people kept asking about problems we had already solved, and nobody could tell that the platform was improving.&lt;/p&gt;

&lt;p&gt;Posting a short note each time a channel-sourced fix shipped changed the tone noticeably. People started reporting smaller things, including things they had previously worked around silently. The quality of what we heard went up because reporting had visibly started to pay.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. Be careful what you optimise
&lt;/h2&gt;

&lt;p&gt;Once you start measuring the channel, the tempting metric is response time. It is easy to collect and it looks like service quality.&lt;/p&gt;

&lt;p&gt;It is the wrong target. A team optimised for response time gets very good at answering quickly, which makes the channel pleasant and permanent. The goal is not to answer faster. It is to receive fewer questions because fewer things need asking.&lt;/p&gt;

&lt;p&gt;The number worth watching is question volume per team per month, or how long the top bucket stays at the top. If the top question is the same one it was six months ago, the platform is not learning, however fast anyone replies.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The question is the data.&lt;/strong&gt; The answer helps one person; the pattern tells you what to build.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep a tally, not a taxonomy.&lt;/strong&gt; A text file and a habit will beat a tool nobody updates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A repeat question is a defect.&lt;/strong&gt; Fix the cause so it cannot be asked, and do not reach for documentation first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the asides and the jokes.&lt;/strong&gt; A normalised workaround is a request nobody will ever file.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Answer in public&lt;/strong&gt; so the data exists at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Announce channel-sourced fixes&lt;/strong&gt; or people stop reporting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not optimise response time.&lt;/strong&gt; Optimise for the question not being needed.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still have not solved is the silent majority. Everything above depends on someone asking, and the teams that struggle most are often the ones least likely to post in a public channel at all. I do not have a reliable way to hear from them short of going and asking, which does not scale.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/17/the-support-channel-is-product-research/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/17/the-support-channel-is-product-research/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devex</category>
      <category>platformengineering</category>
      <category>research</category>
    </item>
    <item>
      <title>Secrets management that developers do not route around</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:44:33 +0000</pubDate>
      <link>https://dev.to/vivek_itp/secrets-management-that-developers-do-not-route-around-3kkm</link>
      <guid>https://dev.to/vivek_itp/secrets-management-that-developers-do-not-route-around-3kkm</guid>
      <description>&lt;p&gt;Every secrets policy I have seen written down says the same sensible things. Do not commit credentials. Do not paste them in chat. Do not email them. Every organisation I have worked in has done all three, regularly, including the people who wrote the policy.&lt;/p&gt;

&lt;p&gt;The instinct is to treat that as a discipline problem and respond with training. I think it is almost always a design problem. People take the shortest path to getting their work done. If the secure path takes twenty minutes and a ticket, and pasting a value into a direct message takes ten seconds, the direct message wins. Not because anyone is careless, but because one route works and the other does not.&lt;/p&gt;

&lt;p&gt;The question worth asking is not "how do we stop people doing this" but "what made the wrong way easier".&lt;/p&gt;

&lt;h2&gt;
  
  
  Where secrets actually leak
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Where it leaks&lt;/th&gt;
&lt;th&gt;What it means&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pasted into chat&lt;/td&gt;
&lt;td&gt;Getting the value the supported way was slower than asking a colleague&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Committed in a config file&lt;/td&gt;
&lt;td&gt;Local development had no other way to provide it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Baked into a container image&lt;/td&gt;
&lt;td&gt;The build needed it and nothing injected it at runtime&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Living in someone's shell history&lt;/td&gt;
&lt;td&gt;The tool requires a value on the command line&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A shared account nobody rotates&lt;/td&gt;
&lt;td&gt;Individual access needed approval and the shared one did not&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every row is a workaround for a path that did not exist or was too slow. None of them is solved by a reminder.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Measure the secure path in minutes
&lt;/h2&gt;

&lt;p&gt;Before changing anything, time it. From "I need a database password for my new service" to "my code can read it", how long, and how many people are involved?&lt;/p&gt;

&lt;p&gt;If that number is measured in days, no amount of policy will hold. Developers will get the value from wherever they can and carry on, and the ones who do it fastest will look like the most effective engineers on the team.&lt;/p&gt;

&lt;p&gt;The target is that the supported path is the fastest path available. Not merely acceptable. Fastest. When the secure route is quicker than asking a colleague, the problem mostly solves itself.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I timed our own onboarding path once, pretending to be a developer on a new service. Requesting access, waiting for approval, being told which path to use, and finally getting a working value took the better part of two days, spread across three people.&lt;/p&gt;

&lt;p&gt;In the same organisation, asking a teammate in a direct message took about four minutes. Nobody needed to be told which one to use. The policy said one thing and the arithmetic said another, and the arithmetic always won.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Secrets should reach the process, not the person
&lt;/h2&gt;

&lt;p&gt;The single biggest reduction in leaks comes from a change in shape: a developer should never hold the value at all.&lt;/p&gt;

&lt;p&gt;If the runtime fetches secrets itself, using the identity of the workload, then there is nothing to paste, nothing to store locally, and nothing to rotate in a person's password manager. The developer references a name in config and never sees the value behind it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="c1"&gt;# service.yaml, what the developer writes&lt;/span&gt;
&lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;DATABASE_URL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;from_secret&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;orders/db/url&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They name what they need. The platform resolves it at start-up using the service's own identity. Whether that is a cloud secrets manager, a vault, or something else matters far less than the fact that no human is in the path.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The change that removed most of the problem&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We moved from handing developers values to injecting them at runtime. The developer's config named the secret; the workload's identity was what authorised the read.&lt;/p&gt;

&lt;p&gt;What surprised me was how much it simplified their side. There was no longer any question about where to store the value locally, because there was no value. Most of the awkward conversations we had been having about handling procedures stopped being relevant, not because anyone was more careful, but because there was nothing left to handle.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Local development is where it breaks
&lt;/h2&gt;

&lt;p&gt;Production is usually fine. Production has identity, injection, and rotation. It is local development that quietly ruins everything, because on a laptop none of that machinery exists and the developer needs something that works right now.&lt;/p&gt;

&lt;p&gt;That is where the committed &lt;code&gt;.env&lt;/code&gt; file comes from. Not from carelessness, but from a real need with no supported answer.&lt;/p&gt;

&lt;p&gt;The options that work, roughly in order of preference:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Do not need the secret.&lt;/strong&gt; Run against local fakes and a local database with a throwaway password. Most development does not require production credentials.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Short-lived personal credentials.&lt;/strong&gt; The developer authenticates as themselves and the tooling fetches a value that expires within hours.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A dedicated development secret,&lt;/strong&gt; clearly separated, low privilege, rotated on a schedule.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;What does not work is asking people to keep a long-lived production value on a laptop and remember not to commit it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The file everyone had&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Almost every service repository had an example environment file, and almost every developer had a real one sitting beside it, ignored by git and sometimes not. It was not a rule anyone had broken; it was the only way to run the thing.&lt;/p&gt;

&lt;p&gt;Giving the command-line tool the ability to fetch short-lived values for the developer's own identity is what actually emptied those files. The instruction to stop keeping them had been in place for a long time and had changed nothing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Rotation has to be invisible
&lt;/h2&gt;

&lt;p&gt;A rotation policy that requires developers to do something will not be followed past the first quarter. Any secret that a human must remember to change is a secret that does not get changed.&lt;/p&gt;

&lt;p&gt;Rotation works when it is a property of the system: the platform issues short-lived credentials, or rotates long-lived ones on a schedule and the application picks up the new value without a redeploy. The developer's involvement should be zero, and the evidence of rotation should be a log line, not a ticket.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpe0otm4bzdlzz01ib3s.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fcpe0otm4bzdlzz01ib3s.png" alt="Diagram" width="795" height="74"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The dotted line back is the part most setups skip. If picking up a rotated value requires a restart that someone has to schedule, rotation stops being routine and becomes a small project, which means it stops happening.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Make the audit trail useful to the developer
&lt;/h2&gt;

&lt;p&gt;Access logging is usually built for auditors and invisible to everyone else. That is a missed opportunity, because the same data answers questions developers genuinely have: which services read this secret, when was it last used, is anything still depending on the old one.&lt;/p&gt;

&lt;p&gt;When the audit view is useful day to day, people look at it, and things get noticed. A secret nobody has read in six months is a secret that can be removed. A service reading a credential it should not need is worth a conversation. None of that surfaces if the log exists only for a yearly review.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The credential nobody was using&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When we made access history visible in the same place developers looked at their service configuration, the first useful thing it produced was a list of secrets with no reads at all. Several belonged to services that had been decommissioned; the credentials were still valid and still granting access to things.&lt;/p&gt;

&lt;p&gt;Nobody had been ignoring a cleanup process. There had never been a way to see the information. Putting it where people already looked was the whole fix.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. Detection is the backstop, not the plan
&lt;/h2&gt;

&lt;p&gt;Scanning for committed secrets is necessary and it is not a strategy. By the time a scanner fires, the value is in git history, possibly pushed, possibly public, and the response is a rotation and an awkward afternoon.&lt;/p&gt;

&lt;p&gt;Use it, absolutely: a pre-commit hook that catches the value before it is committed is far kinder than a pipeline that catches it afterwards, and both are better than nothing. But treat every hit as a signal about the path, not just an incident to close. Someone needed a value and had nowhere good to put it. Ask what they were trying to do, and fix that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Time the supported path.&lt;/strong&gt; If it is slower than asking a colleague, expect people to ask a colleague.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Keep the value away from people.&lt;/strong&gt; Inject at runtime against the workload's identity so there is nothing to hold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Solve local development explicitly.&lt;/strong&gt; Fakes first, then short-lived personal credentials. Never a long-lived production value on a laptop.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Rotate without asking anyone,&lt;/strong&gt; including the pickup of the new value.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show access history to developers,&lt;/strong&gt; not only to auditors. Unused secrets surface themselves.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat every scanner hit as a design report,&lt;/strong&gt; not only an incident.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The part I have not resolved is third-party credentials. Plenty of vendor integrations still hand you a long-lived key, offer no programmatic rotation, and expect a human to copy it from a web console once a year. Everything above assumes the credential can be issued and rotated by a system, and for that category it simply cannot.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/17/secrets-management-nobody-routes-around/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/17/secrets-management-nobody-routes-around/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>secrets</category>
      <category>security</category>
      <category>devex</category>
    </item>
    <item>
      <title>Observability for developers, not just for on-call</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Tue, 15 Sep 2026 16:16:16 +0000</pubDate>
      <link>https://dev.to/vivek_itp/observability-for-developers-not-just-for-on-call-1hi2</link>
      <guid>https://dev.to/vivek_itp/observability-for-developers-not-just-for-on-call-1hi2</guid>
      <description>&lt;p&gt;Almost every dashboard I have built or inherited was designed for an operator. Cluster health, node pressure, error rates across the estate, a wall of graphs that makes sense if you are holding the pager and looking for the one thing that is on fire.&lt;/p&gt;

&lt;p&gt;That is a real audience with real needs. It is just not the audience that uses observability most often. The developer who merged a change twenty minutes ago has a much narrower question, asks it far more frequently, and is usually the person least equipped to answer it with the tools we gave them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two people, two questions
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The operator asks&lt;/th&gt;
&lt;th&gt;The developer asks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What is broken right now, across everything?&lt;/td&gt;
&lt;td&gt;Did the thing I just shipped break anything?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Which service is causing this?&lt;/td&gt;
&lt;td&gt;Is my service healthy, and was it healthy before?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Do I need to page someone?&lt;/td&gt;
&lt;td&gt;Can I go to lunch?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Is this worse than an hour ago?&lt;/td&gt;
&lt;td&gt;Is this worse than before my change?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The right-hand column is answerable with a fraction of the data in the left. It is also asked ten or twenty times more often. Most platforms optimise entirely for the left column and leave the right to whoever can write a query.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. The first question is always about the deploy
&lt;/h2&gt;

&lt;p&gt;Right after a release, a developer wants to know one thing: is this worse than it was before I pressed merge? Not the absolute error rate. The comparison.&lt;/p&gt;

&lt;p&gt;A dashboard showing a line at two percent errors is useless on its own. Was it two percent this morning? Was it zero? The developer usually does not know, and finding out means changing a time range, which means learning the tool, which means most of them do not bother and instead wait to see if anyone complains.&lt;/p&gt;

&lt;p&gt;The fix is to make the deploy the unit of observation. Mark releases on the graphs, default the window to span the last one, and put the before-and-after difference in words, not just in pixels.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had good dashboards and almost nobody outside the platform team opened them. When I asked why, the answer was consistent: people did not know what normal looked like, so a graph told them nothing.&lt;/p&gt;

&lt;p&gt;Adding deploy markers to the service graphs changed that more than any new metric did. Suddenly the question was not "is two percent bad" but "did it change at that vertical line". People could answer that without knowing anything about the system.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Scope to the service, not the cluster
&lt;/h2&gt;

&lt;p&gt;A developer owns a service. They do not own the cluster, cannot act on node pressure, and should not have to scroll past it.&lt;/p&gt;

&lt;p&gt;Every service should have a dashboard that covers only that service and can be reached without choosing anything from a dropdown. Traffic, errors, latency, saturation, restarts, recent deploys. Nothing about the neighbours.&lt;/p&gt;

&lt;p&gt;This sounds obvious and is surprisingly rare, usually because the platform team built one excellent dashboard with a service selector at the top. A selector is a small thing for the person who built it and a real barrier for someone opening the tool for the second time this quarter.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The dropdown nobody used&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our main service dashboard had a variable at the top to pick a service. To us it was elegant: one dashboard, every service. To everyone else it was a page that showed the wrong data by default and had to be configured before it meant anything.&lt;/p&gt;

&lt;p&gt;We generated a per-service dashboard from the same definition instead. Same graphs, same code, no selection step. Usage went up sharply and the questions we got about it changed from "how do I use this" to actual questions about the data.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Link from where they already are
&lt;/h2&gt;

&lt;p&gt;Nobody navigates to an observability tool. They arrive at it from somewhere else, usually in a hurry, usually from a pull request, a deploy notification, or an alert.&lt;/p&gt;

&lt;p&gt;Every one of those places should carry a direct link to the right view of the right service, already scoped and time-ranged. A developer should never have to know the tool's URL, let alone its query language, to answer the question they had.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6b7rpk7brs30i54gsfwd.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F6b7rpk7brs30i54gsfwd.png" alt="Diagram" width="797" height="121"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The links are not a nice touch. They are most of the product. An observability platform that has to be found is an observability platform that does not get used by anyone who is not already fluent in it.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The link in the deploy message&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a long time our deploy notification said which version had gone out and nothing else. Anyone who wanted to check on it had to open the tool, find their service, and set a time range.&lt;/p&gt;

&lt;p&gt;Adding two links to that message, one to the service view scoped to the deploy and one to that version's logs, was probably an afternoon of work. It did more for how often developers looked at their own telemetry than the previous year of dashboard improvements.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Logs must be findable without a query language
&lt;/h2&gt;

&lt;p&gt;Query languages are a wall. They are learnable, and the people who learn them get enormous value, and most developers will not learn one to check on a deploy they are already fairly confident about.&lt;/p&gt;

&lt;p&gt;The default path to logs should be a link, filtered to the service and the version, sorted newest first, with no syntax involved. The query language stays available for the people who want it and the incidents that need it. It just cannot be the entry point.&lt;/p&gt;

&lt;p&gt;The test is simple: can a developer who has never opened the tool get to their service's recent errors in one click from the deploy notification? If not, logs are effectively a platform-team-only feature.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Tell the owner, not just the pager
&lt;/h2&gt;

&lt;p&gt;Alerts usually route to whoever is on call. That is correct for anything user-facing and urgent. It is wrong for the large category of things that are the owning team's problem and nobody else's: a rising error rate on one endpoint, a queue growing slowly, a job that has started failing overnight.&lt;/p&gt;

&lt;p&gt;Those should reach the team that owns the service, in their own channel, during their own hours, without waking anyone. Teams that get told about their own service's problems start fixing them before they become incidents. Teams that never hear anything assume everything is fine, because from where they sit it is.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The alerts that went to the wrong people&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Everything alertable pointed at the on-call rotation, which meant the platform team saw every application-level problem first and had to work out whose it was and who to tell. We were a routing layer made of humans.&lt;/p&gt;

&lt;p&gt;Splitting the alerts into "wakes someone up" and "tells the owning team in their channel" removed a lot of that. The second category was much larger than the first. Most of it had never needed a platform engineer at all; it had just never had anywhere else to go.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. Do not make them learn your cardinality problems
&lt;/h2&gt;

&lt;p&gt;There is a strong temptation to teach developers the internals: why they cannot label a metric with a user ID, what a histogram bucket is, why their query is slow. That knowledge is genuinely useful and it is not their job.&lt;/p&gt;

&lt;p&gt;The platform should provide instrumentation that is hard to misuse, sensible defaults, and a clear error when someone does something expensive. Not a training course. Every hour a product developer spends learning the observability stack's failure modes is an hour not spent on the thing they were hired for, and they will forget most of it before they need it again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Make the deploy the unit of observation.&lt;/strong&gt; Mark releases and default to the window around the last one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;One dashboard per service, no selector.&lt;/strong&gt; Generate them; do not ask people to configure anything.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Link from the pull request, the deploy message, and the alert.&lt;/strong&gt; Nobody navigates to the tool.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Logs in one click, no query language&lt;/strong&gt; for the default path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route non-urgent alerts to the owning team,&lt;/strong&gt; not to the pager.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not teach developers your storage constraints.&lt;/strong&gt; Make the safe thing the default.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still get wrong is knowing when a developer-facing view has become too simple to be honest. A single green tick after a deploy is exactly what people want and it hides a lot. I have not found the line between reassuring and misleading, and I suspect it moves depending on how much the team already trusts the platform.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/15/observability-for-developers/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/15/observability-for-developers/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>observability</category>
      <category>devex</category>
      <category>monitoring</category>
    </item>
    <item>
      <title>Golden paths that people actually take</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Fri, 11 Sep 2026 18:06:57 +0000</pubDate>
      <link>https://dev.to/vivek_itp/golden-paths-that-people-actually-take-2kio</link>
      <guid>https://dev.to/vivek_itp/golden-paths-that-people-actually-take-2kio</guid>
      <description>&lt;p&gt;Every platform team I have been on has built a golden path: a template, a paved road, a "blessed" way to build and ship a service. Every one of them was well used for a few months. The interesting question is never whether people start on the path. It is whether they are still on it a year later, and what made them leave.&lt;/p&gt;

&lt;p&gt;I have watched paths get abandoned quietly, one service at a time, while the adoption slide still said one hundred percent. This post is about why that happens and what I have found actually keeps people on a path without forcing them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why people leave the path
&lt;/h2&gt;

&lt;p&gt;Nobody leaves a golden path because they enjoy maintaining their own pipeline. They leave because, at some specific moment, staying on the path was more work than stepping off it. If you can find those moments, you can fix them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;The moment&lt;/th&gt;
&lt;th&gt;What the developer said&lt;/th&gt;
&lt;th&gt;What it actually signals&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Needed something the template did not support&lt;/td&gt;
&lt;td&gt;"I just added a step to my copy"&lt;/td&gt;
&lt;td&gt;The path was too narrow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Template got an update&lt;/td&gt;
&lt;td&gt;"I'll upgrade later"&lt;/td&gt;
&lt;td&gt;Upgrading was a migration, not a click&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Something broke inside the template&lt;/td&gt;
&lt;td&gt;"I could not see what it was doing"&lt;/td&gt;
&lt;td&gt;The path was opaque&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A new joiner copied an old service&lt;/td&gt;
&lt;td&gt;"That is what the last person did"&lt;/td&gt;
&lt;td&gt;The path was not the obvious starting point&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Platform team was slow to respond&lt;/td&gt;
&lt;td&gt;"We could not wait"&lt;/td&gt;
&lt;td&gt;The path had one gatekeeper&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each row is a design decision the platform team made, or failed to make. None of them is the developer's fault.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. A golden path is a product, not a policy
&lt;/h2&gt;

&lt;p&gt;The fastest way to get everyone onto a path is to mandate it. The fastest way to get everyone quietly off it is the same mandate. A policy gets compliance on the day it is checked and workarounds on every other day.&lt;/p&gt;

&lt;p&gt;A path that people take voluntarily has to be better than the alternative for the person using it, on the day they use it. That is a product question: who is this for, what do they get, and why would they choose it over copying last month's service?&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first path I helped build was rolled out with a mandate: all new services use the template, no exceptions. Adoption was immediate and the dashboard looked great.&lt;/p&gt;

&lt;p&gt;About six months in, I looked at what the services actually contained. Most of them had started from the template and then diverged: an extra step here, a swapped-out base image there, a pipeline job commented out because it was slow. They were on the path in name only. The mandate had been met on day one and ignored from day two, because nothing about the path made staying on it easier than leaving.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. It has to win on day one and on day one hundred
&lt;/h2&gt;

&lt;p&gt;A template is easy to make attractive for a brand new service. Four lines of config and you are deployed. That is day one, and most paths are designed entirely around it.&lt;/p&gt;

&lt;p&gt;Day one hundred is different. The service has a queue consumer now, a scheduled job, a second database, a customer-specific quirk. The question on day one hundred is whether the path still fits, and whether the platform team has shipped anything in the meantime that the service can pick up without effort.&lt;/p&gt;

&lt;p&gt;Paths that only win on day one lose on day one hundred, and day one hundred is where every service lives most of its life.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The template with no upgrade story&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Our first template was a repository you copied. Copying it was trivial. Updating it was impossible, because once copied, each service owned its own copy and nothing linked it back.&lt;/p&gt;

&lt;p&gt;When we shipped an improvement to the template, exactly one service got it: the next new one. Everything already running stayed on whatever version it had started with. Within a year the template had a dozen versions in the wild and no way to tell which service was on which. The path did not have a day one hundred at all.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Escape hatches keep people on the path
&lt;/h2&gt;

&lt;p&gt;The instinct is to make the path strict, so that nobody can deviate. In practice a strict path with no exit means the first unsupported need pushes the whole service off it. Now the developer is not deviating in one place. They are maintaining everything themselves.&lt;/p&gt;

&lt;p&gt;A better design has a clear, supported way to step off the path for one thing while staying on it for everything else. Add a custom step here. Override this one value. Ship a raw manifest alongside the generated ones. The escape hatch is small, visible, and counted.&lt;/p&gt;

&lt;p&gt;Counting matters more than allowing. Every use of the escape hatch is a signal that the path is missing something. If ten services use the same hatch for the same reason, that reason belongs in the path.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The hatch that became a roadmap&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When we added an escape hatch to the deploy template, I expected it to be abused. It mostly was not. What it did instead was give us a list.&lt;/p&gt;

&lt;p&gt;Within a few months the hatch usage clustered around three needs: a sidecar for a particular internal agent, a cron-style job, and a different health check shape. All three became first-class template features, and the services using the hatch for them moved back onto the standard path without being asked. The hatch had not weakened the path. It had told us what to build next.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Upgrades must be a pull request, not a project
&lt;/h2&gt;

&lt;p&gt;If moving from template version 3 to version 4 requires reading a migration guide, the service will stay on version 3. The only upgrade mechanism that works at scale is one where the platform team does the work and the service team reviews it.&lt;/p&gt;

&lt;p&gt;That means the template is not a thing you copy. It is a dependency you reference, with a version. Upgrading is a pull request, opened automatically, with a diff small enough to read and a pipeline run that proves it works. The service team merges it or asks a question. Either way, the effort is minutes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5q7kxk5ikntonz15ukbc.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F5q7kxk5ikntonz15ukbc.png" alt="Diagram" width="798" height="167"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The second branch is the important one. When the upgrade breaks a service, the fix goes into the template so that it does not break the next one. The service team never has to become experts in the thing they are upgrading.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;From copied repo to referenced version&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Moving off the copied-repository model was the single change that made the path survivable. Services referenced a versioned template, and a bot opened a pull request in each one whenever a new version was tagged.&lt;/p&gt;

&lt;p&gt;The first few rounds were rough, because the template had never been upgraded in place before and every hidden assumption surfaced at once. But the fixes went into the template, and after a few releases most upgrade pull requests merged without a comment. The dozen-versions-in-the-wild problem stopped growing, then started shrinking.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Measure drift, not adoption
&lt;/h2&gt;

&lt;p&gt;Adoption is the easy number and the wrong one. A service that started on the path and has been diverging for a year still counts as adopted.&lt;/p&gt;

&lt;p&gt;The number that tells you whether the path is working is drift: how far each service is from the current template. Versions behind, escape hatches in use, files that differ from the generated ones. If drift is low and stable, the path is winning. If drift climbs, people are leaving, whatever the adoption number says.&lt;/p&gt;

&lt;p&gt;Drift is also actionable in a way adoption is not. A service three versions behind is a conversation. Twenty services all using the same override is a feature request.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The dashboard that replaced the adoption slide&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We replaced the adoption percentage with a drift view: every service, which template version it was on, and which escape hatches it used. It was a less flattering picture and a far more useful one.&lt;/p&gt;

&lt;p&gt;It changed what the platform team worked on. Instead of building new features and hoping people would take them, we looked at the services with the highest drift and asked why. The answers were usually specific and usually fixable.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. Talk to the people who left
&lt;/h2&gt;

&lt;p&gt;The services that stepped off the path are the best source of information about it, and they are the ones the platform team is least likely to talk to. Leaving feels like a rejection, so both sides avoid the conversation.&lt;/p&gt;

&lt;p&gt;It should be the opposite. A service that left had a reason, and the reason is almost always something the path should have handled. Ask, without judgement, and fix the reason. Then make it easy to come back.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The one missing feature&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One team left the path entirely over a single need: their service required a specific network policy the template could not express. Rather than use the escape hatch, which did not cover it, they rebuilt their pipeline from scratch and maintained it themselves for months.&lt;/p&gt;

&lt;p&gt;When we finally asked, the fix took a week. We added the option to the template, they moved back, and two other services that had quietly worked around the same limitation moved back with them. Nobody had raised it because nobody thought the platform team would want to hear it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Treat the path as a product.&lt;/strong&gt; People stay on it because it is better, not because it is required.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Design for day one hundred&lt;/strong&gt;, not just the first deploy. Most of a service's life is maintenance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build a small, counted escape hatch.&lt;/strong&gt; Its usage is your roadmap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make upgrades automatic pull requests.&lt;/strong&gt; If upgrading is a project, nobody upgrades.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure drift, not adoption.&lt;/strong&gt; Adoption hides the problem. Drift points at it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask the people who left why.&lt;/strong&gt; Then fix it and make returning easy.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The part I have not figured out is what to do with the services that will never come back: the old ones with no active team, too different to upgrade and too important to ignore. A golden path is for the services that are still moving, and I do not yet have a good answer for the ones that are not.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/11/golden-paths-that-people-actually-take/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/11/golden-paths-that-people-actually-take/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>platformengineering</category>
      <category>goldenpath</category>
      <category>templates</category>
      <category>devex</category>
    </item>
    <item>
      <title>Verifying commit signatures in the CI pipeline</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Fri, 11 Sep 2026 15:09:04 +0000</pubDate>
      <link>https://dev.to/vivek_itp/verifying-commit-signatures-in-the-ci-pipeline-4m0k</link>
      <guid>https://dev.to/vivek_itp/verifying-commit-signatures-in-the-ci-pipeline-4m0k</guid>
      <description>&lt;p&gt;Most teams that "require signed commits" have ticked a box in branch protection and stopped there. That box checks that a signature exists. It does not check that the key belongs to someone you trust, and it does nothing once the code is past the merge button.&lt;/p&gt;

&lt;p&gt;I worked on an application where that was not good enough. The codebase did cryptographic operations, the security policy was zero trust, and the requirement was blunt: every line of code reaching production has to be traceable to a verified person, and every artifact has to be traceable to a verified build. This post is how I made the pipeline enforce that, and the parts that hurt.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "signed" actually proves
&lt;/h2&gt;

&lt;p&gt;A commit signature proves that whoever held a particular private key signed that exact commit content. That is all. It does not prove who holds the key, and it does not prove the key is still trusted today.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;Branch protection alone&lt;/th&gt;
&lt;th&gt;Pipeline verification&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;A signature is present on each commit&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The signature is cryptographically valid&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The key belongs to a known, identity-verified person&lt;/td&gt;
&lt;td&gt;Only if the key is registered with the hosting platform&lt;/td&gt;
&lt;td&gt;Yes, against your own key registry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Every commit in a promotion, not just the tip, is verified&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The build artifact is tied to a verified build&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes, if the build signs its output&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The right-hand column is what zero trust actually asks for. The left-hand column is what most teams settle for.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Pick SSH signing, and own the key registry
&lt;/h2&gt;

&lt;p&gt;We used SSH keys for signing rather than GPG. Every developer already had an SSH key, the tooling is built into git, and there was no separate keyring to teach people.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git config &lt;span class="nt"&gt;--global&lt;/span&gt; gpg.format ssh
git config &lt;span class="nt"&gt;--global&lt;/span&gt; user.signingkey ~/.ssh/id_ed25519.pub
git config &lt;span class="nt"&gt;--global&lt;/span&gt; commit.gpgsign &lt;span class="nb"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important decision was not the key type. It was where the public keys lived. We did not rely on the keys uploaded to the git hosting platform. We kept our own registry: each developer's verified public key stored in a parameter store, under their identity, managed by the platform team.&lt;/p&gt;

&lt;p&gt;That registry is the source of truth. The pipeline reads from it, not from the hosting platform. If a key is not in the parameter store, the commit is not trusted, regardless of what the web UI says.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Why the registry had to be ours&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The git platform will happily show a "Verified" badge next to a commit if the key is attached to any account. That says nothing about whether we had checked who owned the key. For a codebase handling crypto operations, "someone uploaded this key to their profile" was not an acceptable identity check.&lt;/p&gt;

&lt;p&gt;Holding the keys in our own parameter store meant we decided what "trusted" meant, and the pipeline enforced our definition rather than the platform's.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Verify in the pipeline, not only at the branch
&lt;/h2&gt;

&lt;p&gt;Branch protection stayed on. It is a cheap first gate and it rejects unsigned pushes before they waste anyone's time. But the real check ran inside the GitLab runner.&lt;/p&gt;

&lt;p&gt;On every merge request pipeline, the job did the following:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;List the commits new to that merge request.&lt;/li&gt;
&lt;li&gt;For each commit, read the author identity.&lt;/li&gt;
&lt;li&gt;Fetch that person's public key from the parameter store.&lt;/li&gt;
&lt;li&gt;Build an &lt;code&gt;allowed_signers&lt;/code&gt; file from the keys, and run &lt;code&gt;git verify-commit&lt;/code&gt; against it.&lt;/li&gt;
&lt;li&gt;Fail the job on the first commit that does not verify.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Build the allowed-signers file from the registry&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;author &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;git log &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s1"&gt;'%ae'&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$BASE&lt;/span&gt;&lt;span class="s2"&gt;..&lt;/span&gt;&lt;span class="nv"&gt;$HEAD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; | &lt;span class="nb"&gt;sort&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nv"&gt;key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;fetch_public_key_from_parameter_store &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$author&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
  &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$author&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="nv"&gt;$key&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&amp;gt;&lt;/span&gt; allowed_signers
&lt;span class="k"&gt;done

&lt;/span&gt;git config gpg.ssh.allowedSignersFile allowed_signers

&lt;span class="c"&gt;# Verify every new commit, not just the tip&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;sha &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;git rev-list &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$BASE&lt;/span&gt;&lt;span class="s2"&gt;..&lt;/span&gt;&lt;span class="nv"&gt;$HEAD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;git verify-commit &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$sha&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;{&lt;/span&gt; &lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="s2"&gt;"Unverified commit: &lt;/span&gt;&lt;span class="nv"&gt;$sha&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;BASE..HEAD&lt;/code&gt; range is the whole point. Checking only the tip commit lets an unsigned commit ride in underneath a signed one. Every commit in the range has to verify, or the job fails.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Two ranges, two gates&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;On merge request pipelines, the range was "commits new to this MR". That kept the check fast and gave developers a clear signal on their own work.&lt;/p&gt;

&lt;p&gt;On promotion to staging, the range was different: every commit between the previous merge to staging and the current one. That second pass was the one that mattered for the audit trail. It meant nothing could reach staging that had not been verified against the registry, even if it had somehow slipped through an earlier gate.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Sign what you build, not just what you commit
&lt;/h2&gt;

&lt;p&gt;Verifying commits proves where the source came from. It says nothing about whether the artifact in the repository manager was built from that source by a trusted process.&lt;/p&gt;

&lt;p&gt;So the build signed its output too. When the GitLab runner published the JAR to Nexus, it signed the artifact with a KMS key that only the runner's role could use. Deployments verified that signature before installing anything.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnklq80coqjcoa8z1jegp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnklq80coqjcoa8z1jegp.png" alt="Diagram" width="797" height="61"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;With both halves in place, the chain was continuous: a verified person wrote the code, a verified pipeline built it, and the thing running in production could be traced back through both.&lt;/p&gt;

&lt;h2&gt;
  
  
  4. Everything that can break, will
&lt;/h2&gt;

&lt;p&gt;I would like to say the rollout was smooth. It was not. Every failure mode you can think of happened.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rebases and squash merges&lt;/strong&gt; rewrote commits and dropped signatures. The rewritten commit was now unsigned and failed the check.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automation commits&lt;/strong&gt; from bots and pipelines had no key at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Key rotation&lt;/strong&gt; meant a developer's older commits verified against a key that was no longer their current one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;New laptops&lt;/strong&gt; meant a developer signing with a key nobody had registered yet.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The fixes were mostly boring, which is the good kind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When a rewrite dropped signatures, the developer re-signed and force-pushed their branch. Force push had to be allowed on feature branches for this to work, which took some convincing.&lt;/li&gt;
&lt;li&gt;During key rotation, the registry held both the old and the new key for the same person. Old commits verified against the old key, new ones against the new key, and the old key was removed once nothing depended on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The force-push argument&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The instinct on the security side was to forbid force pushes everywhere. But a signature check that rejects rewritten commits, combined with a policy that forbids rewriting them again, leaves a developer with no way out. Their only option is to open a new branch and cherry-pick, which is worse for everyone.&lt;/p&gt;

&lt;p&gt;Allowing force push on feature branches, while keeping it locked on protected branches, was the compromise. The signature check on the protected branch was the real control. The feature branch was a workspace.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Roll it out with a cutoff date, not a flag day
&lt;/h2&gt;

&lt;p&gt;Turning signature verification on for a repository with years of history would have failed on every old commit at once. Instead we did two things.&lt;/p&gt;

&lt;p&gt;First, we verified every developer's key before enforcement. Each person's public key was checked against their identity, the way you would verify a customer before opening an account, and only then was it written to the parameter store.&lt;/p&gt;

&lt;p&gt;Second, we set an enforcement date. The pipeline checked signatures only on commits made after that date and ignored everything before it. That gave people a defined window to get their keys registered and their signing configured, and it meant history did not need rewriting.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;CUTOFF&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"2025-01-01"&lt;/span&gt; +%s&lt;span class="si"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for &lt;/span&gt;sha &lt;span class="k"&gt;in&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;git rev-list &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$BASE&lt;/span&gt;&lt;span class="s2"&gt;..&lt;/span&gt;&lt;span class="nv"&gt;$HEAD&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;do
  &lt;/span&gt;&lt;span class="nv"&gt;ts&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;git show &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;--format&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;%ct &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$sha&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="si"&gt;)&lt;/span&gt;
  &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="o"&gt;[&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$ts&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;-ge&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$CUTOFF&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="k"&gt;then
    &lt;/span&gt;git verify-commit &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$sha&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;exit &lt;/span&gt;1
  &lt;span class="k"&gt;fi
done&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The grace period did most of the work&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Announcing a date and then enforcing it on that date, with nothing retroactive, meant almost nobody was surprised. The people who had not set up signing found out on their first push after the cutoff, with a clear failure, and fixed it in a few minutes. Without the cutoff, the same rollout would have meant either rewriting history or exempting so much that the check meant nothing.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What I would still change
&lt;/h2&gt;

&lt;p&gt;The rule I would keep is the one about owning the registry. It is the difference between "signed" and "signed by someone we checked".&lt;/p&gt;

&lt;p&gt;The part I have not solved is maintenance. Keys expire, people change laptops, people leave. Every one of those is a manual update to the parameter store today. That is fine at a small scale and a liability at a large one. What this needs is automation: key registration tied to identity provisioning, rotation that updates the registry without a ticket, and revocation that happens the moment someone's access is removed. I have not built that yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A "Verified" badge is not identity verification.&lt;/strong&gt; Keep your own registry of public keys and decide for yourself what trusted means.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Verify in the pipeline, over the full commit range.&lt;/strong&gt; Branch protection is a first gate, not the control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign the artifact too.&lt;/strong&gt; Source provenance without build provenance is half a chain.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Expect rewrites, bots, rotation, and new machines to break it.&lt;/strong&gt; Allow re-signing on feature branches and hold old and new keys during rotation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Enforce from a date, not from the beginning of time.&lt;/strong&gt; Verify keys first, then set a cutoff.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The open question is still key maintenance. Until registration, rotation, and revocation are automated, the trust model is only as current as the last person who remembered to update the parameter store.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/11/verifying-commit-signatures-in-the-ci-pipeline/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/11/verifying-commit-signatures-in-the-ci-pipeline/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ci</category>
      <category>security</category>
      <category>signedcommits</category>
      <category>supplychain</category>
    </item>
    <item>
      <title>Error messages are part of the platform</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Wed, 09 Sep 2026 18:47:15 +0000</pubDate>
      <link>https://dev.to/vivek_itp/error-messages-are-part-of-the-platform-30lc</link>
      <guid>https://dev.to/vivek_itp/error-messages-are-part-of-the-platform-30lc</guid>
      <description>&lt;p&gt;Nobody reads a platform's documentation on a good day. On a good day the pipeline is green, the deploy went out, and the developer never thinks about the platform at all. The platform only gets read when something breaks, and what gets read is the error message.&lt;/p&gt;

&lt;p&gt;That makes the error message the most-read text a platform team ever writes. It is also, almost always, the text nobody on the platform team wrote on purpose. It is whatever the underlying tool printed, passed through untouched, with a stack trace attached.&lt;/p&gt;

&lt;p&gt;I spent a long time treating errors as something that happened to the platform rather than something the platform produced. This post is what changed my mind, and what I do differently now.&lt;/p&gt;

&lt;h2&gt;
  
  
  The gap between what it said and what it meant
&lt;/h2&gt;

&lt;p&gt;Here are the kinds of messages developers actually see, next to what the platform team knows they mean. The gap between the two columns is where support requests come from.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What the developer saw&lt;/th&gt;
&lt;th&gt;What it actually meant&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;ImagePullBackOff&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The image tag you deployed was never pushed. Your build job probably failed.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;exit code 137&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The runner ran out of memory. Not your code.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Error: forbidden: User cannot list resource "pods"&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;You are in the wrong namespace, or you were never granted access to this one.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;KeyError: 'DATABASE_URL'&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The secret for this environment was never created.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;x509: certificate signed by unknown authority&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The internal CA is not installed in your base image.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;Error: no matches for kind "Ingress" in version "extensions/v1beta1"&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The template you copied is three years old.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every row in the right column is something the platform team could have printed instead. None of them did, because the message came from a tool underneath the platform and nobody caught it on the way up.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. An error has three jobs
&lt;/h2&gt;

&lt;p&gt;A useful error message answers three questions, in this order:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;What broke?&lt;/strong&gt; In one plain sentence, from the developer's point of view.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why?&lt;/strong&gt; The most likely cause, stated as a fact if it is known and as a guess if it is not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What do I do now?&lt;/strong&gt; A command, a link, or a name. Something the developer can act on in the next minute.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If a message does only the first job, the developer knows they are stuck. If it does all three, the developer is usually unstuck before they think of asking anyone.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first error I rewrote was the image pull failure. The raw version was &lt;code&gt;ImagePullBackOff&lt;/code&gt; with a pod name. Developers would see it, check that their code compiled, re-run the deploy, see it again, and then ask in the channel.&lt;/p&gt;

&lt;p&gt;The rewritten version, printed by the deploy step, read something like: &lt;em&gt;"Deploy failed: image tag &lt;code&gt;abc1234&lt;/code&gt; does not exist in the registry. This usually means the build job for this commit failed or has not finished. Check the build job here: link."&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The first line was the same information. The second and third lines were the difference between a support request and a self-fix. Questions about that error mostly stopped after the change.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Write for the person who did not build the system
&lt;/h2&gt;

&lt;p&gt;Most error messages are written, implicitly, for the person who wrote the code that raised them. They use internal names, assume knowledge of the architecture, and skip context that the author never needed because they already had it in their head.&lt;/p&gt;

&lt;p&gt;A developer on a product team does not have that context. They do not know that "reconciler" means the deploy controller, that "upstream" means the service behind the proxy, or that error code 4012 is the one about quotas.&lt;/p&gt;

&lt;p&gt;The test I use: would a competent engineer who joined last week and has never read our platform code understand this message? If not, it is not finished.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The message with an internal name in it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;We had a validation step that rejected service configs with a message naming the internal component that did the checking. Something like &lt;em&gt;"rejected by policy-gate: rule R7 failed."&lt;/em&gt; Everyone on the platform team knew what R7 was. Nobody else did, and there was nothing in the message to search for.&lt;/p&gt;

&lt;p&gt;The fix was to make the rule print its own explanation: &lt;em&gt;"Service config rejected: &lt;code&gt;memory&lt;/code&gt; must be at most 4Gi for services without a capacity exception. Yours is 8Gi. To request an exception, see: link."&lt;/em&gt; Same rule, same check. The developer just got to read the reason instead of the rule number.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Put the fix in the message
&lt;/h2&gt;

&lt;p&gt;The most valuable thing an error can contain is the next command to run. Not a description of the fix. The fix.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A lint failure should print the command that auto-fixes it.&lt;/li&gt;
&lt;li&gt;A missing secret should print the command or link that creates it.&lt;/li&gt;
&lt;li&gt;A permissions error should name the group that grants the permission and where to request it.&lt;/li&gt;
&lt;li&gt;A deprecated config field should print the new field name and the one-line change.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is also the cheapest place to put documentation, because it is the only documentation that is guaranteed to be read at the moment it is needed.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;# Before
Error: validation failed for field 'healthcheck'

# After
Error: 'healthcheck' must be a path starting with '/', got 'healthz'.
Fix:   change it to '/healthz' in service.yaml.
Docs:  https://internal/docs/service-config#healthcheck
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The link that replaced a wiki page&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For a while the most-linked page in our support channel was a wiki article explaining how to request access to a new environment. Someone would hit a permissions error, ask, and get the link. Dozens of times.&lt;/p&gt;

&lt;p&gt;We put the link in the error message. The wiki page did not change. The number of people who needed to be told about it dropped close to zero, because the tool told them at the exact moment they needed it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Lead with the plain sentence, keep the raw detail
&lt;/h2&gt;

&lt;p&gt;The instinct when wrapping an error is to hide the original. That is a mistake. The original error is often the only thing that helps when the plain-language guess is wrong, and someone on the platform team will eventually need it.&lt;/p&gt;

&lt;p&gt;The structure that works: plain sentence first, cause and action next, raw detail last and clearly labelled.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fplv8opb1jstou6u7odwz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fplv8opb1jstou6u7odwz.png" alt="Diagram" width="552" height="844"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The developer reads the top. The platform engineer, if it gets that far, reads the bottom. Nobody has to scroll through a stack trace to find out that a tag was missing.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The wrapper that hid too much&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;My first attempt at friendlier errors over-corrected. The deploy tool caught every failure and printed a short summary, and dropped the original output entirely. It was clean until the summary was wrong. Then the developer had a confident sentence pointing at the wrong cause and no way to see what had actually happened.&lt;/p&gt;

&lt;p&gt;We put the raw output back, below a separator, under a heading that said it was for debugging. The summaries stayed. The dead ends went away.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Treat the most common errors as bugs
&lt;/h2&gt;

&lt;p&gt;Every platform has a small set of errors that account for most of the confusion. You can find them without any tooling: read a month of the support channel and count. The same five or six messages will come up over and over.&lt;/p&gt;

&lt;p&gt;Each of those is a bug in the platform, even if the underlying tool is behaving correctly. The bug is that the platform let a confusing message reach a developer. Put them on the backlog, in priority order by how often they appear, and fix them the way you would fix any other bug.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The list on the wall&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;At one point I kept a short list of the errors that generated the most repeat questions, ordered by count. It had maybe eight entries. We worked down it during quiet weeks, one rewritten message at a time.&lt;/p&gt;

&lt;p&gt;It was some of the highest-return work the platform team did that year, and almost none of it involved changing what the platform actually did. It only changed what it said.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  6. Test error messages like features
&lt;/h2&gt;

&lt;p&gt;If a message matters, it deserves a test. Not a test that the error is raised, which most codebases already have, but a test that the message says what it should: names the field, includes the fix, links to the right place.&lt;/p&gt;

&lt;p&gt;This sounds excessive until the first time someone refactors the validation code and the helpful message quietly reverts to a generic one. Nobody notices, because errors are not on the happy path, until the support channel fills up again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;The error message is the platform's real interface.&lt;/strong&gt; It is read more than any docs page, at the exact moment it matters.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Three jobs:&lt;/strong&gt; what broke, why, what to do next. A message that does one is not finished.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write for someone who joined last week&lt;/strong&gt; and has never seen the platform code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Put the fix in the message.&lt;/strong&gt; A command or a link beats a description every time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plain sentence on top, raw error on the bottom.&lt;/strong&gt; Never hide the original.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count the repeat errors and fix them as bugs.&lt;/strong&gt; It is cheap, and it compounds.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still have not worked out is errors from tools I do not own. When a managed cloud service or a third-party CLI returns something cryptic, the platform can wrap it, but the wrapping is a guess, and the guess goes stale every time the vendor changes their wording.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/09/error-messages-are-part-of-the-platform/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/09/error-messages-are-part-of-the-platform/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>devex</category>
      <category>errormessages</category>
      <category>tooling</category>
    </item>
    <item>
      <title>Self-service on a slide vs self-service in practice</title>
      <dc:creator>vivek Itp</dc:creator>
      <pubDate>Wed, 09 Sep 2026 18:42:05 +0000</pubDate>
      <link>https://dev.to/vivek_itp/self-service-on-a-slide-vs-self-service-in-practice-1839</link>
      <guid>https://dev.to/vivek_itp/self-service-on-a-slide-vs-self-service-in-practice-1839</guid>
      <description>&lt;p&gt;Every internal platform I have seen described in a slide deck was self-service. Developers would "spin up a new service in minutes" with "no tickets" and "no waiting on the platform team". I have never seen one of those slides that was accurate on the day it was presented.&lt;/p&gt;

&lt;p&gt;That is not because the people writing the slides were lying. It is because self-service is a property of the whole path a developer walks, and the slide only describes the part the platform team built. The parts they did not build, the access request, the DNS entry, the secret that has to be created by someone else, the approval that lives in a different tool, are invisible from the platform side and completely visible from the developer's.&lt;/p&gt;

&lt;p&gt;This post is about measuring that gap honestly, and about what actually closes it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The only definition that matters
&lt;/h2&gt;

&lt;p&gt;Self-service means a developer with a normal amount of access can go from "I need a new service" to "it is running in production and I can see it" without asking a human for anything.&lt;/p&gt;

&lt;p&gt;Not "without asking the platform team". Without asking &lt;em&gt;anyone&lt;/em&gt;. The moment there is a ticket, a Slack message, or a form that someone else has to act on, the path is not self-service, no matter how good the tooling is on either side of that gap.&lt;/p&gt;

&lt;p&gt;That definition is strict on purpose. A path that is self-service except for one step is not ninety percent self-service. It is blocked at that step, and every developer will wait there.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the path usually looks like
&lt;/h2&gt;

&lt;p&gt;Here is the shape of the path I have seen most often, regardless of company or tooling. The platform team owns the middle. The friction lives at the edges.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Step&lt;/th&gt;
&lt;th&gt;Who does it&lt;/th&gt;
&lt;th&gt;Self-service?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Create the repository from a template&lt;/td&gt;
&lt;td&gt;Developer&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get the pipeline running&lt;/td&gt;
&lt;td&gt;Developer, using the shared template&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get a namespace or environment&lt;/td&gt;
&lt;td&gt;Platform team, via ticket&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get a database or queue&lt;/td&gt;
&lt;td&gt;Another team, via a different ticket&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get secrets into the environment&lt;/td&gt;
&lt;td&gt;Security or platform, via request&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get a DNS entry and certificate&lt;/td&gt;
&lt;td&gt;Network team, via email&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Get production deploy permission&lt;/td&gt;
&lt;td&gt;Manager approval, via form&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;First production deploy&lt;/td&gt;
&lt;td&gt;Developer&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The three "yes" rows are what the slide describes. The five "no" rows are what the developer remembers.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this looked like for me&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The first time I actually walked the path myself, start to finish, as if I were a developer on a product team, it took the better part of a week. Not because any single step was slow. Because each step was a request to a different person, and each person had a queue.&lt;/p&gt;

&lt;p&gt;The platform tooling in the middle was genuinely good. Repository template, pipeline, deploy, all worked first time. I still spent most of the week waiting. Nobody on the platform team had ever counted the waiting, because none of it happened in our tools.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  1. Count tickets, not features
&lt;/h2&gt;

&lt;p&gt;The most useful number a platform team can track is not how many features the platform has. It is how many times a developer has to ask another human for something on the way to a first production deploy.&lt;/p&gt;

&lt;p&gt;I call it the ticket count, even though half of those requests are never actual tickets. A Slack message asking for a namespace is a ticket. An email to the network team is a ticket. An approval form is a ticket. If a human has to act before the developer can continue, it counts.&lt;/p&gt;

&lt;p&gt;Then track the second number: how long each of those requests waits. Not how long it takes to fulfil, which is usually minutes. How long it waits in someone's queue before anyone looks at it, which is usually days.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The number that changed the conversation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;When I first put the ticket count in front of the people who owned the platform roadmap, it was somewhere around six human touch points for a new service, with a median wait of a day or two each. That was the first time the roadmap conversation shifted from "what feature should we add" to "which of these requests can we make disappear".&lt;/p&gt;

&lt;p&gt;It also surfaced something uncomfortable: the platform team was not the bottleneck. We were the fastest queue in the chain. The slow ones belonged to teams that had never been told they were part of the developer experience at all.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  2. Most tickets exist because of a default nobody set
&lt;/h2&gt;

&lt;p&gt;When you look at why each request exists, a pattern shows up. Most of them are not there for safety. They are there because nobody decided what the default should be, so a human decides it every time.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A namespace request exists because nobody decided that every repository gets a namespace automatically.&lt;/li&gt;
&lt;li&gt;A resource quota request exists because nobody decided what a reasonable starting quota is.&lt;/li&gt;
&lt;li&gt;A database request exists because nobody decided that a service can provision its own small database within limits.&lt;/li&gt;
&lt;li&gt;A production access request exists because nobody decided that the team that owns a service can deploy it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of these is a policy question that a human answers by hand, dozens of times, usually the same way. Writing the answer down once and letting the tooling apply it is the whole job.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The request that was always approved&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;One of the requests in our chain was for a new namespace with a default quota. When I looked at the history, every single one had been approved, almost always with the same quota. The approval step existed because at some point someone had worried about cluster capacity, and the worry had turned into a permanent human gate.&lt;/p&gt;

&lt;p&gt;We replaced it with a rule: every repository created from the template gets a namespace and a standard quota at creation time, with a documented way to ask for more. The approval disappeared, the ticket disappeared, and the cluster did not run out of capacity. The worry had been real. The gate had never been the right answer to it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  3. Pre-approve the safe path, and only the safe path
&lt;/h2&gt;

&lt;p&gt;The objection to removing gates is always the same: what stops someone doing something dangerous? The answer is not a human. It is a narrow path that is safe by construction, with the gates kept only for leaving that path.&lt;/p&gt;

&lt;p&gt;In practice that means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The template produces a service that is safe to deploy as-is: no public exposure, conservative resources, no elevated permissions.&lt;/li&gt;
&lt;li&gt;Everything on that path is pre-approved. Nobody reviews it, because the review already happened when the template was written.&lt;/li&gt;
&lt;li&gt;Anything off the path, a public endpoint, a privileged container, a larger quota, still needs a human. But that human is now reviewing exceptions, not routine.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is the part that makes security and platform teams comfortable with removing gates. They are not removing review. They are moving it from every request to the template, where it happens once and applies everywhere.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Where the pre-approval broke down&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The thing I got wrong the first time was making the safe path too narrow. The template covered a stateless HTTP service and nothing else. Anything with a queue consumer, a scheduled job, or a database needed the exception route, which meant most real services went through the exception route, which meant we had rebuilt the ticket queue with extra steps.&lt;/p&gt;

&lt;p&gt;The fix was widening the path until it covered what most teams actually built, and treating each new exception request as a signal that the path was still too narrow. The exception route is supposed to be rare. If it is not, the template is wrong, not the developers.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  4. Own the whole path, or at least measure it
&lt;/h2&gt;

&lt;p&gt;The hardest part is that most of the tickets belong to teams the platform team does not control. Networking, security, database operations, a change board. Telling those teams to remove their gates does not work. Showing them where they sit in the developer's week sometimes does.&lt;/p&gt;

&lt;p&gt;What I have seen work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Publish the path as a single diagram&lt;/strong&gt;, with every human touch point marked and the median wait for each. Put it somewhere visible.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ask each owning team one question:&lt;/strong&gt; what would need to be true for this step to be automatic? Usually the answer is a policy nobody has written down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Offer to build the automation&lt;/strong&gt; for them. Most teams are not against self-service. They do not have time to build it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz113exnovc6kd8up736a.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fz113exnovc6kd8up736a.png" alt="Diagram" width="795" height="74"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Each green node was, at one point, a ticket to a different team. None of them stopped being someone's responsibility. They stopped needing a human in the loop for the routine case.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The team that was never asked&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The slowest step on our path was a DNS and certificate request that went to a networking team by email. From the platform side it looked like an immovable external dependency. When I finally talked to that team, they had no idea their queue was on the critical path for every new service. They had automation for internal zones already. It just was not connected to anything developers could reach.&lt;/p&gt;

&lt;p&gt;Connecting it took a couple of weeks and one conversation. The step went from a multi-day wait to a couple of minutes, and the platform team did not build most of it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5. Measure it the way a developer feels it
&lt;/h2&gt;

&lt;p&gt;If the goal is honest self-service, the metric has to be something a developer would recognise. The ones I keep coming back to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Time from repository creation to first production deploy.&lt;/strong&gt; Wall clock, not effort. This is the number that captures the waiting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Human touch points on that path.&lt;/strong&gt; The ticket count from above. Target: zero for the standard path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Exception rate.&lt;/strong&gt; What fraction of new services need the off-path route. If it is high, the path is too narrow.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Support requests per new service.&lt;/strong&gt; How many times the developer had to ask a question that the tooling should have answered.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of these need a new tool. They need someone to walk the path occasionally, as a developer would, and write down what happened.&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Self-service means no human in the loop&lt;/strong&gt;, not just no platform engineer in the loop. One gate blocks the whole path.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Count the tickets, including the ones that are not tickets.&lt;/strong&gt; Slack messages, emails, and approval forms all count.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Most gates are undecided defaults.&lt;/strong&gt; Decide the default once, encode it, and the gate disappears.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pre-approve the safe path and keep review for exceptions.&lt;/strong&gt; Then widen the path until exceptions are rare.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The slow steps usually belong to someone else.&lt;/strong&gt; Show them the path. Offer to build the automation.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What I still do not have a good answer for is the last step on most paths: the production approval that exists for compliance reasons rather than technical ones. I have made it faster and made it clearer, but I have not yet made it go away, and I am not sure it should.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://vivek-itp.github.io/mydevexblog/blog/2026/09/07/self-service-on-a-slide-vs-in-practice/" rel="noopener noreferrer"&gt;https://vivek-itp.github.io/mydevexblog/blog/2026/09/07/self-service-on-a-slide-vs-in-practice/&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>platformengineering</category>
      <category>selfservice</category>
      <category>internaldeveloperplatform</category>
    </item>
  </channel>
</rss>
