<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Nnamdi Felix Ibe</title>
    <description>The latest articles on DEV Community by Nnamdi Felix Ibe (@ndcodes).</description>
    <link>https://dev.to/ndcodes</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3994627%2F3c1e9fc3-98cd-45c4-9e34-3ef72ece6604.jpg</url>
      <title>DEV Community: Nnamdi Felix Ibe</title>
      <link>https://dev.to/ndcodes</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ndcodes"/>
    <language>en</language>
    <item>
      <title>Day 53: A Shared Volume Needs a Shared Path, and Ubuntu 22.04 Has Two Names on Azure</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Wed, 30 Sep 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-53-a-shared-volume-needs-a-shared-path-and-ubuntu-2204-has-two-names-on-azure-4pah</link>
      <guid>https://dev.to/ndcodes/day-53-a-shared-volume-needs-a-shared-path-and-ubuntu-2204-has-two-names-on-azure-4pah</guid>
      <description>&lt;p&gt;Day 53 of DevOps, Day 3 of Azure. Both tasks today came down to the same thing being known by more than one name.&lt;/p&gt;

&lt;p&gt;Two containers shared a volume and still could not see the same file, because they disagreed about where it lived. And an Ubuntu release on Azure turned out to have two valid image names, with an alias quietly choosing one of them for me.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one Azure task. Fix a broken nginx and PHP-FPM pod, then create an Azure VM from the CLI using an image alias. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two containers, one volume, two paths
&lt;/h2&gt;

&lt;p&gt;The pod, &lt;code&gt;nginx-phpfpm&lt;/code&gt;, runs two containers. nginx takes the requests and hands PHP files to PHP-FPM, and both mount one shared volume for the site's files. nginx's configuration lives in a ConfigMap. The site was not working.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe pod nginx-phpfpm
kubectl get configmap nginx-config &lt;span class="nt"&gt;-o&lt;/span&gt; yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ConfigMap set nginx's document root to &lt;code&gt;/var/www/html&lt;/code&gt;, and &lt;code&gt;describe&lt;/code&gt; showed where each container mounted the shared volume. The nginx side matched the root. The PHP-FPM container mounted the volume somewhere else.&lt;/p&gt;

&lt;p&gt;Why does that break the site? Because nginx does not hand PHP-FPM a file. It hands it a path. Nginx passes the script's location in a FastCGI parameter, &lt;code&gt;SCRIPT_FILENAME&lt;/code&gt;, which nginx's own documentation describes as what PHP uses to determine the script name. PHP-FPM then opens that path inside its own container. With the volume mounted elsewhere, the path pointed at a directory where the site's files were not.&lt;/p&gt;

&lt;p&gt;The volume was shared correctly. The two containers just did not agree on where it was.&lt;/p&gt;

&lt;h2&gt;
  
  
  A pod you fix by replacing it
&lt;/h2&gt;

&lt;p&gt;The fix was one line of YAML, and it still needed a new pod. The Kubernetes docs list what an update to an existing pod may change: container images, a couple of timing fields, tolerations and scheduling gates. Volume mounts are not on the list.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get pod nginx-phpfpm &lt;span class="nt"&gt;-o&lt;/span&gt; yaml &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; po.yaml
&lt;span class="c"&gt;# in po.yaml, set the php-fpm-container mountPath to /var/www/html&lt;/span&gt;
kubectl delete pod nginx-phpfpm
kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; po.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;kubectl replace --force -f po.yaml&lt;/code&gt; does the delete and the re-create in one step.&lt;/p&gt;

&lt;p&gt;Then the file the task asked for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl &lt;span class="nb"&gt;cp&lt;/span&gt; /home/thor/index.php nginx-phpfpm:/var/www/html &lt;span class="nt"&gt;-c&lt;/span&gt; php-fpm-container
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I copied it in through the PHP-FPM container, and that is fine, because after the fix both containers mount the same volume at the same path. A shared volume does not care which container a file arrives through. The copy has to come after the new pod exists, though. The file needs to land in the volume the new pod mounts, and if that volume is an &lt;code&gt;emptyDir&lt;/code&gt;, its data is deleted along with the old pod.&lt;/p&gt;

&lt;p&gt;One thing the task did not test, and worth knowing: if the ConfigMap had been the broken part, editing it would not have been enough. A container that mounts a ConfigMap with &lt;code&gt;subPath&lt;/code&gt; never receives updates, and nginx's documentation says configuration changes are not applied until nginx is told to reload or is restarted.&lt;/p&gt;

&lt;h2&gt;
  
  
  An alias picks one of two names
&lt;/h2&gt;

&lt;p&gt;The Azure task was another VM, this time with the image given as an alias, &lt;code&gt;Ubuntu2204&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--image&lt;/span&gt; Ubuntu2204 &lt;span class="nt"&gt;--size&lt;/span&gt; Standard_B2s &lt;span class="nt"&gt;--admin-username&lt;/span&gt; azureuser &lt;span class="nt"&gt;--generate-ssh-keys&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The alias resolved to &lt;code&gt;Canonical:0001-com-ubuntu-server-jammy:22_04-lts-gen2:latest&lt;/code&gt;. Yesterday's 24.04 image was &lt;code&gt;Canonical:ubuntu-24_04-lts:server:latest&lt;/code&gt;. Same publisher, two different naming schemes.&lt;/p&gt;

&lt;p&gt;My notes explained that as Canonical changing its naming between 22.04 and 24.04. Canonical's own documentation says otherwise. The &lt;code&gt;0001-com-ubuntu&lt;/code&gt; offers are its legacy format, and "With Ubuntu 22.04 LTS, Canonical began deploying a unified offer structure". Canonical lists 22.04 under the new scheme too, as &lt;code&gt;Canonical:ubuntu-22_04-lts:server:latest&lt;/code&gt;. So 22.04 has two valid names, and the alias happens to point at the older one.&lt;/p&gt;

&lt;p&gt;That makes the lesson stronger rather than weaker. You cannot derive an image URN from a version number, because one release can have more than one. Use an alias, or look the URN up, and filter the search: Microsoft notes that &lt;code&gt;az vm image list --all&lt;/code&gt; "can take several minutes to produce the entire list". The alias list itself varies by Azure CLI version and cloud, so an alias that works on one machine can be missing on another.&lt;/p&gt;

&lt;p&gt;And check what booted, not just what was requested:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &lt;span class="nt"&gt;-o&lt;/span&gt; &lt;span class="nv"&gt;StrictHostKeyChecking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;no azureuser@&lt;span class="nv"&gt;$VM_IP&lt;/span&gt; &lt;span class="s1"&gt;'hostname; grep PRETTY_NAME /etc/os-release'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;devops-vm
PRETTY_NAME="Ubuntu 22.04.5 LTS"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The image SKU in the VM's configuration says what Azure was asked for. &lt;code&gt;/etc/os-release&lt;/code&gt; says what is actually running. Only the second one is evidence.&lt;/p&gt;

&lt;p&gt;One smaller thing carried over from yesterday. The temporary disk at &lt;code&gt;/mnt&lt;/code&gt; was 8 GiB this time instead of 4, because it comes with the VM size, and Microsoft's size table lists exactly those figures for B1s and B2s. Some newer sizes have no temporary disk at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Same thing, different names
&lt;/h2&gt;

&lt;p&gt;A volume that both containers mount is not a path they both agree on. A release of Ubuntu is not a single image name. In both cases the thing each side depended on existed. They were using different names for it, and nothing flagged the difference until something looked in the wrong place.&lt;/p&gt;

&lt;p&gt;So here is the Day 53 question. Where in your setup do two components have to agree on a path or a name, with nothing checking that they do?&lt;/p&gt;

&lt;p&gt;Day 53 down. Forty-seven to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>azure</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 52: Undo Rolls Forward, and the Disk Already Mounted Is the One Not to Trust</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Mon, 28 Sep 2026 22:03:38 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-52-undo-rolls-forward-and-the-disk-already-mounted-is-the-one-not-to-trust-32bb</link>
      <guid>https://dev.to/ndcodes/day-52-undo-rolls-forward-and-the-disk-already-mounted-is-the-one-not-to-trust-32bb</guid>
      <description>&lt;p&gt;Day 52 of DevOps, Day 2 of Azure. Both tasks today were about commands that do something a little different from what their names suggest.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;kubectl rollout undo&lt;/code&gt; sounds like a rewind. It is not one; it rolls forward to an old template. And &lt;code&gt;az vm create&lt;/code&gt; sounds like it creates a VM. It does, along with a network, a firewall rule, a public address and one more disk than I asked for, and that extra disk is the one not to keep anything on.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one Azure task. Roll a Deployment back to its previous version, then create an Azure VM with a specific image, size and disk type. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Undo rolls forward
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl rollout undo deployment nginx-deployment
kubectl rollout status deployment nginx-deployment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Yesterday's rolling update kept the old ReplicaSet on purpose. Today's task is what it was kept for: a release has a bug, so go back to the previous revision.&lt;/p&gt;

&lt;p&gt;Under the hood, kubectl's rollback code takes the pod template stored in the older ReplicaSet and patches it into the Deployment's &lt;code&gt;.spec.template&lt;/code&gt;. A changed template is exactly what starts a rollout, as yesterday's post covered, so an undo is a new rollout with the same rolling behaviour as any other update. The Kubernetes docs put it briefly: each rollback updates the revision of the Deployment.&lt;/p&gt;

&lt;p&gt;Two consequences are worth knowing.&lt;/p&gt;

&lt;p&gt;Only the template goes back. The docs say "only the Deployment's Pod template part is rolled back". If someone scaled the Deployment or changed its strategy since that revision, those changes stay.&lt;/p&gt;

&lt;p&gt;And the history has a limit. Revisions are the old ReplicaSets, and &lt;code&gt;.spec.revisionHistoryLimit&lt;/code&gt; keeps ten of them by default. Set it to zero, and in the docs' words, "a new Deployment rollout cannot be undone".&lt;/p&gt;

&lt;p&gt;One habit I am adding is &lt;code&gt;kubectl rollout history&lt;/code&gt; before the undo, to see what there is to go back to. Its CHANGE-CAUSE column shows &lt;code&gt;&amp;lt;none&amp;gt;&lt;/code&gt; unless you set the &lt;code&gt;kubernetes.io/change-cause&lt;/code&gt; annotation, and the old &lt;code&gt;--record&lt;/code&gt; flag that used to fill it is deprecated. A history where every revision says &lt;code&gt;&amp;lt;none&amp;gt;&lt;/code&gt; tells you nothing at the moment you most need it.&lt;/p&gt;

&lt;p&gt;The point I would underline, though, is this one. Undo fixes the cluster, not the source of truth. If the buggy image is still in a manifest file, the next &lt;code&gt;kubectl apply&lt;/code&gt; of that file rolls the bug straight back out. The rollback is only finished when the file, or the commit, goes back too.&lt;/p&gt;

&lt;h2&gt;
  
  
  One command, a lot of resources
&lt;/h2&gt;

&lt;p&gt;The Azure task was a single VM: Ubuntu 24.04, &lt;code&gt;Standard_B1s&lt;/code&gt;, a 30 GB Standard HDD disk, SSH access.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;az vm create&lt;/code&gt; does much more than its name says. Alongside the VM, it created a virtual network and subnet, a network security group with an SSH rule, a public IP address, a network interface, an OS disk and a data disk. Microsoft describes this as the CLI using default values to create any required supporting resources.&lt;/p&gt;

&lt;p&gt;The catch is on the way out. Microsoft's page on deleting VMs says that by default, the disks, NICs and public IPs associated with a VM are kept when the VM is deleted. One command builds it all, and the matching delete leaves most of it behind, where a leftover disk or public IP is still billed.&lt;/p&gt;

&lt;p&gt;Coming from AWS, the concept with no direct equivalent is the resource group. Every Azure resource lives in exactly one, and deleting the group deletes everything in it. In a lab that is the cleanest teardown Azure offers.&lt;/p&gt;

&lt;p&gt;A naming trap while I am here: "Standard HDD" in the portal is &lt;code&gt;Standard_LRS&lt;/code&gt; in the CLI, and Standard SSD is &lt;code&gt;StandardSSD_LRS&lt;/code&gt;. One token apart, a different storage medium. &lt;code&gt;LRS&lt;/code&gt; is not the disk type at all. It is the redundancy: locally redundant storage, three copies within one data centre.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three disks, and one to leave alone
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;lsblk&lt;/code&gt; on the new VM showed three disks, not two:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;sda       30G  disk
├─sda1    29G  part /
sdb        4G  disk
└─sdb1     4G  part /mnt
sdc       30G  disk
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;sda&lt;/code&gt; is the OS disk, and &lt;code&gt;sdc&lt;/code&gt; is the data disk, attached but raw. &lt;code&gt;sdb&lt;/code&gt; is the one I did not ask for. It is the temporary disk, and its size comes with the VM size, 4 GiB on a B1s. It arrives formatted and mounted at &lt;code&gt;/mnt&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It looks like free storage, ready to use. Microsoft says data on it "might be lost during a maintenance event, when you redeploy a VM, or when you stop the VM", though it does survive a normal restart. The Linux VM FAQ puts it more bluntly: "Don't use the temporary disk, mounted under (/mnt) to store data."&lt;/p&gt;

&lt;p&gt;The nearest AWS equivalent is instance store, and the rule is the same: nothing you need to keep goes there.&lt;/p&gt;

&lt;p&gt;The data disk arriving raw is normal, the same as an extra EBS volume. Making it usable means partitioning, formatting, mounting, and an &lt;code&gt;/etc/fstab&lt;/code&gt; entry by UUID with &lt;code&gt;nofail&lt;/code&gt;. Microsoft recommends the UUID because device paths like &lt;code&gt;/dev/sdc1&lt;/code&gt; are not persistent and change on reboot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reason I had wrong
&lt;/h2&gt;

&lt;p&gt;The lab host runs as &lt;code&gt;root&lt;/code&gt;, and &lt;code&gt;az vm create&lt;/code&gt; takes its default admin username from the local account. &lt;code&gt;root&lt;/code&gt; is on Microsoft's list of disallowed VM usernames, along with &lt;code&gt;admin&lt;/code&gt;, &lt;code&gt;administrator&lt;/code&gt;, &lt;code&gt;user&lt;/code&gt; and &lt;code&gt;test&lt;/code&gt;. So my notes said the create would fail unless I set a username myself.&lt;/p&gt;

&lt;p&gt;The CLI reference says otherwise. The default is the current OS username, and "If the default value is system reserved, then default value will be set to azureuser." Running as root, the CLI would have picked &lt;code&gt;azureuser&lt;/code&gt; on its own. What does get rejected is a reserved name you pass explicitly.&lt;/p&gt;

&lt;p&gt;I still set &lt;code&gt;--admin-username azureuser&lt;/code&gt;, and I would again, because it makes the result independent of who runs the command. But the reason in my notes was wrong, and a wrong reason is worth correcting even when the command it produced was right.&lt;/p&gt;

&lt;h2&gt;
  
  
  Read what the command did
&lt;/h2&gt;

&lt;p&gt;An undo that rolls forward. A create that builds a small network, and a delete that leaves most of it standing. A disk that is mounted and ready, and not safe to use. In each case the name was a fair summary and a poor specification. The rollout history, the resource list and the documentation had the real answer.&lt;/p&gt;

&lt;p&gt;So here is the Day 52 question. The last time you rolled something back, did you also change the file that would have deployed it again?&lt;/p&gt;

&lt;p&gt;Day 52 down. Forty-eight to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>azure</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 51: A Rollout Only Starts When the Template Changes, and Azure Keeps Only Half Your Key</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Sun, 27 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-51-a-rollout-only-starts-when-the-template-changes-and-azure-keeps-only-half-your-key-5b8o</link>
      <guid>https://dev.to/ndcodes/day-51-a-rollout-only-starts-when-the-template-changes-and-azure-keeps-only-half-your-key-5b8o</guid>
      <description>&lt;p&gt;The AWS track finished yesterday, so from today the cloud half of each post comes from the Azure track instead. The series keeps its name, so everything stays in one place. Day 51 of DevOps, Day 1 of Azure.&lt;/p&gt;

&lt;p&gt;Both tasks today turn on what gets kept. A rolling update keeps the old version around so you can go back to it. Azure keeps one half of your SSH key and hands the other half to you exactly once.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one Azure task. Roll a Deployment to a new image without downtime, then create an SSH key pair for Azure VMs. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  A rollout starts only when the template changes
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl describe deployment nginx-deployment
kubectl &lt;span class="nb"&gt;set &lt;/span&gt;image deployment/nginx-deployment nginx-container&lt;span class="o"&gt;=&lt;/span&gt;nginx:1.17
kubectl rollout status deployment/nginx-deployment
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;describe&lt;/code&gt; comes first for a reason. &lt;code&gt;set image&lt;/code&gt; takes the container name, not the Deployment name, so &lt;code&gt;nginx-container&lt;/code&gt; has to match the container in the pod template exactly. Get it wrong and nothing updates, because there is no such container.&lt;/p&gt;

&lt;p&gt;Kubernetes is precise about what starts a rollout: a Deployment's rollout is triggered if and only if the pod template changes, and other updates, such as scaling, do not trigger one. Changing the image changes the template. Changing the replica count does not.&lt;/p&gt;

&lt;p&gt;The default strategy is RollingUpdate, with two settings that both default to 25%. &lt;code&gt;maxUnavailable&lt;/code&gt; is how far below the desired count you can drop during the update, rounded down. &lt;code&gt;maxSurge&lt;/code&gt; is how far above it you can go, rounded up.&lt;/p&gt;

&lt;p&gt;Those rounding directions are what make a single replica safe. With one pod, 25% rounds down to zero unavailable and up to one extra, so Kubernetes starts the new pod before removing the old one. Even a one-pod Deployment updates without a gap.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;kubectl rollout status&lt;/code&gt; then watches until it finishes, and Kubernetes documents its exit status as 0 on success and 1 when the Deployment exceeds its progress deadline. That is what belongs in a script after an update, instead of a guessed sleep.&lt;/p&gt;

&lt;h2&gt;
  
  
  The old version is kept on purpose
&lt;/h2&gt;

&lt;p&gt;The update does not modify the existing ReplicaSet. It makes a new one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl get replicasets
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The new ReplicaSet scales up, the old one scales down to zero, and the old one stays. Kubernetes keeps ten by default, and that retained ReplicaSet is exactly what &lt;code&gt;kubectl rollout undo&lt;/code&gt; goes back to. Tomorrow's task uses it.&lt;/p&gt;

&lt;p&gt;There were three ways to make this change: &lt;code&gt;kubectl set image&lt;/code&gt;, &lt;code&gt;kubectl edit&lt;/code&gt;, or editing the manifest and running &lt;code&gt;kubectl apply&lt;/code&gt;. All three start the same rollout. They do not leave the same thing behind. The first two change the live object, so a manifest file on disk is now out of date, and the next time someone applies that file it quietly puts the old image back. Only the &lt;code&gt;apply&lt;/code&gt; route keeps the file as the truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Azure keeps only the public half
&lt;/h2&gt;

&lt;p&gt;The Azure task was an SSH key pair for virtual machines.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;az sshkey create &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$KEY_NAME&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="nt"&gt;--resource-group&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$RG&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The resource this creates is a &lt;code&gt;Microsoft.Compute/sshPublicKeys&lt;/code&gt; object, and its only key property is the public key. There is nowhere in it for a private key to live. Microsoft documents that each newly created key is also stored locally, and the output tells you where:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Private key is saved to "/root/.ssh/1757564000_123456".
Public key is saved to "/root/.ssh/1757564000_123456.pub".
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That file is the only copy of the private key that will ever exist. No API returns it. Lose it and the key pair is useless for login, and the fix is a new pair and resetting access on every VM that used the old one.&lt;/p&gt;

&lt;p&gt;The filename is the other trap. It is not &lt;code&gt;id_rsa&lt;/code&gt;, so &lt;code&gt;ssh&lt;/code&gt; will not find it by default. Every login needs &lt;code&gt;-i&lt;/code&gt; with the full path, or the file needs renaming to something memorable straight away.&lt;/p&gt;

&lt;p&gt;Two things in my own notes turned out to be wrong when I checked them.&lt;/p&gt;

&lt;p&gt;I had written that Windows VMs do not accept Ed25519 keys. Microsoft's documentation lists the same two supported types for Windows as for Linux, RSA of at least 2048 bits and Ed25519. The real difference is that Azure does not provision SSH public keys to Windows machines automatically at all, of any type.&lt;/p&gt;

&lt;p&gt;And I had uploaded a key with &lt;code&gt;--public-key @~/.ssh/nautilus-key.pub&lt;/code&gt;. The &lt;code&gt;@&lt;/code&gt; tells the Azure CLI to read the value from a file, which is right. But bash only expands a tilde at the start of a word, and after &lt;code&gt;@&lt;/code&gt; it is not at the start, so the shell passes the tilde through untouched. Microsoft's examples use absolute paths. &lt;code&gt;@$HOME/.ssh/nautilus-key.pub&lt;/code&gt; works because the shell expands &lt;code&gt;$HOME&lt;/code&gt; anywhere in the word.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coming from AWS
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;AWS&lt;/th&gt;
&lt;th&gt;Azure&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Command&lt;/td&gt;
&lt;td&gt;&lt;code&gt;aws ec2 create-key-pair&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;az sshkey create&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Private key delivery&lt;/td&gt;
&lt;td&gt;Returned in &lt;code&gt;KeyMaterial&lt;/code&gt; in the response&lt;/td&gt;
&lt;td&gt;Written to a local file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scope&lt;/td&gt;
&lt;td&gt;Available only in the Region where created&lt;/td&gt;
&lt;td&gt;Resource group plus location&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Default type&lt;/td&gt;
&lt;td&gt;RSA&lt;/td&gt;
&lt;td&gt;RSA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The one-chance rule is the same on both. AWS keeps the public key and returns the private key inside the JSON response, so forgetting to capture &lt;code&gt;KeyMaterial&lt;/code&gt; means it scrolls past once and is gone. Azure writing it to a file is harder to lose, and the filename is easier to misplace.&lt;/p&gt;

&lt;p&gt;One more thing worth knowing. Deleting the Azure key resource does not lock anyone out of VMs made with it. The public key was copied into each VM when it was created, and the VM has no ongoing link back to the resource. The resource is a convenience for provisioning, not a control point for access.&lt;/p&gt;

&lt;h2&gt;
  
  
  What gets kept, and by whom
&lt;/h2&gt;

&lt;p&gt;Kubernetes keeps the old ReplicaSet so you can go back. Azure keeps the public key so it can hand it to new VMs. The private key is kept by exactly one party, you, in a file with an unhelpful name.&lt;/p&gt;

&lt;p&gt;So here is the Day 51 question. For the credentials your systems depend on, do you know where every copy of the private half lives right now?&lt;/p&gt;

&lt;p&gt;Day 51 down. Forty-nine to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>azure</category>
      <category>security</category>
    </item>
    <item>
      <title>Day 50: One Number Picks a Pod's QoS Class, and a Bigger Disk Is Three Separate Jobs</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Sat, 26 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-50-one-number-picks-a-pods-qos-class-and-a-bigger-disk-is-three-separate-jobs-2odp</link>
      <guid>https://dev.to/ndcodes/day-50-one-number-picks-a-pods-qos-class-and-a-bigger-disk-is-three-separate-jobs-2odp</guid>
      <description>&lt;p&gt;Halfway. And today's AWS task is the last one in the 50-day AWS track, so half of the series is finished. More on that at the end.&lt;/p&gt;

&lt;p&gt;The two tasks themselves were both about layers that look like one thing. A resource spec where a single value changes how Kubernetes treats the whole pod. A disk that has to be made bigger three separate times before the operating system notices.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one AWS task. Set resource requests and limits on a pod, then expand an EC2 root volume without stopping the instance. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Requests, limits, and the one number
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;requests&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;15Mi"&lt;/span&gt;
    &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;100m"&lt;/span&gt;
  &lt;span class="na"&gt;limits&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;memory&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;memory&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;limit&amp;gt;"&lt;/span&gt;
    &lt;span class="na"&gt;cpu&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;100m"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A request is what the scheduler reserves when it places the pod. A limit is the ceiling once it is running. And the two limits are enforced in very different ways. Kubernetes documents CPU limits as enforced by throttling, so a container over its CPU limit just runs slower. Memory limits are enforced by the kernel with out-of-memory kills, so a container over its memory limit may be terminated, though only when the kernel detects memory pressure, not necessarily at once. CPU over the limit is slow. Memory over the limit is dead.&lt;/p&gt;

&lt;p&gt;The units deserve a second look too. &lt;code&gt;100m&lt;/code&gt; is a tenth of a CPU, one hundred millicpu. &lt;code&gt;15Mi&lt;/code&gt; is 15 mebibytes. And Kubernetes' own warning about case is worth quoting in spirit: &lt;code&gt;400m&lt;/code&gt; of memory is a request for 0.4 bytes, when whoever typed it almost certainly meant &lt;code&gt;400Mi&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Now the one number. I wrote the memory limit down as &lt;code&gt;15Mi&lt;/code&gt;, the same as the request, but I have not confirmed that was the figure the task gave, and it matters more than any other value in the file.&lt;/p&gt;

&lt;p&gt;Kubernetes assigns each pod a QoS class from its requests and limits. A pod is Guaranteed only if every container has both CPU and memory requests and limits, all above zero, with each limit equal to its request. It is Burstable if it misses that but has at least one request or limit, and BestEffort if it has none at all.&lt;/p&gt;

&lt;p&gt;The CPU request and limit here are both &lt;code&gt;100m&lt;/code&gt;. So if the memory limit is &lt;code&gt;15Mi&lt;/code&gt;, every limit equals its request, and the pod is Guaranteed. If the memory limit is anything higher, the same pod is Burstable. One value, two classes.&lt;/p&gt;

&lt;p&gt;And the class has consequences. The documentation is explicit: when a node runs out of resources, Kubernetes evicts BestEffort pods first, then Burstable, and Guaranteed last. You can read the result with &lt;code&gt;kubectl describe pod&lt;/code&gt;, on the &lt;code&gt;QoS Class&lt;/code&gt; line.&lt;/p&gt;

&lt;p&gt;One small thing I appreciated. My first draft of this manifest had a typo, &lt;code&gt;rrequests&lt;/code&gt;. kubectl's &lt;code&gt;--validate&lt;/code&gt; flag defaults to &lt;code&gt;strict&lt;/code&gt;, which rejects unknown fields rather than dropping them, so that typo fails at apply instead of producing a pod with no requests and a very confusing QoS class.&lt;/p&gt;

&lt;h2&gt;
  
  
  A bigger disk is three jobs
&lt;/h2&gt;

&lt;p&gt;The AWS task was to grow a root volume from 8 GiB to 12 GiB and have the instance see the space, without disrupting it.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;EBS volume&lt;/td&gt;
&lt;td&gt;&lt;code&gt;aws ec2 modify-volume&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The virtual disk is bigger&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Partition&lt;/td&gt;
&lt;td&gt;&lt;code&gt;growpart&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;The partition uses the new space&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Filesystem&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;xfs_growfs&lt;/code&gt; or &lt;code&gt;resize2fs&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The filesystem uses the bigger partition&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each one is separate, and AWS's procedure says so: before you can extend a filesystem, you must extend the partition, if the volume has one. Do only the first, and you have a 12 GiB disk holding an 8 GiB partition holding an 8 GiB filesystem, and &lt;code&gt;df&lt;/code&gt; still says 8G.&lt;/p&gt;

&lt;p&gt;I could see the middle state directly, which is what convinced me the layers are real. After &lt;code&gt;growpart&lt;/code&gt;, &lt;code&gt;lsblk&lt;/code&gt; showed the partition at 12G while &lt;code&gt;df&lt;/code&gt; still showed the filesystem at 8G. Only &lt;code&gt;xfs_growfs&lt;/code&gt; closed the gap.&lt;/p&gt;

&lt;p&gt;Three details from that sequence worth keeping.&lt;/p&gt;

&lt;p&gt;You do not have to wait for the modification to finish. AWS documents that size increases take effect once the modification reaches the &lt;code&gt;optimizing&lt;/code&gt; state, usually within seconds, and that you can extend the partition and filesystem as soon as it does. Waiting for &lt;code&gt;completed&lt;/code&gt; burns time for nothing.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;growpart&lt;/code&gt; takes two arguments. The AWS procedure asks you to note the space between the device and the partition number: &lt;code&gt;growpart /dev/xvda 1&lt;/code&gt;, not &lt;code&gt;growpart /dev/xvda1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;And the two filesystem tools take different arguments. &lt;code&gt;xfs_growfs&lt;/code&gt; wants the mount point, &lt;code&gt;resize2fs&lt;/code&gt; wants the partition device. Check &lt;code&gt;df -hT&lt;/code&gt; first, because Amazon Linux 2023 is xfs, Ubuntu is ext4, and using the wrong one produces an error that reads like corruption.&lt;/p&gt;

&lt;h2&gt;
  
  
  The rule I had wrong
&lt;/h2&gt;

&lt;p&gt;My notes said a volume cannot be modified again for six hours after a change. That rule is out of date.&lt;/p&gt;

&lt;p&gt;AWS's current documentation says you must wait for a modification to reach &lt;code&gt;completed&lt;/code&gt; before starting another on the same volume, and that you can modify a volume at most four times in a rolling 24-hour period. The six hours seems to come from a different sentence on the same page, that a 1 TiB volume can typically take up to six hours to modify.&lt;/p&gt;

&lt;p&gt;The practical point survives, just for a different reason. EBS volumes cannot be shrunk at all, and XFS cannot be shrunk either. And you cannot iterate freely: the next change waits for the last to finish, which on a large volume can be hours, and you only get four a day. Pick the size carefully the first time.&lt;/p&gt;

&lt;p&gt;This is exactly why I now check every note against the current documentation before it goes out. The rule I had was plausible, specific, and wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  Halfway, and the end of the AWS track
&lt;/h2&gt;

&lt;p&gt;Fifty tasks of AWS are done, and the cloud half of this series moves to Azure from tomorrow, under the same series name.&lt;/p&gt;

&lt;p&gt;Looking back across the fifty, three habits came up often enough to be worth naming.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Verify the deliverable, not a status field.&lt;/strong&gt; An ECS task can read &lt;code&gt;RUNNING&lt;/code&gt; and be unreachable. A route can look right and read &lt;code&gt;blackhole&lt;/code&gt;. A modification can be &lt;code&gt;optimising&lt;/code&gt; while the filesystem is still 8G. The check that means something is the one that exercises the actual promise: the &lt;code&gt;curl&lt;/code&gt;, the &lt;code&gt;cmp&lt;/code&gt;, the &lt;code&gt;df&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The console does things the CLI makes you do yourself.&lt;/strong&gt; A DB subnet group, an instance profile, a listener, a resource-based permission on a Lambda. Each was invisible in the console and a separate, silent failure on the CLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Errors name the call that failed, not the cause.&lt;/strong&gt; A missing zip reported three commands after a broken heredoc. &lt;code&gt;AccessDenied&lt;/code&gt; that meant a public access block. &lt;code&gt;Unable to validate the following destination configurations&lt;/code&gt; meant a missing Lambda permission. The fix was rarely where the message pointed.&lt;/p&gt;

&lt;p&gt;None of that is specific to AWS, which is a good sign for the Azure half.&lt;/p&gt;

&lt;p&gt;So here is the Day 50 question. What is a rule you are confident about that you have not checked against the documentation since you learned it?&lt;/p&gt;

&lt;p&gt;Day 50 down. Fifty to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>aws</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Day 49: A Deployment Keeps Pods Alive, and Routable Is Not Reachable</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Fri, 25 Sep 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-49-a-deployment-keeps-pods-alive-and-routable-is-not-reachable-2noh</link>
      <guid>https://dev.to/ndcodes/day-49-a-deployment-keeps-pods-alive-and-routable-is-not-reachable-2noh</guid>
      <description>&lt;p&gt;Today was layers. A Deployment that does not run pods itself, but owns a ReplicaSet that does. A peering connection that makes a subnet routable, a security group that decides whether it is reachable, and an SSH hop in between that did not carry the key I handed it.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one AWS task. Run an application as a Deployment, then move a log file from a private VPC to S3 through a peered public one. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The layer you forget is there
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl create deployment httpd &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;httpd:latest
kubectl get deployments
kubectl get replicasets
kubectl get pods
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Kubernetes describes a Deployment as providing declarative updates for Pods and ReplicaSets, and the ReplicaSet is the part that answers yesterday's question. A ReplicaSet replaces pods that are deleted or terminated for any reason. Delete one of this Deployment's pods and a replacement appears within seconds, with a different name.&lt;/p&gt;

&lt;p&gt;The names give the structure away. A Deployment's pods are called something like &lt;code&gt;httpd-&amp;lt;hash&amp;gt;-&amp;lt;suffix&amp;gt;&lt;/code&gt;, where the hash identifies the ReplicaSet that made them. The name changes every time a pod is replaced, so nothing should ever depend on it.&lt;/p&gt;

&lt;p&gt;What the Deployment actually uses to find its pods is a label. &lt;code&gt;kubectl create deployment&lt;/code&gt; gives them &lt;code&gt;app=httpd&lt;/code&gt;, and &lt;code&gt;kubectl get pods -l app=httpd&lt;/code&gt; lists exactly this Deployment's pods and nothing else.&lt;/p&gt;

&lt;p&gt;Two facts worth having before Day 51. Scaling is not an update: Kubernetes is explicit that a rollout is triggered only when the pod template changes, so &lt;code&gt;kubectl scale&lt;/code&gt; just changes a count. And old ReplicaSets are kept after an update, ten by default, which is what makes rolling back possible at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four things have to say yes
&lt;/h2&gt;

&lt;p&gt;The AWS task had an instance with no internet access writing a log file, which had to end up in S3. The route was across a VPC peering connection to an instance that did have internet access, and from there to the bucket.&lt;/p&gt;

&lt;p&gt;The private instance had no public IP, no NAT and no internet route. The only way to reach it at all was through the public one, over the peering connection, so the networking had to be finished before a single line of configuration could reach the private side. The plumbing was not a step, it was the prerequisite.&lt;/p&gt;

&lt;p&gt;And peering turned out to need four separate things, each documented, each capable of failing on its own:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;CIDRs that do not overlap.&lt;/strong&gt; AWS will not peer VPCs with matching or overlapping ranges, and there is no translation option. The private VPC was &lt;code&gt;10.10.0.0/16&lt;/code&gt;, so the new one became &lt;code&gt;10.20.0.0/16&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Acceptance.&lt;/strong&gt; Even within one account the connection sits in &lt;code&gt;pending-acceptance&lt;/code&gt; until someone accepts it. You can create and accept it yourself, but you do have to accept it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Routes on both sides.&lt;/strong&gt; AWS requires a route in the route tables for both instances' subnets. One direction only gives you packets that arrive and replies that cannot get home, and the symptom is a timeout identical to having no route at all.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security groups.&lt;/strong&gt; The one people miss after the routes are right. AWS's own guidance is to update the security group rules so traffic to and from the peer VPC is not restricted.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The fourth is the title. Peering made the private instance routable from the public VPC. Without a security group rule allowing &lt;code&gt;10.20.0.0/16&lt;/code&gt;, SSH would have been dropped at the instance with the routing working perfectly. Routes and security groups are separate layers and both have to agree.&lt;/p&gt;

&lt;p&gt;It is also worth knowing that peering does not chain. AWS says plainly that it does not support transitive peering relationships, so a third VPC peered to the public one would still have no path to the private one.&lt;/p&gt;

&lt;h2&gt;
  
  
  The key that did not make the hop
&lt;/h2&gt;

&lt;p&gt;This is the one that actually cost time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ssh &lt;span class="nt"&gt;-i&lt;/span&gt; /root/.ssh/datacenter-key.pem &lt;span class="nt"&gt;-J&lt;/span&gt; ubuntu@&lt;span class="nv"&gt;$PUB_IP&lt;/span&gt; ubuntu@10.10.1.114 &lt;span class="s1"&gt;'hostname'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ubuntu@32.196.116.111: Permission denied &lt;span class="o"&gt;(&lt;/span&gt;publickey&lt;span class="o"&gt;)&lt;/span&gt;&lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seconds earlier, the same key had logged into that same public instance directly. The error names the jump host, which is the only clue that the private instance was never reached.&lt;/p&gt;

&lt;p&gt;The ssh manual has the answer, in a note under &lt;code&gt;-J&lt;/code&gt;: configuration directives supplied on the command line generally apply to the destination host and not to any jump hosts, and you should use &lt;code&gt;~/.ssh/config&lt;/code&gt; to configure jump hosts. So &lt;code&gt;-i&lt;/code&gt; went to the private instance. The public one was only offered the default key files and anything in an ssh agent, and none of those was right.&lt;/p&gt;

&lt;p&gt;The fix is to stop passing it on the command line:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ssh"&gt;&lt;code&gt;&lt;span class="k"&gt;Host&lt;/span&gt; pubec2
    &lt;span class="k"&gt;HostName&lt;/span&gt; &lt;span class="m"&gt;32&lt;/span&gt;.196.116.111
    &lt;span class="k"&gt;User&lt;/span&gt; ubuntu
    &lt;span class="k"&gt;IdentityFile&lt;/span&gt; /root/.ssh/datacenter-key.pem

&lt;span class="k"&gt;Host&lt;/span&gt; privec2
    &lt;span class="k"&gt;HostName&lt;/span&gt; &lt;span class="m"&gt;10&lt;/span&gt;.10.1.114
    &lt;span class="k"&gt;User&lt;/span&gt; ubuntu
    &lt;span class="k"&gt;IdentityFile&lt;/span&gt; /root/.ssh/datacenter-key.pem
    &lt;span class="k"&gt;ProxyJump&lt;/span&gt; pubec2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every hop carries its own identity, and &lt;code&gt;ssh privec2&lt;/code&gt; works as though the private instance were directly reachable.&lt;/p&gt;

&lt;h2&gt;
  
  
  One more silent failure
&lt;/h2&gt;

&lt;p&gt;The log shipping ran from &lt;code&gt;/etc/cron.d&lt;/code&gt;, and that directory has a trap:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight conf"&gt;&lt;code&gt;* * * * * &lt;span class="n"&gt;root&lt;/span&gt; &lt;span class="n"&gt;scp&lt;/span&gt; ... /&lt;span class="n"&gt;var&lt;/span&gt;/&lt;span class="n"&gt;log&lt;/span&gt;/&lt;span class="n"&gt;boots&lt;/span&gt;.&lt;span class="n"&gt;log&lt;/span&gt; &lt;span class="n"&gt;ubuntu&lt;/span&gt;@&lt;span class="m"&gt;10&lt;/span&gt;.&lt;span class="m"&gt;20&lt;/span&gt;.&lt;span class="m"&gt;1&lt;/span&gt;.&lt;span class="m"&gt;134&lt;/span&gt;:/&lt;span class="n"&gt;home&lt;/span&gt;/&lt;span class="n"&gt;ubuntu&lt;/span&gt;/&lt;span class="n"&gt;boots&lt;/span&gt;.&lt;span class="n"&gt;log&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six fields before the command, not five. The crontab manual describes five time fields followed by a username in the system crontab, and says jobs in &lt;code&gt;cron.d&lt;/code&gt; are system jobs that need that username. A &lt;code&gt;crontab -e&lt;/code&gt; entry has no user field because the user is whoever owns the crontab. Put a five-field line in &lt;code&gt;/etc/cron.d&lt;/code&gt;, and cron reads &lt;code&gt;scp&lt;/code&gt; as a username, logs a failure to syslog, and never runs anything.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checking each stage
&lt;/h2&gt;

&lt;p&gt;The file landed. I checked both halves separately, the copy on the public instance and the object in S3, because on a two-stage pipeline that one extra &lt;code&gt;ls&lt;/code&gt; halves the search when something breaks. Nothing on the public instance points at the peering, the security group or the first cron job. A file there but not in S3 points at the role or the second.&lt;/p&gt;

&lt;p&gt;I should also name what the task required and what I would not repeat. Both instances used the same key pair, so the private instance ended up holding a private key that also opens the public one. Compromise either and you have both. In anything real, this would be SSM Session Manager and CloudWatch Logs, and no instance would hold another's credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Each layer has its own no
&lt;/h2&gt;

&lt;p&gt;The Deployment works because a layer I rarely look at does the replacing. The peering works because four independent layers each agreed. And the SSH hop failed because an option I thought was global only applied to one layer of the connection.&lt;/p&gt;

&lt;p&gt;So here is the Day 49 question. When something is unreachable, how many separate layers are you checking before you decide which one is wrong?&lt;/p&gt;

&lt;p&gt;Day 49 down. Fifty-one to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>aws</category>
      <category>networking</category>
    </item>
    <item>
      <title>Day 48: Nobody Replaces a Bare Pod, and a Stack Knows How to Take Itself Apart</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Wed, 23 Sep 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-48-nobody-replaces-a-bare-pod-and-a-stack-knows-how-to-take-itself-apart-30dl</link>
      <guid>https://dev.to/ndcodes/day-48-nobody-replaces-a-bare-pod-and-a-stack-knows-how-to-take-itself-apart-30dl</guid>
      <description>&lt;p&gt;Kubernetes starts today, and it starts with the one object that has nobody looking after it. On the other side, the same Lambda I built by hand on Day 33, rebuilt as a template that knows how to create itself and how to tear itself down.&lt;/p&gt;

&lt;p&gt;Both are about who is responsible for a thing once it exists.&lt;/p&gt;

&lt;p&gt;One Kubernetes task, one AWS task. Deploy a pod, then deploy a Lambda through CloudFormation. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two routes to a pod manifest
&lt;/h2&gt;

&lt;p&gt;The fast route generates the YAML:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl run pod-httpd &lt;span class="nt"&gt;--image&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;httpd:latest &lt;span class="nt"&gt;--labels&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"app=httpd_app"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dry-run&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;client &lt;span class="nt"&gt;-o&lt;/span&gt; yaml &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; pod.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;--dry-run=client -o yaml&lt;/code&gt; prints the object kubectl would send without sending it. The Kubernetes quick reference uses exactly this pattern to generate a spec file, and it saves typing all the boilerplate.&lt;/p&gt;

&lt;p&gt;The careful route writes it by hand:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Pod&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pod-httpd&lt;/span&gt;
  &lt;span class="na"&gt;labels&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;app&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;httpd_app&lt;/span&gt;
&lt;span class="na"&gt;spec&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;containers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;httpd-container&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;httpd:latest&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;They do not produce the same file, and that caught me. &lt;code&gt;kubectl run&lt;/code&gt; names the container after the pod, so the generated manifest has a container called &lt;code&gt;pod-httpd&lt;/code&gt;. The task wanted &lt;code&gt;httpd-container&lt;/code&gt;. The generated file is a starting point to edit, not a finished answer.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;kubectl apply &lt;span class="nt"&gt;-f&lt;/span&gt; pod.yaml
kubectl describe pod pod-httpd
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two small things worth having. kubectl validates manifests strictly by default: &lt;code&gt;--validate&lt;/code&gt; defaults to &lt;code&gt;strict&lt;/code&gt;, which fails the request on unknown fields rather than silently dropping them, so a misspelled key is an error you see. And when a pod will not start, the reason is in the Events section at the bottom of &lt;code&gt;describe&lt;/code&gt;, not in the one-word status from &lt;code&gt;get&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nobody is watching this pod
&lt;/h2&gt;

&lt;p&gt;Here is the part that matters. Delete &lt;code&gt;pod-httpd&lt;/code&gt;, or lose the node it runs on, and it is gone. Nothing brings it back.&lt;/p&gt;

&lt;p&gt;Kubernetes documents this directly: a ReplicaSet replaces pods that are deleted or terminated for any reason, such as node failure, unlike pods a user created directly. It goes as far as recommending a ReplicaSet even when you only need one pod.&lt;/p&gt;

&lt;p&gt;So a bare pod is useful for exactly what I did with it, which is learn what a pod is. Anything that should still be running tomorrow needs something above it whose job is to notice when it is not. That is tomorrow's task.&lt;/p&gt;

&lt;h2&gt;
  
  
  The same Lambda, declared instead of commanded
&lt;/h2&gt;

&lt;p&gt;On Day 33 this Lambda was six CLI calls and a retry loop, the loop because &lt;code&gt;create-function&lt;/code&gt; failed when the IAM role had been created seconds earlier and had not propagated yet. Today it is one template and one &lt;code&gt;create-stack&lt;/code&gt;, and it deployed first time with no retry anywhere.&lt;/p&gt;

&lt;p&gt;I want to be careful about what that proves. CloudFormation documents that in its default mode it waits for each resource to reach a fully stabilised state before reporting it complete, which fits what happened. It does not document anything specific to IAM propagation for a role a Lambda uses in the same stack. So "no retry loop needed" is what I saw, not something I can point to as a promise. That distinction matters the day it fails.&lt;/p&gt;

&lt;p&gt;What the template genuinely removes is the ordering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;Role&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kt"&gt;!GetAtt&lt;/span&gt; &lt;span class="s"&gt;LambdaExecutionRole.Arn&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AWS's rule is that when one resource refers to another with &lt;code&gt;!Ref&lt;/code&gt;, &lt;code&gt;!GetAtt&lt;/code&gt; or &lt;code&gt;!Sub&lt;/code&gt;, the referenced one is created first and deleted last. That single reference is the whole dependency graph. &lt;code&gt;DependsOn&lt;/code&gt; is for the cases where a real dependency exists that nothing in the template points at.&lt;/p&gt;

&lt;p&gt;And one detail that looks arbitrary until you read why:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;Handler&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;index.lambda_handler&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When code is supplied inline, CloudFormation places it in a file named &lt;code&gt;index&lt;/code&gt;, so the handler's module part has to be &lt;code&gt;index&lt;/code&gt;. Only the function name after the dot is yours.&lt;/p&gt;

&lt;p&gt;Two more habits carried in from yesterday. The template uses managed policies only, because Day 47's account refused inline ones. And before deploying, one &lt;code&gt;aws iam get-role&lt;/code&gt; to make sure no role called &lt;code&gt;lambda_execution_role&lt;/code&gt; already existed, since the template names it explicitly and a collision would leave the stack in &lt;code&gt;ROLLBACK_COMPLETE&lt;/code&gt;, which only a delete gets you out of.&lt;/p&gt;

&lt;h2&gt;
  
  
  Taking it apart
&lt;/h2&gt;

&lt;p&gt;Day 33's teardown was to delete the function, detach the policy, delete the role, in that order, and get it wrong at each step.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws cloudformation delete-stack &lt;span class="nt"&gt;--stack-name&lt;/span&gt; datacenter-lambda-app
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same references that ordered the build order the teardown, in reverse. The stack knows what it made and how to unmake it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who is responsible once it exists
&lt;/h2&gt;

&lt;p&gt;A bare pod has nobody. A CloudFormation stack is its own record of what exists and how it fits together. Most of the difference between a lab and a system is whether something is in that role.&lt;/p&gt;

&lt;p&gt;So here is the Day 48 question. For the thing you deployed most recently, if it disappeared tonight, what would you notice, and what would put it back?&lt;/p&gt;

&lt;p&gt;Day 48 down. Fifty-two to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>kubernetes</category>
      <category>aws</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 47: Copy Requirements Before the Code, and SQS Has No Priority Feature</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Mon, 21 Sep 2026 21:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-47-copy-requirements-before-the-code-and-sqs-has-no-priority-feature-5363</link>
      <guid>https://dev.to/ndcodes/day-47-copy-requirements-before-the-code-and-sqs-has-no-priority-feature-5363</guid>
      <description>&lt;p&gt;Neither of today's tasks has a setting for the thing it delivers. Docker has no switch that makes rebuilds fast, and SQS has no switch that makes a queue high priority. In both cases, the feature is the order you put things in.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Package a Python app as an image, then build priority queues with SQS and SNS in a CloudFormation template. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;p&gt;This is also the last Docker task. Days 35 to 47 went from installing the engine to a real application image, and Kubernetes starts tomorrow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two COPY lines, and which one goes first
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;FROM&lt;/span&gt;&lt;span class="s"&gt; python:3.11-slim&lt;/span&gt;
&lt;span class="k"&gt;WORKDIR&lt;/span&gt;&lt;span class="s"&gt; /app&lt;/span&gt;

&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; src/requirements.txt .&lt;/span&gt;
&lt;span class="k"&gt;RUN &lt;/span&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--no-cache-dir&lt;/span&gt; &lt;span class="nt"&gt;-r&lt;/span&gt; requirements.txt

&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; src/ .&lt;/span&gt;

&lt;span class="k"&gt;EXPOSE&lt;/span&gt;&lt;span class="s"&gt; 3004&lt;/span&gt;
&lt;span class="k"&gt;CMD&lt;/span&gt;&lt;span class="s"&gt; ["python", "server.py"]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The obvious version copies all of &lt;code&gt;src/&lt;/code&gt; and then installs. It works. It also reinstalls every dependency every time you touch a line of &lt;code&gt;server.py&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Docker's cache guidance explains why. A change causes a rebuild for every step that follows it, so expensive steps belong near the beginning, and installing dependencies in an earlier layer means there is no need to rebuild that layer when a project file changes. Copy the manifest on its own, install, and only then copy the code. Now editing &lt;code&gt;server.py&lt;/code&gt; invalidates the last &lt;code&gt;COPY&lt;/code&gt; and nothing before it.&lt;/p&gt;

&lt;p&gt;The same logic explains &lt;code&gt;--no-cache-dir&lt;/code&gt;, and it is a nice detail. pip's own documentation recommends leaving its cache on unless you have caching at a higher level, and it gives layered caches in container builds as the example. That is precisely this situation. The Docker layer is the cache; a pip cache sitting inside it is just weight in the image.&lt;/p&gt;

&lt;p&gt;One thing worth knowing about that last line. &lt;code&gt;CMD&lt;/code&gt; in exec form is a JSON array and each element is taken literally, so a stray character in &lt;code&gt;"server.py"&lt;/code&gt; is not a syntax error. It is a different filename. The container starts, Python cannot find the file, and it exits immediately. &lt;code&gt;docker logs&lt;/code&gt; shows it; &lt;code&gt;docker ps&lt;/code&gt; without &lt;code&gt;-a&lt;/code&gt; shows nothing at all.&lt;/p&gt;

&lt;p&gt;Then the run, with the port mapping reading host first, container second:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; 8097:3004 &lt;span class="nt"&gt;--name&lt;/span&gt; pythonapp_nautilus nautilus/python-app
curl http://localhost:8097/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  CloudFormation, and the permission that was not there
&lt;/h2&gt;

&lt;p&gt;The AWS task was the first infrastructure-as-code task of the run: two SQS queues, an SNS topic routing by priority, and a Lambda that drains high before low, all in one template.&lt;/p&gt;

&lt;p&gt;The first deploy rolled back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;User: ... is not authorized to perform: iam:PutRolePolicy on resource: role lambda_execution_role
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The template gave the role permissions two ways. CloudFormation documents &lt;code&gt;Policies&lt;/code&gt; as adding an inline policy embedded in the role, and &lt;code&gt;ManagedPolicyArns&lt;/code&gt; as attaching standalone managed policies. The error named the IAM action the inline route used, &lt;code&gt;PutRolePolicy&lt;/code&gt;. Attaching a managed policy is a different action, and this account allowed one and refused the other. The role had been created and the managed policy attached before the refusal, which is how you tell exactly which half was blocked.&lt;/p&gt;

&lt;p&gt;Moving everything to managed policies fixed it, at a cost worth stating. The managed policies available are &lt;code&gt;AmazonSQSFullAccess&lt;/code&gt; and &lt;code&gt;AmazonSNSFullAccess&lt;/code&gt;, far broader than a function that needs to receive and delete messages on two queues. In an account that allows it, a separate &lt;code&gt;AWS::IAM::ManagedPolicy&lt;/code&gt; resource keeps the permissions scoped while still attaching rather than embedding.&lt;/p&gt;

&lt;p&gt;Two CloudFormation lessons came with the failure.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;validate-template&lt;/code&gt; had passed. The documentation is blunt that it is designed to check only the syntax of your template, not whether the property values are valid, and that to check operational validity you have to create the stack. Passing validation means the YAML parses.&lt;/p&gt;

&lt;p&gt;And the failed stack was stuck. &lt;code&gt;ROLLBACK_COMPLETE&lt;/code&gt; only exists after a failed stack creation, and AWS says the only operation available in that state is a delete. So the first deploy of a new stack is the one that costs a full delete-and-recreate cycle when it goes wrong. A failed update is kinder, returning to the previous working state as &lt;code&gt;UPDATE_ROLLBACK_COMPLETE&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The priority is in the consumer
&lt;/h2&gt;

&lt;p&gt;SQS has no priority feature. The whole mechanism is two queues and a consumer that looks at one first:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;poll_messages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;high_priority_queue&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;No more messages to poll&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;poll_messages&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;low_priority_queue&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What gets messages into the right queue is the SNS filter policy on each subscription:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;FilterPolicy&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;priority&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;high&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AWS describes it simply: if the message attributes satisfy the filter policy, SNS sends the message to that subscriber, and otherwise it does not. Without it, SNS fan-out sends every message to every queue and the exercise means nothing.&lt;/p&gt;

&lt;p&gt;So I checked the queue depths before invoking anything. Two and two, not four and four. That proves the routing independently of the Lambda, which matters, because four and four would mean a filter problem that no amount of debugging the function could fix.&lt;/p&gt;

&lt;p&gt;Then four invocations:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;invoke 1: "Message 'High Priority message 2' deleted"
invoke 2: "Message 'High Priority message 1' deleted"
invoke 3: "Message 'Low Priority message 2' deleted"
invoke 4: "Message 'Low Priority message 1' deleted"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;High before low, exactly as required. And message 2 before message 1 in both queues, which is worth sitting with rather than dismissing.&lt;/p&gt;

&lt;p&gt;AWS documents standard queues as delivering at least once, with messages occasionally arriving out of order because of the distributed architecture, while making a best-effort attempt to keep the send order. So this design gives strict ordering between priorities and none within one. FIFO queues would fix that, with a &lt;code&gt;.fifo&lt;/code&gt; name and a &lt;code&gt;MessageGroupId&lt;/code&gt; on every send, and even then the guarantee is scoped to a message group. Messages in different groups can still arrive out of order relative to each other.&lt;/p&gt;

&lt;h2&gt;
  
  
  The feature is the arrangement
&lt;/h2&gt;

&lt;p&gt;The Dockerfile is fast to rebuild because of which line comes first. The priority system works because of which queue is polled first. Neither is configured anywhere; both are the order.&lt;/p&gt;

&lt;p&gt;So here is the Day 47 question. In the last thing you built, which behaviour depends on the order of steps rather than on any setting, and would anyone reading it later know?&lt;/p&gt;

&lt;p&gt;Day 47 down. Fifty-three to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 46: Compose Wires the Whole Stack, and Permission to Invoke Is Not Permission to Act</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Sun, 20 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-46-compose-wires-the-whole-stack-and-permission-to-invoke-is-not-permission-to-act-1o1</link>
      <guid>https://dev.to/ndcodes/day-46-compose-wires-the-whole-stack-and-permission-to-invoke-is-not-permission-to-act-1o1</guid>
      <description>&lt;p&gt;Two things today needed wiring on both sides, and in both cases getting one side right produces something that looks finished and does nothing.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Stand up a PHP and MariaDB stack with Compose, then build an event pipeline where an S3 upload triggers a Lambda that copies the object and logs it. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two services, one command
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;web&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;php:7.2-apache&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;php_host&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;8085:80"&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/var/www/html:/var/www/html&lt;/span&gt;

  &lt;span class="na"&gt;db&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mariadb:latest&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;mysql_host&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;3306:3306"&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/var/lib/mysql:/var/lib/mysql&lt;/span&gt;
    &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;MYSQL_DATABASE&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;database_host&lt;/span&gt;
      &lt;span class="na"&gt;MYSQL_USER&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;&amp;lt;username&amp;gt;&lt;/span&gt;
      &lt;span class="na"&gt;MYSQL_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;password&amp;gt;"&lt;/span&gt;
      &lt;span class="na"&gt;MYSQL_RANDOM_ROOT_PASSWORD&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;yes"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;sudo &lt;/span&gt;docker compose &lt;span class="nt"&gt;-f&lt;/span&gt; /opt/sysops/docker-compose.yml up &lt;span class="nt"&gt;-d&lt;/span&gt;
curl http://localhost:8085
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One &lt;code&gt;up -d&lt;/code&gt; creates a project network, both containers and both mounts. The two containers can reach each other by service name over that network, which is the Day 42 embedded DNS point finally doing the job it exists for. Nothing here declares that relationship; it comes free with the project network.&lt;/p&gt;

&lt;p&gt;Worth separating two names that look interchangeable and are not. The &lt;code&gt;db&lt;/code&gt; service is addressable from &lt;code&gt;web&lt;/code&gt; as &lt;code&gt;db&lt;/code&gt;, the service key. &lt;code&gt;container_name&lt;/code&gt; only fixes what &lt;code&gt;docker ps&lt;/code&gt; and &lt;code&gt;docker exec&lt;/code&gt; see. Pointing a connection string at &lt;code&gt;mysql_host&lt;/code&gt; works because the container name resolves too, but &lt;code&gt;db&lt;/code&gt; is the portable one.&lt;/p&gt;

&lt;p&gt;Note also that this mount targets &lt;code&gt;/var/www/html&lt;/code&gt; while yesterday's &lt;code&gt;httpd&lt;/code&gt; file wanted &lt;code&gt;/usr/local/apache2/htdocs&lt;/code&gt;. Same idea, different image convention, and it is worth reading the image's documentation rather than assuming a path you have seen before.&lt;/p&gt;

&lt;p&gt;The one that catches people: those MariaDB environment variables are the image's initialisation contract, and they only apply on first start. The image initialises the database only when the data directory is empty. Mount a &lt;code&gt;/var/lib/mysql&lt;/code&gt; that already has data and every one of those variables is ignored, in silence. A stack that "will not pick up the new password" is nearly always this.&lt;/p&gt;

&lt;p&gt;And publishing 3306 is what the task asked for, not what you would otherwise do. The web container reaches the database over the project network with no published port at all. Putting 3306 on the host exposes the database to anything that can reach the host.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every permission it needed, and never called
&lt;/h2&gt;

&lt;p&gt;The AWS task was six resources in one path: an upload to a public bucket fires an S3 event, Lambda copies the object into a private bucket and writes an audit row to DynamoDB.&lt;/p&gt;

&lt;p&gt;The execution role covered &lt;code&gt;s3:GetObject&lt;/code&gt;, &lt;code&gt;s3:PutObject&lt;/code&gt; and &lt;code&gt;dynamodb:PutItem&lt;/code&gt;, scoped to exactly one bucket each and one table. Correct, least-privilege, and completely insufficient, because none of it lets S3 call the function.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Policy&lt;/th&gt;
&lt;th&gt;Attached to&lt;/th&gt;
&lt;th&gt;Answers&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Identity-based&lt;/td&gt;
&lt;td&gt;The role&lt;/td&gt;
&lt;td&gt;What is Lambda allowed to do?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resource-based&lt;/td&gt;
&lt;td&gt;The function&lt;/td&gt;
&lt;td&gt;Who is allowed to invoke Lambda?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws lambda add-permission &lt;span class="nt"&gt;--function-name&lt;/span&gt; devops-copyfunction &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--statement-id&lt;/span&gt; s3invoke &lt;span class="nt"&gt;--action&lt;/span&gt; lambda:InvokeFunction &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--principal&lt;/span&gt; s3.amazonaws.com &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--source-arn&lt;/span&gt; arn:aws:s3:::&lt;span class="nv"&gt;$PUB_BUCKET&lt;/span&gt; &lt;span class="nt"&gt;--source-account&lt;/span&gt; &lt;span class="nv"&gt;$ACCT&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Get only the identity side right and nothing errors. There are no logs, because the function never ran. This is the same shape as trust policy versus permissions policy on Day 33 and execution role versus task role on Day 38, and it keeps arriving in a new costume.&lt;/p&gt;

&lt;p&gt;The ordering is load-bearing too. &lt;code&gt;add-permission&lt;/code&gt; has to come before &lt;code&gt;put-bucket-notification-configuration&lt;/code&gt;, because S3 validates the destination when you set a notification. Without the permission you get:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Unable to validate the following destination configurations
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Which mentions neither Lambda nor permissions nor what to change. Same category as &lt;code&gt;AccessDenied&lt;/code&gt; on &lt;code&gt;PutBucketPolicy&lt;/code&gt; actually meaning &lt;code&gt;BlockPublicPolicy&lt;/code&gt;: the error names the call that failed rather than the setting that caused it. The console hides this by issuing both calls when you add a trigger through the UI.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scope the permission, or anyone can use it
&lt;/h2&gt;

&lt;p&gt;Lambda's resource-based permissions support a deliberately small set of condition keys: &lt;code&gt;aws:SourceArn&lt;/code&gt;, &lt;code&gt;aws:SourceAccount&lt;/code&gt; and &lt;code&gt;aws:PrincipalOrgID&lt;/code&gt;. Leave the first two off and the statement allows any S3 bucket, in any account, to invoke your function.&lt;/p&gt;

&lt;p&gt;That is the confused deputy problem in one line. Someone creates a bucket, points a notification at your function ARN, and your code runs on their objects with your permissions. The condition block you want to see in the response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"Condition"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"StringEquals"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"AWS:SourceAccount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"245695940513"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"ArnLike"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"AWS:SourceArn"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"arn:aws:s3:::devops-public-22976"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Three details worth stealing
&lt;/h2&gt;

&lt;p&gt;128 MB handles any file size here, and the reason is not generosity. &lt;code&gt;s3.copy_object&lt;/code&gt; is server-side: S3 copies the object internally and the bytes never pass through the Lambda execution environment. Had the code done &lt;code&gt;get_object&lt;/code&gt; then &lt;code&gt;put_object&lt;/code&gt; instead, the file would flow through Lambda memory and 128 MB would cap the file size hard. Same apparent outcome, completely different resource profile.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;put-bucket-notification-configuration&lt;/code&gt; takes the whole configuration document. There is no append. On a bucket already sending events to SQS or SNS, a straight write removes them silently, so the safe pattern on anything you did not create is read, merge, write.&lt;/p&gt;

&lt;p&gt;And a hyphen is not valid in a Python module name. The provided file is &lt;code&gt;lambda-function.py&lt;/code&gt;, the handler string is &lt;code&gt;module.function&lt;/code&gt;, and &lt;code&gt;lambda-function&lt;/code&gt; is not importable. Renaming it during the build step is a one-character fix that otherwise surfaces as an import error at invoke time, minutes after everything else looked fine.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug in the error handler
&lt;/h2&gt;

&lt;p&gt;Worth recording, because it is a realistic failure mode rather than a criticism of the lab:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;source_bucket&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;Records&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s3&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;bucket&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="bp"&gt;...&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;log_entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;SourceBucket&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;source_bucket&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;      &lt;span class="c1"&gt;# unbound if the first line threw
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If that first line raises, on a malformed or non-S3 event, &lt;code&gt;source_bucket&lt;/code&gt; was never assigned. The &lt;code&gt;except&lt;/code&gt; block then throws &lt;code&gt;NameError&lt;/code&gt; while building the error log, and the original exception is lost.&lt;/p&gt;

&lt;p&gt;A handler written to record failures, unable to record the one class of failure that happens before its variables exist. Initialising them to &lt;code&gt;None&lt;/code&gt; above the &lt;code&gt;try&lt;/code&gt; fixes it. Not needed for the task to pass, and exactly how an error handler ends up hiding the error it was written to surface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Both sides, or nothing moves
&lt;/h2&gt;

&lt;p&gt;Compose wires two containers into one project, and the relationship is implicit. Lambda needs two policies pointing in opposite directions, and neither implies the other.&lt;/p&gt;

&lt;p&gt;The tell in both cases is the same: a component that is individually correct and collectively inert. A &lt;code&gt;db&lt;/code&gt; container running perfectly that &lt;code&gt;web&lt;/code&gt; cannot name. A function with exactly the right permissions that nothing is allowed to call.&lt;/p&gt;

&lt;p&gt;So here is the Day 46 question. When you grant access to something, do you check both halves, what it can do and who can reach it, or only the half that was in the ticket?&lt;/p&gt;

&lt;p&gt;Day 46 down. Fifty-four to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>serverless</category>
    </item>
    <item>
      <title>Day 45: RUN Cannot See Your Build Context, and NAT Only Goes One Way</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Sat, 19 Sep 2026 08:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-45-run-cannot-see-your-build-context-and-nat-only-goes-one-way-4dj5</link>
      <guid>https://dev.to/ndcodes/day-45-run-cannot-see-your-build-context-and-nat-only-goes-one-way-4dj5</guid>
      <description>&lt;p&gt;Both of today's tasks are about direction. A Dockerfile instruction that cannot reach backwards to the machine you are building on. A gateway that lets traffic out and never lets it in.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Fix a Dockerfile that will not build, then give a private-subnet instance outbound internet access. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The error is about location, not spelling
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /opt/docker/
docker build &lt;span class="nt"&gt;-t&lt;/span&gt; nautilus:latest_img &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cp: cannot stat 'certs/server.crt': No such file or directory
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The file is right there. &lt;code&gt;ls&lt;/code&gt; shows &lt;code&gt;certs/server.crt&lt;/code&gt; sitting next to the Dockerfile. The line that failed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;RUN &lt;/span&gt;&lt;span class="nb"&gt;cp &lt;/span&gt;certs/server.crt /usr/local/apache2/conf/server.crt
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;cp&lt;/code&gt; is not wrong and the path is not wrong. The instruction is. &lt;code&gt;RUN&lt;/code&gt; executes a command inside the image filesystem, where &lt;code&gt;certs/&lt;/code&gt; has never existed. The file is in the build context, on the host, and &lt;code&gt;RUN&lt;/code&gt; cannot see it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; certs/server.crt /usr/local/apache2/conf/server.crt&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; certs/server.key /usr/local/apache2/conf/server.key&lt;/span&gt;
&lt;span class="k"&gt;COPY&lt;/span&gt;&lt;span class="s"&gt; html/index.html /usr/local/apache2/htdocs/&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;COPY&lt;/code&gt; takes files from the build context and puts them into the image. That is the one instruction that spans the boundary, and it is the only way anything on your machine gets into the build.&lt;/p&gt;

&lt;p&gt;What makes this a good puzzle is that the same file also has four &lt;code&gt;RUN sed&lt;/code&gt; lines that are completely fine:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight docker"&gt;&lt;code&gt;&lt;span class="k"&gt;RUN &lt;/span&gt;&lt;span class="nb"&gt;sed&lt;/span&gt; &lt;span class="nt"&gt;-i&lt;/span&gt; &lt;span class="s2"&gt;"s/Listen 80/Listen 8080/g"&lt;/span&gt; /usr/local/apache2/conf/httpd.conf
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Those edit &lt;code&gt;httpd.conf&lt;/code&gt;, which is already in the image because the base image put it there. Editing a file that exists inside the image is exactly what &lt;code&gt;RUN&lt;/code&gt; is for. Reaching for a file on the host is not. Four correct &lt;code&gt;RUN&lt;/code&gt; lines and three wrong ones, in the same file, differing only by which side of the boundary the file lives on.&lt;/p&gt;

&lt;p&gt;One detail to check before blaming the Dockerfile. The trailing &lt;code&gt;.&lt;/code&gt; on &lt;code&gt;docker build&lt;/code&gt; is the build context, and &lt;code&gt;COPY&lt;/code&gt; paths resolve against it. Build from the wrong directory and &lt;code&gt;COPY certs/server.crt&lt;/code&gt; fails identically, for the opposite reason: the instruction is right and the context does not contain the file.&lt;/p&gt;

&lt;h2&gt;
  
  
  The NAT gateway goes in the subnet it is not serving
&lt;/h2&gt;

&lt;p&gt;This one reads backwards until it clicks. A NAT gateway serving a private subnet is placed in the public subnet, and AWS states it as a requirement: you create a public NAT gateway in a public subnet and must associate an Elastic IP with it at creation.&lt;/p&gt;

&lt;p&gt;The reason is that a NAT gateway is itself a client of the internet. It takes traffic from private instances, rewrites the source address to its own Elastic IP, and forwards it out. To forward anything it needs its own route to an internet gateway, and that route lives in the public subnet's table.&lt;/p&gt;

&lt;p&gt;Put it in the private subnet and it creates without complaint, reaches &lt;code&gt;available&lt;/code&gt;, and drops every packet it forwards. Nothing in &lt;code&gt;describe-nat-gateways&lt;/code&gt; suggests a problem.&lt;/p&gt;

&lt;p&gt;The finished path is two tables and two hops:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;EC2 (private subnet)
  -&amp;gt; priv-rt: 0.0.0.0/0 -&amp;gt; nat-gateway
    -&amp;gt; NAT gateway (public subnet, holds the Elastic IP)
      -&amp;gt; pub-rt: 0.0.0.0/0 -&amp;gt; internet gateway
        -&amp;gt; internet
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The private subnet never references the IGW directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  One way only
&lt;/h2&gt;

&lt;p&gt;"Internet access" is doing a lot of work in the task description, so it is worth being precise. AWS says instances behind a NAT gateway cannot receive unsolicited inbound connections from the internet.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Internet gateway&lt;/th&gt;
&lt;th&gt;NAT gateway&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Outbound from instance&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inbound to instance&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Instance needs a public IP&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Costs money at rest&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The instance still has no public IP and never will. It can install packages, call AWS APIs and upload to S3, and nothing on the internet can open a connection to it. That asymmetry is the entire reason to use a NAT gateway rather than just moving the instance to a public subnet.&lt;/p&gt;

&lt;p&gt;The Elastic IP belongs to the gateway, not the instance, so every private instance behind it shares one source address. Which is also what makes it useful when a third party wants an IP to allowlist.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two things the task was right to insist on
&lt;/h2&gt;

&lt;p&gt;It said explicitly not to touch the VPC's Main route table, and that is not fussiness. The Main table is the fallback for any subnet with no explicit association, so a &lt;code&gt;0.0.0.0/0&lt;/code&gt; route there grants internet access to every unassociated subnet in the VPC, including ones created months later by someone who assumed a new subnet would be isolated. A dedicated table with an explicit association makes the intent readable and keeps everything else isolated by default.&lt;/p&gt;

&lt;p&gt;And AZ placement is not cosmetic. AWS documents each NAT gateway as created in a specific availability zone and made redundant within that zone. Which means if resources across several AZs share one gateway and its zone goes down, the resources in the healthy zones lose internet access too. The recommendation is one per AZ with routing to match. Here one private subnet in one AZ meant one gateway in the same AZ, which is both correct and cheapest.&lt;/p&gt;

&lt;p&gt;While we are on cost: a NAT gateway bills hourly whether anything flows through it or not, plus a per-gigabyte processing charge. For this workload, an instance in a private subnet talking only to S3, a Gateway VPC Endpoint would have done the job for nothing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 create-vpc-endpoint &lt;span class="nt"&gt;--vpc-id&lt;/span&gt; &lt;span class="nv"&gt;$VPC&lt;/span&gt; &lt;span class="nt"&gt;--service-name&lt;/span&gt; com.amazonaws.us-east-1.s3 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--route-table-ids&lt;/span&gt; &lt;span class="nv"&gt;$PRIV_RT&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Free, keeps the traffic on the AWS network, and covers S3 and DynamoDB only. The task asked for NAT, so NAT is what got built, but the cheaper answer exists.&lt;/p&gt;

&lt;h2&gt;
  
  
  Checks that come before the wait
&lt;/h2&gt;

&lt;p&gt;Two habits carried over from Day 40, both of which turn an ambiguous outcome into a clear one.&lt;/p&gt;

&lt;p&gt;Check the route state before waiting on the result. &lt;code&gt;active&lt;/code&gt; rather than &lt;code&gt;blackhole&lt;/code&gt; proves the target resolves, so an empty bucket three minutes later means the cron job has not fired yet rather than that the path is broken.&lt;/p&gt;

&lt;p&gt;And establish the negative first. The bucket was confirmed empty before anything was built, which is what turns the later listing into evidence. A leftover file from a previous attempt would have looked exactly like success.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which side of the line
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;RUN&lt;/code&gt; cannot reach the host. A NAT gateway cannot be reached from the internet. In both cases the thing that feels like a limitation is the boundary doing its job, and in both cases the failure mode is a command that looks reasonable and fails somewhere you were not looking.&lt;/p&gt;

&lt;p&gt;So here is the Day 45 question. In the last thing you built, do you know which direction each connection can be opened from, or only that it works?&lt;/p&gt;

&lt;p&gt;Day 45 down. Fifty-five to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>networking</category>
    </item>
    <item>
      <title>Day 44: Compose Stopped Needing a Version, and the CLI Gives You No Grace Period</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Fri, 18 Sep 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-44-compose-stopped-needing-a-version-and-the-cli-gives-you-no-grace-period-4ip7</link>
      <guid>https://dev.to/ndcodes/day-44-compose-stopped-needing-a-version-and-the-cli-gives-you-no-grace-period-4ip7</guid>
      <description>&lt;p&gt;Two defaults today, both of them different from what a reasonable person would assume. Compose no longer wants the &lt;code&gt;version&lt;/code&gt; key that every tutorial still opens with. And an Auto Scaling group created from the CLI has a health check grace period of zero, where the console gives you 300 seconds.&lt;/p&gt;

&lt;p&gt;One is a warning. The other is an infinite loop.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Write a Compose file serving httpd from a host directory, then rebuild Day 36's load-balanced instance as an Auto Scaling group. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The version key is obsolete
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;services&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;web&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;httpd:latest&lt;/span&gt;
    &lt;span class="na"&gt;container_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;httpd&lt;/span&gt;
    &lt;span class="na"&gt;ports&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5000:80"&lt;/span&gt;
    &lt;span class="na"&gt;volumes&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;/opt/sysops:/usr/local/apache2/htdocs&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No &lt;code&gt;version: "3.9"&lt;/code&gt; at the top, and that is deliberate. Compose's documentation says the top-level &lt;code&gt;version&lt;/code&gt; property exists only for backward compatibility, that it is informative, and that using it produces a warning telling you it is obsolete. Compose validates against the most recent schema regardless of what you write there.&lt;/p&gt;

&lt;p&gt;So the line every older tutorial opens with now buys nothing and costs you a warning on every run.&lt;/p&gt;

&lt;p&gt;Three other things in that short file are worth more than they look.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;ports&lt;/code&gt; and &lt;code&gt;volumes&lt;/code&gt; both read host first, container second, matching &lt;code&gt;-p&lt;/code&gt; and &lt;code&gt;-v&lt;/code&gt; on the command line. Quote the port mapping, because YAML will read an unquoted &lt;code&gt;5000:80&lt;/code&gt; as a sexagesimal number in some parsers and hand you a value you did not write.&lt;/p&gt;

&lt;p&gt;The document root is &lt;code&gt;/usr/local/apache2/htdocs&lt;/code&gt;, not &lt;code&gt;/var/www/html&lt;/code&gt;. That second path belongs to the Debian and Ubuntu Apache packages; the official &lt;code&gt;httpd&lt;/code&gt; image lays itself out differently. Mount over the wrong one and you get a running container cheerfully serving the image's default page.&lt;/p&gt;

&lt;p&gt;And &lt;code&gt;container_name&lt;/code&gt; fixes the name rather than letting Compose generate &lt;code&gt;&amp;lt;project&amp;gt;-&amp;lt;service&amp;gt;-1&lt;/code&gt;. It makes &lt;code&gt;docker exec&lt;/code&gt; predictable and makes scaling that service past one replica impossible, since two containers cannot share a name.&lt;/p&gt;

&lt;h2&gt;
  
  
  The default that would have cost me the task
&lt;/h2&gt;

&lt;p&gt;The AWS half was Day 36 again with the hand-launched instance replaced by an Auto Scaling group, a launch template and a target tracking policy. Two things in it are genuinely worth carrying.&lt;/p&gt;

&lt;p&gt;First, launch templates do not encode user data for you.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# run-instances: CLI v2 base64-encodes this automatically&lt;/span&gt;
&lt;span class="nt"&gt;--user-data&lt;/span&gt; file:///root/ud.sh

&lt;span class="c"&gt;# create-launch-template: you do it yourself&lt;/span&gt;
&lt;span class="nv"&gt;UD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;base64&lt;/span&gt; &lt;span class="nt"&gt;-w&lt;/span&gt; 0 /root/ud.sh&lt;span class="si"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;run-instances&lt;/code&gt; has a dedicated parameter with special handling. &lt;code&gt;create-launch-template&lt;/code&gt; takes one opaque blob where &lt;code&gt;UserData&lt;/code&gt; is a string field the API expects already encoded. Pass raw text and the call succeeds. The template is created, the ASG launches an instance, cloud-init cannot make sense of the payload and moves on. You get a running instance with no nginx, no error in any API response, and a target group that never goes healthy. The only trace is in a log file on an instance you have not set up SSH for.&lt;/p&gt;

&lt;p&gt;Second, and this is the one I would not have found without reading the documentation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--health-check-grace-period&lt;/span&gt; 300
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The grace period is the minimum time a new instance stays in service before being terminated for failing a health check. AWS documents the default as 300 seconds in the console and 0 seconds via the CLI or SDK, where 0 turns it off entirely.&lt;/p&gt;

&lt;p&gt;That matters because the task also required &lt;code&gt;--health-check-type ELB&lt;/code&gt;, which makes the target group's HTTP check drive instance replacement rather than the default EC2 status check. The default only asks whether the VM is running, so nginx could crash completely and the ASG would see a perfectly healthy instance.&lt;/p&gt;

&lt;p&gt;Put those together with no grace period and the sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Instance launches&lt;/li&gt;
&lt;li&gt;Health check fails, because nginx is still installing&lt;/li&gt;
&lt;li&gt;ASG terminates it as unhealthy&lt;/li&gt;
&lt;li&gt;ASG launches a replacement to satisfy desired capacity&lt;/li&gt;
&lt;li&gt;Back to step 2, forever&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;An infinite launch-and-terminate loop that burns money and never converges, and it shows up in the ASG activity history rather than anywhere you would think to look. Preventing exactly that is what AWS says the grace period is for.&lt;/p&gt;

&lt;p&gt;Worth noting the warning runs the other way too. Set it too high and a genuinely broken instance stays in service longer, blunting the health checks you turned on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three smaller things
&lt;/h2&gt;

&lt;p&gt;The ASG registers targets, you do not. Naming &lt;code&gt;--target-group-arns&lt;/code&gt; on the group means it registers every instance it launches and deregisters every one it terminates. Running &lt;code&gt;register-targets&lt;/code&gt; by hand against an ASG-managed group is actively wrong, because the next scaling action silently undoes it.&lt;/p&gt;

&lt;p&gt;Early &lt;code&gt;unhealthy&lt;/code&gt; is normal and the reason field tells you which kind it is. &lt;code&gt;Elb.RegistrationInProgress&lt;/code&gt; and &lt;code&gt;Elb.InitialHealthChecking&lt;/code&gt; mean wait. &lt;code&gt;Target.Timeout&lt;/code&gt; means no TCP connection, so a security group or nothing listening. &lt;code&gt;Target.ResponseCodeMismatch&lt;/code&gt; means it answered with the wrong status. Only the last two are unambiguously broken.&lt;/p&gt;

&lt;p&gt;And escape &lt;code&gt;\$Latest&lt;/code&gt;. It is a literal AWS keyword, and unescaped inside double quotes bash expands it as a shell variable, finds nothing, and sends &lt;code&gt;Version=&lt;/code&gt; to an API that rejects it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reading the default rather than assuming it
&lt;/h2&gt;

&lt;p&gt;The Compose version key is a default that used to be required and now warns. The grace period is a default that differs depending on whether you used the console or the CLI, in a direction that turns a working configuration into a loop.&lt;/p&gt;

&lt;p&gt;Neither is discoverable by doing the obvious thing and watching it work. The Compose file runs fine with the version key, just noisily. The ASG would have failed in a way that looks like an application problem.&lt;/p&gt;

&lt;p&gt;So here is the Day 44 question. For the last resource you created from a script rather than a console, do you know which of its defaults differ between the two?&lt;/p&gt;

&lt;p&gt;Day 44 down. Fifty-six to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 43: A Port Is Closed Until You Publish It, and an AZ Name Is Not an AZ</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Wed, 16 Sep 2026 17:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-43-a-port-is-closed-until-you-publish-it-and-an-az-name-is-not-an-az-20h8</link>
      <guid>https://dev.to/ndcodes/day-43-a-port-is-closed-until-you-publish-it-and-an-az-name-is-not-an-az-20h8</guid>
      <description>&lt;p&gt;Today both tasks turned on reading a label correctly. A port number on the host is not the port inside the container, and they do not have to match. An availability zone name is not the availability zone, and in someone else's account it points somewhere else entirely.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Publish a container port on a different host port, then provision an EKS control plane. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  Publishing, and the direction of the colon
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker pull nginx:stable
docker run &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="nt"&gt;--name&lt;/span&gt; demo &lt;span class="nt"&gt;-p&lt;/span&gt; 6000:80 nginx:stable
curl http://localhost:6000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Host first, container second. &lt;code&gt;-p 6000:80&lt;/code&gt; listens on 6000 on the host and forwards to 80 inside. Reverse it and you get a host listening on 80, forwarding to a container port nothing is bound to, and a connection that opens and returns nothing.&lt;/p&gt;

&lt;p&gt;The two numbers are independent, which is the part worth internalising. nginx is still on 80 inside the container and has no idea the host calls it 6000. Nothing in the image changes to publish it elsewhere.&lt;/p&gt;

&lt;p&gt;And without &lt;code&gt;-p&lt;/code&gt; the container is not broken, it is just not reachable from outside. Docker puts it plainly: ports on bridge networks are accessible from the Docker host and from other containers on the same network, and are not accessible from outside the host or, by default, from containers on other networks.&lt;/p&gt;

&lt;p&gt;That splits cleanly against yesterday's task. The network decides which containers can see each other. Publishing decides whether the host's port space does. Two different questions, and &lt;code&gt;EXPOSE&lt;/code&gt; in a Dockerfile answers neither, since it only documents which port the application listens on.&lt;/p&gt;

&lt;p&gt;One habit worth forming early: &lt;code&gt;-p 127.0.0.1:6000:80&lt;/code&gt; binds to loopback only. Plain &lt;code&gt;-p 6000:80&lt;/code&gt; binds on every interface, which on a public host means the internet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The error names a zone that does not generalise
&lt;/h2&gt;

&lt;p&gt;The EKS task was a control plane to a specific configuration. Named cluster, latest stable Kubernetes, default VPC, three availability zones, Auto Mode off, private endpoint.&lt;/p&gt;

&lt;p&gt;Passing all six subnets in the default VPC produced this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cannot create cluster because EKS does not support creating control plane instances in us-east-1e
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Clear enough, and I nearly wrote it down as the rule. It is not the rule.&lt;/p&gt;

&lt;p&gt;AWS documents the restriction against availability zone &lt;strong&gt;IDs&lt;/strong&gt;, not zone names. Subnets cannot sit in &lt;code&gt;use1-az3&lt;/code&gt; in us-east-1, &lt;code&gt;usw1-az2&lt;/code&gt; in us-west-1, or &lt;code&gt;cac1-az3&lt;/code&gt; in ca-central-1. Zone names are mapped to zone IDs independently per account, so &lt;code&gt;us-east-1e&lt;/code&gt; is this account's name for &lt;code&gt;use1-az3&lt;/code&gt;, and in your account &lt;code&gt;use1-az3&lt;/code&gt; may well be called something else.&lt;/p&gt;

&lt;p&gt;Which means "avoid us-east-1e" is a note that works in exactly one account and silently avoids the wrong zone everywhere else. The portable version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws ec2 describe-availability-zones &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'AvailabilityZones[].{Name:ZoneName,Id:ZoneId}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is the kind of thing that only shows up when you go and read the documentation for something you already appeared to understand. The error told me the truth about my account and nothing about the rule.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shorthand cannot hold a list inside a struct
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--resources-vpc-config&lt;/span&gt; &lt;span class="nv"&gt;subnetIds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;subnet-a,subnet-b,subnet-c,endpointPublicAccess&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nb"&gt;false&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;AWS CLI shorthand separates struct fields with commas. It also separates list items with commas. So that string is genuinely ambiguous, and the parser cannot tell where &lt;code&gt;subnetIds&lt;/code&gt; ends and the next field begins.&lt;/p&gt;

&lt;p&gt;JSON removes the question:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--resources-vpc-config&lt;/span&gt; &lt;span class="s2"&gt;"{&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;subnetIds&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:[&lt;/span&gt;&lt;span class="nv"&gt;$SUBNET_JSON&lt;/span&gt;&lt;span class="s2"&gt;],&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;endpointPublicAccess&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:false,&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;endpointPrivateAccess&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;:true}"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is now the third list convention in this challenge. &lt;code&gt;elbv2 create-load-balancer --subnets&lt;/code&gt; wants spaces. &lt;code&gt;ecs create-service&lt;/code&gt; wants commas inside brackets. &lt;code&gt;eks create-cluster&lt;/code&gt; wants a JSON array. There is no rule, only the habit of checking, and the working heuristic is to reach for JSON the moment a struct contains a list.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two smaller things, and one honest caveat
&lt;/h2&gt;

&lt;p&gt;Resolve the version rather than hardcoding it. &lt;code&gt;aws eks describe-cluster-versions&lt;/code&gt; returns what is currently valid, and Kubernetes ships three minor releases a year, so a hardcoded version is wrong within months. Same habit as SSM parameters for AMIs and orderable instance options for RDS engine versions: ask AWS what is valid now.&lt;/p&gt;

&lt;p&gt;The service principal for an EKS cluster role is &lt;code&gt;eks.amazonaws.com&lt;/code&gt;. That is the fourth different principal in this run, after &lt;code&gt;lambda.amazonaws.com&lt;/code&gt;, &lt;code&gt;ec2.amazonaws.com&lt;/code&gt; and &lt;code&gt;ecs-tasks.amazonaws.com&lt;/code&gt;. Mostly it is &lt;code&gt;&amp;lt;service&amp;gt;.amazonaws.com&lt;/code&gt;, with ECS as the exception. A wrong one gives you a role that looks right in the console and can never be assumed.&lt;/p&gt;

&lt;p&gt;And the caveat, because the task says "ready for workloads" and that sentence is doing some work. The cluster is ACTIVE and has no compute attached. No managed node group, no Fargate profile, nothing. Schedule a pod today and it sits &lt;code&gt;Pending&lt;/code&gt; forever with no nodes available. That is the correct scope, since the task specified nothing about nodes and inventing a node group would have been inventing requirements. But ACTIVE is not schedulable, and the node role that comes next needs three managed policies and is a different role from the cluster role.&lt;/p&gt;

&lt;h2&gt;
  
  
  Labels are not the things they label
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;6000&lt;/code&gt; on the host and &lt;code&gt;80&lt;/code&gt; in the container are the same service under two names. &lt;code&gt;us-east-1e&lt;/code&gt; and &lt;code&gt;use1-az3&lt;/code&gt; are one zone under two names, one of which travels and one of which does not.&lt;/p&gt;

&lt;p&gt;The port case is harmless because you chose both numbers. The zone case is not, because the name was chosen for you, per account, and the error message hands you the local one as though it were universal.&lt;/p&gt;

&lt;p&gt;So here is the Day 43 question. The last error message you turned into a rule, was it telling you about the system, or about your particular instance of it?&lt;/p&gt;

&lt;p&gt;Day 43 down. Fifty-seven to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>containers</category>
    </item>
    <item>
      <title>Day 42: A Custom Network Gives You DNS, and DynamoDB Only Declares Its Key</title>
      <dc:creator>Nnamdi Felix Ibe</dc:creator>
      <pubDate>Mon, 14 Sep 2026 21:00:00 +0000</pubDate>
      <link>https://dev.to/ndcodes/day-42-a-custom-network-gives-you-dns-and-dynamodb-only-declares-its-key-232</link>
      <guid>https://dev.to/ndcodes/day-42-a-custom-network-gives-you-dns-and-dynamodb-only-declares-its-key-232</guid>
      <description>&lt;p&gt;Both of today's tasks declare far less than you would expect, and both have a rule waiting in the part you did not declare. A Docker network is three flags and changes how containers find each other. A DynamoDB table declares one attribute out of three, and the one word the task chose for another is reserved.&lt;/p&gt;

&lt;p&gt;One Docker task, one AWS task. Create a user-defined network, then build a DynamoDB table and prove two items have the right status. The tasks come from the KodeKloud Engineer platform.&lt;/p&gt;

&lt;h2&gt;
  
  
  The reason to create a network is not isolation
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;docker network &lt;span class="nb"&gt;ls
&lt;/span&gt;docker network create &amp;lt;name&amp;gt; &lt;span class="nt"&gt;--driver&lt;/span&gt; &amp;lt;driver&amp;gt; &lt;span class="nt"&gt;--subnet&lt;/span&gt; &amp;lt;CIDR&amp;gt; &lt;span class="nt"&gt;--ip-range&lt;/span&gt; &amp;lt;CIDR inside it&amp;gt;
docker network inspect &amp;lt;name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Six drivers ship with Docker, and they are not variations on a theme. &lt;code&gt;bridge&lt;/code&gt; is the default, &lt;code&gt;host&lt;/code&gt; removes network isolation from the host entirely, &lt;code&gt;none&lt;/code&gt; isolates completely, &lt;code&gt;overlay&lt;/code&gt; joins multiple daemons for Swarm, and &lt;code&gt;ipvlan&lt;/code&gt; and &lt;code&gt;macvlan&lt;/code&gt; put containers onto the physical network, the latter making them appear as devices on the host's own network.&lt;/p&gt;

&lt;p&gt;But the thing that makes a user-defined network worth creating is smaller and more useful than any of that. Containers on the default bridge cannot refer to each other by name. Containers on a network you created use Docker's embedded DNS server and can reach each other by container name.&lt;/p&gt;

&lt;p&gt;That is the whole feature. Every multi-container stack that connects to &lt;code&gt;db&lt;/code&gt; rather than &lt;code&gt;172.18.0.3&lt;/code&gt; is relying on it, including the Compose files coming up later this week, which get a project network for free and never mention DNS anywhere.&lt;/p&gt;

&lt;p&gt;Two smaller things worth knowing. &lt;code&gt;--subnet&lt;/code&gt; is the CIDR the network occupies, and &lt;code&gt;--ip-range&lt;/code&gt; allocates container addresses from a sub-range inside it, which is how you keep the rest of the block for addresses you assign yourself. And &lt;code&gt;macvlan&lt;/code&gt; and &lt;code&gt;ipvlan&lt;/code&gt; are not drop-in swaps for a bridge: they attach to a physical interface and need the host's network to cooperate.&lt;/p&gt;

&lt;p&gt;Finish with &lt;code&gt;inspect&lt;/code&gt;, not &lt;code&gt;ls&lt;/code&gt;. &lt;code&gt;ls&lt;/code&gt; proves a name exists. &lt;code&gt;inspect&lt;/code&gt; shows the driver, the addressing and what is attached, which is what the task actually specified.&lt;/p&gt;

&lt;h2&gt;
  
  
  You declare the keys and nothing else
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws dynamodb create-table &lt;span class="nt"&gt;--table-name&lt;/span&gt; datacenter-tasks &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--attribute-definitions&lt;/span&gt; &lt;span class="nv"&gt;AttributeName&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;taskId,AttributeType&lt;span class="o"&gt;=&lt;/span&gt;S &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--key-schema&lt;/span&gt; &lt;span class="nv"&gt;AttributeName&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;taskId,KeyType&lt;span class="o"&gt;=&lt;/span&gt;HASH &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--billing-mode&lt;/span&gt; PAY_PER_REQUEST
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The items carry &lt;code&gt;taskId&lt;/code&gt;, &lt;code&gt;description&lt;/code&gt; and &lt;code&gt;status&lt;/code&gt;. Only one of those appears in the command, and that is correct. AWS defines &lt;code&gt;AttributeDefinitions&lt;/code&gt; as an array of attributes that describe the key schema for the table and indexes, and nothing beyond that. Everything else comes into existence when an item carries it.&lt;/p&gt;

&lt;p&gt;This reliably catches people arriving from SQL, who read &lt;code&gt;--attribute-definitions&lt;/code&gt; as a column list and try to declare all three. Adding &lt;code&gt;description&lt;/code&gt; there without using it in a key is an error, because DynamoDB has no use for the definition.&lt;/p&gt;

&lt;p&gt;One habit before the first write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws dynamodb &lt;span class="nb"&gt;wait &lt;/span&gt;table-exists &lt;span class="nt"&gt;--table-name&lt;/span&gt; datacenter-tasks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;create-table&lt;/code&gt; returns &lt;code&gt;CREATING&lt;/code&gt;, and writing to a table in that state fails with &lt;code&gt;ResourceNotFoundException&lt;/code&gt;, which reads exactly like a typo in the table name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Every value carries its type, and numbers are strings
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"taskId"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"S"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"1"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"S"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Learn DynamoDB"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="nl"&gt;"status"&lt;/span&gt;&lt;span class="p"&gt;:{&lt;/span&gt;&lt;span class="nl"&gt;"S"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"completed"&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;S&lt;/code&gt; string, &lt;code&gt;N&lt;/code&gt; number, &lt;code&gt;B&lt;/code&gt; binary, plus &lt;code&gt;BOOL&lt;/code&gt;, &lt;code&gt;L&lt;/code&gt;, &lt;code&gt;M&lt;/code&gt;, &lt;code&gt;NULL&lt;/code&gt; and the set types. The one that surprises people is that &lt;code&gt;N&lt;/code&gt; values are written as JSON strings: &lt;code&gt;{"N": "42"}&lt;/code&gt;, never &lt;code&gt;{"N": 42}&lt;/code&gt;. AWS gives the reason directly, which is that numbers are sent across the network as strings to maximise compatibility across languages and libraries. They are still treated as numbers once they arrive.&lt;/p&gt;

&lt;p&gt;Note &lt;code&gt;taskId&lt;/code&gt; is &lt;code&gt;{"S":"1"}&lt;/code&gt;, the character, not the integer. The key schema declared it &lt;code&gt;S&lt;/code&gt;, so &lt;code&gt;{"N":"1"}&lt;/code&gt; would be rejected outright as a key type mismatch.&lt;/p&gt;

&lt;h2&gt;
  
  
  The word the task chose is reserved
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws dynamodb scan &lt;span class="nt"&gt;--table-name&lt;/span&gt; datacenter-tasks &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter-expression&lt;/span&gt; &lt;span class="s2"&gt;"status = :v"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--expression-attribute-values&lt;/span&gt; &lt;span class="s1"&gt;'{":v":{"S":"completed"}}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Invalid FilterExpression: Attribute name is a reserved keyword; reserved keyword: status
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;DynamoDB reserves several hundred words in its expression language and they are ordinary English. &lt;code&gt;status&lt;/code&gt;, &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;size&lt;/code&gt;, &lt;code&gt;count&lt;/code&gt;, &lt;code&gt;timestamp&lt;/code&gt;, &lt;code&gt;year&lt;/code&gt;, &lt;code&gt;data&lt;/code&gt;, &lt;code&gt;value&lt;/code&gt;, &lt;code&gt;source&lt;/code&gt;, &lt;code&gt;region&lt;/code&gt;, &lt;code&gt;owner&lt;/code&gt;, &lt;code&gt;hash&lt;/code&gt; and &lt;code&gt;key&lt;/code&gt; are all on the published list. Almost any attribute name you would reach for first is on it somewhere.&lt;/p&gt;

&lt;p&gt;The fix is an expression attribute name, which is a placeholder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws dynamodb scan &lt;span class="nt"&gt;--table-name&lt;/span&gt; datacenter-tasks &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--filter-expression&lt;/span&gt; &lt;span class="s2"&gt;"#s = :v"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--expression-attribute-names&lt;/span&gt; &lt;span class="s1"&gt;'{"#s":"status"}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--expression-attribute-values&lt;/span&gt; &lt;span class="s1"&gt;'{":v":{"S":"completed"}}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s1"&gt;'Items[].taskId.S'&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two substitution mechanisms, and conflating them is the usual next mistake. &lt;code&gt;#name&lt;/code&gt; stands in for an attribute name and its value is a bare string. &lt;code&gt;:value&lt;/code&gt; stands in for a value and its value is a typed descriptor. &lt;code&gt;#&lt;/code&gt; is what to look at, &lt;code&gt;:&lt;/code&gt; is what to compare against.&lt;/p&gt;

&lt;p&gt;Worth noticing that &lt;code&gt;put-item&lt;/code&gt; needed none of this. Reserved words only matter inside expressions, so an attribute literally called &lt;code&gt;status&lt;/code&gt; stores perfectly well and only becomes awkward the moment you query on it. Which points to the real lesson, upstream of all of it: do not name an attribute &lt;code&gt;status&lt;/code&gt;. &lt;code&gt;taskStatus&lt;/code&gt; costs nothing and deletes the problem from every future query.&lt;/p&gt;

&lt;p&gt;One last thing about that verification, because it is doing something the &lt;code&gt;get-item&lt;/code&gt; before it did not. &lt;code&gt;get-item&lt;/code&gt; shows the item and lets me read the status off the screen. The filtered scan makes DynamoDB assert the status matches and hand back the ID. Same answer, different authority. That said, &lt;code&gt;FilterExpression&lt;/code&gt; is applied after the read, so a scan reads the whole table and pays for all of it before discarding what does not match. Fine for two items, a bad habit at a million.&lt;/p&gt;

&lt;h2&gt;
  
  
  The undeclared part still has rules
&lt;/h2&gt;

&lt;p&gt;The Docker network and the DynamoDB table are both mostly implicit. You do not wire containers to each other, you put them on a network and naming resolves. You do not declare a schema, you write items, and the attributes appear.&lt;/p&gt;

&lt;p&gt;In both cases, the implicit part is not lawless. Name resolution only works on a network you created. Undeclared attributes are fine until one of them collides with a reserved word you had no reason to know about.&lt;/p&gt;

&lt;p&gt;So here is the Day 42 question. In the system you work on, what behaviour are you relying on that nothing in your configuration actually states?&lt;/p&gt;

&lt;p&gt;Day 42 down. Fifty-eight to go.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>docker</category>
      <category>aws</category>
      <category>database</category>
    </item>
  </channel>
</rss>
