Day 60 of DevOps, Day 10 of Azure. Both tasks today had an object in the middle that the task never names. A pod does not mount a PersistentVolume; it mounts a claim, and the claim finds the volume. A VM does not hold a public IP; an IP configuration on its network interface does.
One Kubernetes task, one Azure task. Give a web server pod storage that outlives it, then attach an existing public IP to an existing VM. The tasks come from the KodeKloud Engineer platform.
Task 1: Persistent storage for a pod
The task: a 3Gi PersistentVolume called pv-xfusion on a host path, a claim called pvc-xfusion, a pod called pod-xfusion running httpd:latest with the claim mounted, and a NodePort Service called web-xfusion on port 30008.
Step 1: Create the PersistentVolume
apiVersion: v1
kind: PersistentVolume
metadata:
name: pv-xfusion
spec:
capacity:
storage: 3Gi
accessModes:
- ReadWriteOnce
storageClassName: manual
hostPath:
path: /mnt/security
kubectl apply -f k3s-pv.yaml
kubectl get pv
The volume shows STATUS Available. It exists, and nothing is using it yet.
Step 2: Claim it
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
name: pvc-xfusion
spec:
accessModes:
- ReadWriteOnce
resources:
requests:
storage: 3Gi
storageClassName: manual
kubectl apply -f k3s-pvc.yaml
kubectl get pvc
Now both show Bound, and the claim's VOLUME column reads pv-xfusion. I never told the claim which volume to use. Kubernetes matched them on three things: the same storageClassName, a compatible access mode, and enough capacity.
Step 3: Mount the claim in a pod
apiVersion: v1
kind: Pod
metadata:
name: pod-xfusion
labels:
app: xfusion
spec:
containers:
- name: container-xfusion
image: httpd:latest
volumeMounts:
- name: xfusion-volume
mountPath: /usr/share/nginx/html/data
resources:
limits:
memory: "128Mi"
cpu: "500m"
requests:
memory: "64Mi"
cpu: "250m"
volumes:
- name: xfusion-volume
persistentVolumeClaim:
claimName: pvc-xfusion
kubectl apply -f k3s-pod.yaml
The pod names the claim, not the volume. That is the whole design: the pod says what it needs, and what actually provides the storage can change without the pod's spec changing.
Step 4: Expose the pod
apiVersion: v1
kind: Service
metadata:
name: web-xfusion
labels:
app: xfusion
spec:
type: NodePort
ports:
- port: 80
targetPort: 80
nodePort: 30008
selector:
app: xfusion
kubectl apply -f k3s-svc.yaml
kubectl get pods
kubectl get svc web-xfusion
The pod is Running, and the Service shows 80:30008/TCP.
Step 5: Prove the data outlives the pod
kubectl exec pod-xfusion -- sh -c 'echo "Hello from persistent storage" > /usr/share/nginx/html/data/test.html'
kubectl delete pod pod-xfusion
kubectl apply -f k3s-pod.yaml
kubectl exec pod-xfusion -- cat /usr/share/nginx/html/data/test.html
Hello from persistent storage
It is a new pod, and the file is still there, because it was never in the pod. It sits in /mnt/security on the node, held by a volume and a claim that both outlived the pod.
A path from the wrong image
Look at the mount path again: /usr/share/nginx/html/data. That sits under /usr/share/nginx/html, the nginx image's document root. This pod runs httpd, and the official httpd image serves from /usr/local/apache2/htdocs/.
So the storage worked exactly as the test shows, and Apache was never serving it. A request to the Service on port 30008 would get the image's default page, and nothing from the volume. The persistence test passed because it only tested the volume.
If the volume is meant to hold the site, the mount for this image is:
volumeMounts:
- name: xfusion-volume
mountPath: /usr/local/apache2/htdocs
That has its own consequence, the one from Day 55: a mount hides what the image had at that path. Apache's default page would be covered, and the site would serve whatever is on the volume.
What the docs recommend
Keep hostPath for single-node testing. The Kubernetes docs list it as "for single node testing only; WILL NOT WORK in a multi-node cluster", and say of the volume type in general: "If you can avoid using a hostPath volume, you should." On a cluster with several nodes, my recreated pod could land on a different node, where /mnt/security is a different directory.
Let the cluster provision storage. In production nobody writes the PersistentVolume by hand. Dynamic provisioning "automatically provisions storage when users create PersistentVolumeClaim objects", through a StorageClass. In my manifests manual is only a matching name; no StorageClass object called manual has to exist.
Know the fourth access mode. My notes listed three. The docs have four, the newest being ReadWriteOncePod, where "the volume can be mounted as read-write by a single Pod". ReadWriteOnce is looser than it sounds: it is per node, so several pods on the same node can share the volume.
Decide what happens to the data afterwards. For a PersistentVolume created by hand, the reclaim policy defaults to Retain. Delete the claim, and the volume becomes Released, which is not the same as available: "it is not yet available for another claim because the previous claimant's data remains on the volume".
Delete in order. A claim that a pod is still using is not removed straight away; its removal "is postponed until the PVC is no longer actively used by any Pods". Pod first, then claim, then volume.
One correction to my notes: their cloud volume examples were out of date. The in-tree awsElasticBlockStore and azureDisk volume types were removed in Kubernetes v1.27, and the docs' current AWS example uses the CSI provisioner ebs.csi.aws.com.
Task 2: Attach a public IP to a VM
The task: attach the existing public IP devops-pip to the existing VM devops-vm-pip.
Step 1: Derive the NIC and IP configuration names
A public IP does not attach to a VM, or even to a NIC. Microsoft: "you associate the public IP address to an IP configuration of a network interface attached to a VM."
VM -> NIC -> IP configuration -> public IP
So the command needs two names the task never mentions, and I read both from the API:
NIC_ID=$(az vm show -g $RG -n devops-vm-pip --query "networkProfile.networkInterfaces[0].id" -o tsv)
NIC=$(basename $NIC_ID)
IPCFG=$(az network nic show --ids $NIC_ID --query "ipConfigurations[0].name" -o tsv)
basename works because every Azure resource ID ends in /<name>. [0] was safe because I had confirmed the VM had one NIC.
The IP configuration turned out to be ipconfigdevops-vm-pip, not the ipconfig1 that most examples use. The CLI's source explains the difference: az vm create names it ipconfig plus the VM name, while az network nic create uses ipconfig1. Microsoft's own pages are not consistent either. The page on associating a public IP uses ipconfig1 in its example, and the page on dissociating one uses ipconfigmyVM. Guessing is how this step fails.
Step 2: Attach the public IP
az network nic ip-config update -g $RG --nic-name $NIC -n $IPCFG --public-ip-address devops-pip
The VM stayed VM running before, during and after. Yesterday's NIC needed the VM deallocated. A public IP is a setting on the NIC's IP configuration, so the VM itself is not modified.
Step 3: Verify from both ends
az vm show -g $RG -n devops-vm-pip -d --query "{PubIPs:publicIps}"
az network public-ip show -g $RG -n devops-pip --query "{IP:ipAddress,AttachedTo:ipConfiguration.id}"
"PubIPs": "172.178.12.220"
"IP": "172.178.12.220"
"AttachedTo": ".../networkInterfaces/devops-vm-pipVMNic/ipConfigurations/ipconfigdevops-vm-pip"
The VM reports the address, and the public IP reports its binding, down to the IP configuration. These are two views of one relationship, and they agree.
Before the attach, that first query returned "", an empty string, not null. publicIps is not a stored property. The CLI computes it by joining the VM's public addresses with commas, and joining nothing gives an empty string. A script that tests for null there will never match.
Step 4: Test reachability separately
Assigned is not reachable. Microsoft: "Before you can connect to a public IP address from the internet, you must open the necessary ports/protocols in your network security groups." Day 57 covered why: a Standard public IP is closed to inbound traffic by default.
timeout 5 bash -c "</dev/tcp/$PIP/22" && echo PORT22_REACHABLE || echo PORT22_BLOCKED
PORT22_REACHABLE
So a rule allowing port 22 was already in place. Had it printed PORT22_BLOCKED, the attachment would still have been correct, with a VM nobody could reach.
/dev/tcp is a Bash feature, which is why this works on a host with no nc or telnet. The Bash manual: "Bash attempts to open the corresponding TCP socket." timeout keeps it from hanging on a port that silently drops packets.
What the docs recommend
Do not look for the public IP inside the VM. Microsoft: "a virtual machine's operating system is unaware of any public IP address assigned to it". ip addr in the guest shows only the private address, and Azure translates between the two.
My notes said to ask the instance metadata service instead. That needs a qualifier. The metadata service has a publicIpAddress field, but for a Standard SKU address, which is the only kind left, Microsoft's page points to the Load Balancer metadata endpoint. I have not tested what the plain field returns.
Dissociate before deleting. Microsoft only allows deleting a public IP that is not associated with any IP configuration, and its page removes the association like this:
az network nic ip-config update -g "$RG" --nic-name "$NIC" -n "$IPCFG" --public-ip-address null
az network public-ip delete -g "$RG" -n devops-pip
My notes used --remove publicIPAddress for the first line. Microsoft's current page uses --public-ip-address null. Either way, the VM loses its address the moment that line runs, so do not run it over an SSH session that arrives through the same address.
What I learned
Stated briefly, the tasks were to mount a volume in a pod and to attach a public IP to a VM. Both statements skip a layer. The claim is what lets storage change underneath a pod. The IP configuration is what lets one NIC carry several addresses. And in the first task, the part that was off was neither of those. It was a path, which no status column checks.
So here is the Day 60 question. When did a test of yours last pass while checking the wrong thing?
Day 60 down. Forty to go.
Top comments (5)
That βtest passed while checking the wrong thingβ part hit home π
Iβve definitely had cases where everything was green, only to realize later that my test was proving something slightly different from what I thought it was proving.
The httpd/nginx path example is a great reminder: green doesnβt always mean correct.
Nice one, Nnamdi. Day 60 already β keep going! π
Thank you! π And the uncomfortable part is that the test wasn't wrong. It proved the volume persisted, which it did, perfectly. It just never asked whether Apache could see it.
So nothing failed. I'd have shipped it, and the only thing that would have told me was someone opening the page.
Thanks for the motivation Mustafa.
60 Days woww I like it! π₯°
All the best for 40Days :D
Thank you! π Sixty down, and the posts have got longer rather than shorter, which wasn't the plan at all.
Forty to go. π
Some comments may only be visible to logged-in visitors. Sign in to view all comments.