Day 61 of DevOps, Day 11 of Azure. Both tasks today were about something that has to finish before the real work starts. An init container must exit before the application container is allowed to run. And a VM has to restart, sometimes on different hardware, before a new size takes effect.
One Kubernetes task, one Azure task. Run an init container that prepares a file for the main container, then resize a VM from one size to the next. The tasks come from the KodeKloud Engineer platform.
Task 1: An init container that prepares a file
The task: a Deployment called ic-deploy-devops with one replica. An init container, ic-msg-devops, writes a message to a file on a shared volume. The main container, ic-main-devops, prints that file every five seconds.
Step 1: Write the Deployment
apiVersion: apps/v1
kind: Deployment
metadata:
name: ic-deploy-devops
labels:
app: ic-devops
spec:
replicas: 1
selector:
matchLabels:
app: ic-devops
template:
metadata:
labels:
app: ic-devops
spec:
initContainers:
- name: ic-msg-devops
image: debian:latest
command: ['/bin/bash', '-c', 'echo Init Done - Welcome to xFusionCorp Industries > /ic/media']
volumeMounts:
- name: ic-volume-devops
mountPath: /ic
containers:
- name: ic-main-devops
image: debian:latest
command: ['/bin/bash', '-c', 'while true; do cat /ic/media; sleep 5; done']
volumeMounts:
- name: ic-volume-devops
mountPath: /ic
volumes:
- name: ic-volume-devops
emptyDir: {}
There are two lists of containers and one volume. Both containers mount the same emptyDir at /ic, and that volume is the only thing connecting them. The init container writes /ic/media and exits. The main container reads it.
Step 2: Apply and watch the pod start
kubectl apply -f k3s-deployment.yaml
kubectl get deployments.apps
kubectl get pods -w
The pod passes through Init:0/1, then PodInitializing, then Running. The Kubernetes docs read Init:N/M as: "The Pod has M Init Containers, and N have completed so far." So Init:0/1 means the one init container has not finished yet.
Step 3: Confirm the init container finished
kubectl logs <pod-name> -c ic-msg-devops
kubectl describe pod <pod-name> | grep -A 20 "Init Containers"
The logs are empty, and that is correct. The echo is redirected into the file, so nothing goes to stdout. The evidence is in describe: State Terminated, Reason Completed, Exit Code 0.
Step 4: Confirm the main container reads the file
kubectl logs <pod-name> -c ic-main-devops
Init Done - Welcome to xFusionCorp Industries
Init Done - Welcome to xFusionCorp Industries
A new line appears every five seconds. To skip looking up the pod name, kubectl accepts the Deployment instead:
kubectl logs deployment/ic-deploy-devops -c ic-main-devops
Step 5: Delete the pod and watch it run again
kubectl delete pod -l app=ic-devops
kubectl get pods -w
The Deployment creates a replacement, and the replacement goes through Init:0/1 again. The file was on an emptyDir, so it went with the old pod, and the init container writes it afresh for the new one.
Two things my notes had wrong
My notes said the pod shows Init:Running while the init container runs. That is not a status the docs list. Their table has Init:N/M, Init:Error and Init:CrashLoopBackOff, and the count form is what kubectl prints while an init container is still going.
They also said init containers run again "even if main containers are restarted". The Pod Lifecycle page says otherwise: "init containers run only once (if successful), during Pod startup." By default, a main container crashing and restarting does not bring the init container back. A new pod does, which is what Step 5 showed. The same page describes one opt-in exception, a feature-gated restart rule called RestartAllContainers, which re-runs init containers too.
What the docs recommend
The basic contract comes first. "Init containers always run to completion", and "Each init container must complete successfully before the next one starts." If one fails, "the kubelet repeatedly restarts that init container until it succeeds", and the application container never starts in the meantime.
Write them to be safe to repeat. The docs: "Because init containers can be restarted, retried, or re-executed, init container code should be idempotent." Mine is, because > overwrites the file. An init container that appended with >>, or created something that must not already exist, would behave differently on its second run.
Do not expect probes. Regular init containers do not support livenessProbe, readinessProbe, startupProbe or lifecycle. They are meant to exit, so there is nothing to keep checking.
Count their resources. The scheduler does not add init containers to the total. It takes the highest request among them, compares it with the sum of the application containers' requests, and uses the larger. A heavy setup step can make a pod need more room than the application itself ever uses.
Know the related form from Day 55. A native sidecar is declared in the same initContainers list, with one extra line, restartPolicy: Always. That line is the whole difference: with it, the container starts first and keeps running; without it, the container must exit before anything else starts.
And debian:latest was the task's choice. Outside a lab, pin the tag.
Task 2: Resize a VM
The task: change nautilus-vm from Standard_B1s to Standard_B2s and leave it running.
Step 1: Ask what this VM can resize to
az vm list-vm-resize-options -g $RG -n nautilus-vm --query "[?name=='Standard_B2s']" -o table
Name Cores
------------ -------
Standard_B2s 2
A VM runs on a hardware cluster that supports a particular set of sizes. Microsoft: "Deallocation may be required if the new size isn't available on the hardware cluster currently hosting the VM." This query lists the sizes this one VM can be resized to. A row means Standard_B2s is on that list. An empty result means it is not, and that is the case where deallocating first may be needed.
Step 2: Resize
The query listed Standard_B2s. I deallocated anyway:
az vm deallocate -g $RG -n nautilus-vm
az vm resize -g $RG -n nautilus-vm --size Standard_B2s
az vm start -g $RG -n nautilus-vm
The shorter route would have been the middle command by itself:
az vm resize -g $RG -n nautilus-vm --size Standard_B2s
My notes recorded the three-command version as the slower path, taken when a cheaper one was available. That is fair, with one qualification: Microsoft's own example script does the same. It checks the resize options and then "deallocates the VM, resizes it, and starts it again". The one-command version does not release the VM from its hardware first, so it should mean less downtime, although Microsoft's page gives no timings.
Neither avoids a restart. Microsoft: "Even when deallocation isn't required, changing the size of a running VM will cause it to restart."
Step 3: Verify size and state separately
az vm show -g $RG -n nautilus-vm -d --query "{Size:hardwareProfile.vmSize,Power:powerState,Prov:provisioningState}"
az vm get-instance-view -g $RG -n nautilus-vm --query "{Status:instanceView.statuses[].displayStatus,Agent:instanceView.vmAgent.statuses[].displayStatus}"
{ "Size": "Standard_B2s", "Power": "VM running", "Prov": "Succeeded" }
{ "Agent": ["Ready"], "Status": ["Provisioning succeeded", "VM running"] }
There are two requirements, so there are two checks. Size covers the resize. VM running with agent Ready covers the running state, and as on Day 58, the agent is the half that proves the guest came back up.
One caveat about the first check comes from Microsoft's resize page. If a resize fails, "the VM model will still display the requested size, but the VM will continue running on its previous size until the resize is successfully allocated". So the size field reports what was asked for. A CPU count from inside the guest would settle it, using Day 58's Run Command:
az vm run-command invoke -g $RG -n nautilus-vm --command-id RunShellScript --scripts "nproc"
On a Standard_B2s it should print 2.
The column that vanished
My first version of the Step 1 query asked for three fields:
--query "[?name=='Standard_B2s'].{Name:name,Cores:numberOfCores,MemoryMB:memoryInMb}" -o table
Two columns came back. MemoryMB was not blank. It was gone.
Two things combined. The property is spelt memoryInMB in the REST reference, with a capital B, so my memoryInMb matched nothing and returned null. Then -o table dropped the column, because every value in it was null. A wrong field name produced output that looked complete, with no error anywhere.
It is Day 58's lesson again, one step worse. Write queries against -o json, where a null is at least visible, and switch to -o table once the field names are confirmed.
What the docs recommend
Check before you deallocate. az vm list-vm-resize-options is scoped to one VM and answered on the server. It is the opposite of the az vm list-skus call that ended two lab sessions on Day 59.
Plan for the restart. Anything in memory is lost, and the temporary disk from Day 52 cannot be relied on. If you deallocate, Microsoft adds that it "also releases any dynamic IP addresses assigned to the VM. The OS and data disks are not affected."
Know what the new size changes. From Microsoft's size table:
| Standard_B1s | Standard_B2s | |
|---|---|---|
| vCPUs | 1 | 2 |
| Memory | 1 GiB | 4 GiB |
| Temporary disk | 4 GiB | 8 GiB |
| Max data disks | 2 | 4 |
| Max NICs | 2 | 3 |
The limits matter in reverse. A VM using three NICs, or more than two data disks, would not fit back into a B1s.
I removed one claim. My notes said B-series CPU credits are lost when a VM is deallocated. I could not find that in Microsoft's current B-series pages, which describe the credits a VM starts with and say nothing about stopping it. So it is not in this post as a fact.
What I learned
An init container is a rule about order: this must succeed before that may begin. A resize has the same shape, with the restart as the step you cannot skip. In both, the question worth asking early is what has to finish first, and what it costs while it does.
So here is the Day 61 question. What in your system assumes a setup step runs exactly once, and what happens the day it runs twice?
Day 61 down. Thirty-nine to go.
Top comments (1)
Wow, time flies! Weβre already at 61. I wonder if I will have finished writing the final stratagem by your 100-day mark. π