DEV Community

Cover image for Day 59: A Broken Deployment and a NIC
Nnamdi Felix Ibe
Nnamdi Felix Ibe

Posted on AI-assisted

Day 59: A Broken Deployment and a NIC

Day 59 of DevOps, Day 9 of Azure. Both tasks today came down to what you check first. In Kubernetes, the events showed me one fault while a second sat unread in the same file. In Azure, the first check I ran was so heavy that it ended the lab session, twice.

One Kubernetes task, one Azure task. Fix a Redis Deployment whose pods will not start, then attach an existing network interface to an existing VM. The tasks come from the KodeKloud Engineer platform.

Task 1: Fix a broken Redis Deployment

The task: the pods of redis-deployment are not running. Find out why and fix it.

Step 1: Look at the pod

kubectl get pods
kubectl describe pod <redis-deployment-pod-name>
Enter fullscreen mode Exit fullscreen mode

The Events section at the bottom of describe had the cause:

MountVolume.SetUp failed for volume "config" : configmap "redis-cofig" not found
Enter fullscreen mode Exit fullscreen mode

Step 2: Compare with what exists

kubectl get cm
Enter fullscreen mode Exit fullscreen mode

The ConfigMap in the cluster is redis-config. The Deployment asks for redis-cofig. One letter is missing.

Step 3: Fix the reference

kubectl get deployment redis-deployment -o yaml > redis-deployment.yaml
Enter fullscreen mode Exit fullscreen mode
volumes:
- name: config
  configMap:
    name: redis-config    # was redis-cofig
Enter fullscreen mode Exit fullscreen mode
kubectl apply -f redis-deployment.yaml
kubectl get pods
Enter fullscreen mode Exit fullscreen mode

Step 4: Find the second fault

The new pod got past the volume and stopped again, this time with ImagePullBackOff. It was the same file, with a second typo:

containers:
- name: redis-container
  image: redis:alpine    # was redis:alpin
Enter fullscreen mode Exit fullscreen mode

Step 5: Recreate and verify

I deleted the Deployment and created it again from the fixed file:

kubectl delete deployment redis-deployment
kubectl apply -f redis-deployment.yaml
Enter fullscreen mode Exit fullscreen mode
kubectl get pods
kubectl logs <redis-pod-name>
kubectl exec <redis-pod-name> -- redis-cli ping
Enter fullscreen mode Exit fullscreen mode
PONG
Enter fullscreen mode Exit fullscreen mode

Running says the container started. PONG says Redis answers, which is the thing the task was really about.

What I would do differently

Read the whole spec before fixing anything. Both typos were in the Deployment from the start. The events only showed the first failure the pod hit, so the image problem stayed hidden until the volume was fixed. One careful read of kubectl get deployment redis-deployment -o yaml would have found both.

Skip the delete. A Deployment rolls out new pods whenever its pod template changes, so applying the fixed file was enough. Deleting it also deleted its ReplicaSets, because Kubernetes cascades deletion to dependents by default, and with them the revision history that kubectl rollout undo relies on, as on Day 52. The app was already down, so it cost nothing here. On a live one, it is an outage.

Use the direct tool for a one-line fix, then wait for the rollout:

kubectl set image deployment/redis-deployment redis-container=redis:alpine
kubectl rollout status deployment/redis-deployment
Enter fullscreen mode Exit fullscreen mode

And preview a change before making it:

kubectl diff -f redis-deployment.yaml
kubectl apply --dry-run=server -f redis-deployment.yaml
Enter fullscreen mode Exit fullscreen mode

The first shows the live object against the version you are about to apply. The second sends the request to the server without saving anything.

What the docs recommend

Start where I started. The Kubernetes debugging guide: "The first step in debugging a Pod is taking a look at it", with kubectl describe pods. It then splits the early failures in two. A pod stuck in Pending "can not be scheduled onto a node". A pod stuck in Waiting has been scheduled and cannot run, and "The most common cause of Waiting pods is a failure to pull the image." Its three checks for that: is the image name correct, has the image been pushed, and can you pull it by hand.

A ConfigMap has to exist before anything references it. The docs: "You must create the ConfigMap object before you reference it in a Pod specification", and unless the reference is marked optional, "the Pod won't start".

ImagePullBackOff means Kubernetes "could not pull a container image", and it keeps retrying with a growing delay, up to 300 seconds.

Do not read the STATUS column as the pod's state. My notes listed ContainerCreating, CrashLoopBackOff and ImagePullBackOff as pod states. The docs define five phases: Pending, Running, Succeeded, Failed and Unknown. The rest is what kubectl displays, and the docs say not to confuse the two.

Manage an object one way. I exported a live object and ran kubectl apply on it, which works, but the docs warn that an object "should be managed using only one technique", and that mixing them "results in undefined behavior". If you do export, strip the server-set fields such as resourceVersion, uid and status first.

Read events early. The API server keeps them for one hour by default.

Task 2: Attach a second NIC to a VM

The task: attach the existing NIC xfusion-nic to the existing VM xfusion-vm, and make sure the VM has finished initialising.

Yesterday's disk went onto a running VM. A NIC cannot. Microsoft: "To add a NIC to an existing VM, first deallocate the VM with az vm deallocate." That makes the order of work matter, because the VM is down from the first command to the last.

Step 1: Check what is cheap to check

Everything that could make the attach fail, I checked while the VM was still running:

az network nic show -g $RG -n xfusion-nic --query "{Subnet:ipConfigurations[0].subnet.id,AttachedTo:virtualMachine.id,Loc:location}"
az network nic show -g $RG -n xfusion-vmVMNic --query "ipConfigurations[0].subnet.id" -o tsv
Enter fullscreen mode Exit fullscreen mode
xfusion-nic:        .../virtualNetworks/xfusion-vmVNET/subnets/xfusion-vmSubnet
xfusion-vmVMNic:    .../virtualNetworks/xfusion-vmVNET/subnets/xfusion-vmSubnet
Enter fullscreen mode Exit fullscreen mode

Same virtual network, AttachedTo was null, and both were in eastus. The first of those is a hard rule. Microsoft: "the network interfaces must all be connected to the same virtual network", though they can sit in different subnets.

Step 2: Deallocate, add, start

az vm deallocate -g $RG -n xfusion-vm
az vm nic add    -g $RG --vm-name xfusion-vm --nics xfusion-nic
az vm start      -g $RG -n xfusion-vm
Enter fullscreen mode Exit fullscreen mode

The command is deallocate, not stop. az vm stop powers the VM off but leaves it, in Microsoft's words, "allocated on a host but not running", and the CLI's own help adds: "The VM will continue to be billed." Deallocating is the state in which the VM "has released the lease on the underlying hardware", and it is the one a NIC change needs.

Step 3: Verify

az network nic show -g $RG -n xfusion-nic --query "{AttachedTo:virtualMachine.id,Primary:primary}"
az vm get-instance-view -g $RG -n xfusion-vm --query "{Status:instanceView.statuses[].displayStatus,Agent:instanceView.vmAgent.statuses[].displayStatus}"
Enter fullscreen mode Exit fullscreen mode

The NIC pointed at the VM with "Primary": false, and the VM was back to VM running with agent Ready.

That false is new. Before the attach, the VM's only NIC showed "Primary": null. With one NIC the flag was unset; with two, the original kept the primary role and the new one arrived as secondary. Microsoft: "By default, the first network interface attached to a VM is the primary network interface."

The check that cost two lab sessions

There is a fourth constraint: each VM size supports a limited number of NICs. To look it up, I ran:

az vm list-skus -l southcentralus --resource-type virtualMachines --query "..."
Enter fullscreen mode Exit fullscreen mode

The lab terminal disconnected, twice. Each disconnect meant a new lab with new resource names.

My notes said the command downloads the SKU catalogue for the region and filters it on the client. The CLI's source shows it is heavier than that. It sends no filter at all, so it fetches the SKU list for the whole subscription, every region, and applies --location and --resource-type afterwards. --query only shrinks what is printed.

The fix was to drop the check. Azure applies the limit itself when a NIC is added. Microsoft: "A VM can only have as many network interfaces attached to it as the VM size supports." I stayed inside the limit, so I have not seen what that failure looks like. The cheap way to know the number in advance is Microsoft's size table, which gives Standard_B1s a maximum of 2 NICs. This second NIC was the last one the VM could take.

What the docs recommend

Plan a NIC change as downtime. This is where Azure and AWS differ. AWS: "You can attach a network interface to an instance when it's running (hot attach), when it's stopped (warm attach), or when the instance is being launched (cold attach)." Azure has only the deallocated route.

Expect side effects from deallocating. Microsoft notes that it "also releases any dynamic IP addresses assigned to the VM", and the temporary disk from Day 52 cannot be relied on afterwards. Disks and networking resources are still charged while the VM is deallocated.

Finish the job in the guest. Attaching the NIC does not make it carry traffic properly. Microsoft's Linux guide: "To send to or from a secondary network interface, you have to manually add persistent routes to the operating system for each secondary network interface." My notes said current Ubuntu images sort this out by themselves. I could not find that in Microsoft's documentation, so I have left it out.

Know which NIC is primary. By default, a VM "sends all outbound traffic to the IP address that's assigned to the primary IP configuration of the primary network interface". Adding a NIC does not move that role. --primary-nic on az vm nic set does.

What I learned

In the first task I checked too little before acting, and found the second fault only by hitting it. In the second, I checked too much, and the check itself took the session down. The useful rule sits between the two: read everything that is cheap to read, and let the operation enforce what is expensive to ask about.

So here is the Day 59 question. When you troubleshoot, do you stop reading at the first error you find, or at the end of the file?

Day 59 down. Forty-one to go.

Top comments (0)