DEV Community

Cover image for Day 58: Grafana and a Managed Disk
Nnamdi Felix Ibe
Nnamdi Felix Ibe

Posted on AI-assisted

Day 58: Grafana and a Managed Disk

Day 58 of DevOps, Day 8 of Azure. Both tasks today had a gap between "created" and "working". A Deployment reports 1/1 before Grafana is serving a single page. A VM reports Provisioning succeeded before the agent inside it is ready, and a disk reports Attached while the operating system sees a blank device.

One Kubernetes task, one Azure task. Deploy Grafana and expose it on a node port, then attach an existing managed disk to an existing VM. The tasks come from the KodeKloud Engineer platform.

Task 1: Deploy Grafana behind a NodePort Service

The task: a Deployment named grafana-deployment-datacenter running a Grafana image, and a NodePort Service exposing it on port 32000.

Step 1: Write the Deployment

apiVersion: apps/v1
kind: Deployment
metadata:
  name: grafana-deployment-datacenter
  labels:
    app: grafana
spec:
  replicas: 1
  selector:
    matchLabels:
      app: grafana
  template:
    metadata:
      labels:
        app: grafana
    spec:
      containers:
      - name: grafana
        image: grafana/grafana:latest
        ports:
        - containerPort: 3000
          name: http-grafana
          protocol: TCP
        resources:
          requests:
            memory: "128Mi"
            cpu: "100m"
          limits:
            memory: "256Mi"
            cpu: "200m"
Enter fullscreen mode Exit fullscreen mode

Port 3000 is Grafana's own default. Its configuration docs say of http_port: "The port to bind to, defaults to 3000."

Step 2: Add the Service to the same file

---
apiVersion: v1
kind: Service
metadata:
  name: grafana-service
  labels:
    app: grafana
spec:
  type: NodePort
  selector:
    app: grafana
  ports:
  - port: 3000
    targetPort: 3000
    nodePort: 32000
    protocol: TCP
    name: http
Enter fullscreen mode Exit fullscreen mode

The --- line lets one file hold both objects. Nothing in the Service names the Deployment. As on Day 56, the only link is the label: app: grafana appears in the Deployment's selector, in the pod template and in the Service's selector, and all three have to agree.

Step 3: Apply and check both objects

kubectl apply -f k3s-deployment.yaml
kubectl get deployments.apps
kubectl get svc
Enter fullscreen mode Exit fullscreen mode

The Deployment should show READY 1/1, and the Service should show TYPE NodePort with PORT(S) 3000:32000/TCP.

Step 4: Open the login page

kubectl get nodes -o wide
# then browse to http://<node-ip>:32000

# or, when the node IP is not reachable from where you sit:
kubectl port-forward deployment/grafana-deployment-datacenter 3000:3000
# then browse to http://localhost:3000
Enter fullscreen mode Exit fullscreen mode

Grafana's sign-in docs give the first login as admin for both username and password, followed by a prompt to change the password.

Step 5: Look at the pod

kubectl logs deployment/grafana-deployment-datacenter
kubectl describe pod -l app=grafana
kubectl top pod -l app=grafana
Enter fullscreen mode Exit fullscreen mode

Selecting by Deployment or by label saves copying the generated pod name. kubectl top only works where Metrics Server is installed; the reference says it "must be installed and running in the cluster for this command to work".

Why 1/1 is not the finish line

My manifest has no readiness probe. The Kubernetes docs say that if a container does not provide a particular probe, "the kubelet always considers the result as Success". So this pod counts as ready the moment its container starts, whether or not Grafana has finished starting inside it. 1/1 told me the container was running. Only the login page told me Grafana was.

What the docs recommend

Grafana publishes its own Kubernetes manifest, and mine falls short of it in several places:

My manifest Grafana's example
Memory 128Mi request, 256Mi limit 750Mi request
CPU 100m request, 200m limit 250m request
Readiness probe None GET /robots.txt on port 3000
Liveness probe None TCP check on port 3000
Storage None 1Gi claim mounted at /var/lib/grafana
targetPort 3000 http-grafana, the port's name

Resources come first. Grafana's Kubernetes page lists minimums of 750 MiB of memory and 250m of CPU. My memory limit is about a third of that. The pod started in the lab, but a limit under the documented minimum is how you meet the out-of-memory kill from Day 50.

Storage matters more. Grafana's docs: "By default, Grafana uses an embedded SQLite version 3 database to store configuration, users, dashboards, and other data." In the container that lives under /var/lib/grafana. With nothing mounted there, every dashboard and user disappears when the pod is replaced.

The probes are short to add. These are the two from Grafana's example, without their timing settings:

readinessProbe:
  httpGet:
    path: /robots.txt
    port: 3000
livenessProbe:
  tcpSocket:
    port: 3000
Enter fullscreen mode Exit fullscreen mode

The named port is a small thing worth copying. The Kubernetes docs: "Port definitions in Pods have names, and you can reference these names in the targetPort attribute of a Service." With targetPort: http-grafana, the number 3000 lives in one place.

Three smaller points remain. Port 32000 sits in the dynamic band of the node port range from Day 56, the band Kubernetes hands out automatically, so a hand-picked port there can already be taken. grafana/grafana:latest is what Grafana's example uses too, though Kubernetes' advice is to avoid :latest in production. And the admin password can come from a Secret through the GF_SECURITY_ADMIN_PASSWORD environment variable, a name that follows Grafana's GF_<SECTION NAME>_<KEY> rule. Grafana documents the setting behind it, admin_password, as "Set once on first-run".

Task 2: Attach a managed disk to a VM

The task: attach the existing disk datacenter-disk to the existing VM datacenter-vm, and make sure the VM has finished initialising.

Step 1: Check the two resources fit

Nothing is being created, so the risk is the two not fitting together. Two read-only calls cover it:

az vm show -g $RG -n datacenter-vm --query "{Loc:location,Zones:zones,Size:hardwareProfile.vmSize}"
az disk show -g $RG -n datacenter-disk --query "{Loc:location,Zones:zones,State:diskState,Sku:sku.name}"
Enter fullscreen mode Exit fullscreen mode
Check VM Disk Result
Region eastus eastus Match
Zone null null Match
Disk state Unattached Free to attach
Premium support Standard_B1s Standard_LRS Not needed

Each row is a way the attach can fail. Microsoft's disks FAQ: "All managed disks, even shared disks, must be in the same region as the VM they're attaching to." A zonal disk attaches "only to VMs in that zone". A disk already attached elsewhere is not free. And Premium SSDs can be used "only with compatible VM series".

One state is worth knowing in advance. A disk on a deallocated VM reads Reserved, not Attached. It is still bound to that VM, and a filter on Attached will miss it.

Step 2: Confirm the VM has finished initialising

az vm wait -g $RG -n datacenter-vm --created
az vm get-instance-view -g $RG -n datacenter-vm --query "{Status:instanceView.statuses[].displayStatus,Agent:instanceView.vmAgent.statuses[].displayStatus}"
Enter fullscreen mode Exit fullscreen mode
{ "Status": ["Provisioning succeeded", "VM running"], "Agent": ["Ready"] }
Enter fullscreen mode Exit fullscreen mode

"Initialised" has two halves. Provisioning succeeded is Azure's side. Agent Ready is the guest's side: the Azure Linux Agent inside the VM has started and reported back. az vm wait --created only waits for the first, so the instance view is the real check.

Step 3: Attach the disk

az vm disk attach -g $RG --vm-name datacenter-vm --name datacenter-disk
Enter fullscreen mode Exit fullscreen mode

The VM stayed VM running throughout. Azure picked the LUN: the REST reference says "If not specified, lun would be auto assigned", and the CLI takes the lowest free number, here 0.

Step 4: Verify from Azure, then from inside the guest

az vm show -g $RG -n datacenter-vm --query "storageProfile.dataDisks[].{Name:name,Lun:lun,Caching:caching}"
az vm run-command invoke -g $RG -n datacenter-vm --command-id RunShellScript --scripts "lsblk"
Enter fullscreen mode Exit fullscreen mode
sda   30G  disk     <- OS disk
sdb    4G  disk     <- temporary disk at /mnt
sdc   30G  disk     <- datacenter-disk, raw
Enter fullscreen mode Exit fullscreen mode

Run Command needs no SSH key and no open port. It uses the VM agent to run the script, so a result also proves the agent works. Microsoft documents its limits: about 20 seconds minimum, one script at a time, and only the last 4,096 bytes of output.

The disk is attached and raw, with no partition or filesystem. The task asked for attachment only, so I left it there.

The query that returned null

Checking the disk's size from both sides gave two different answers:

az disk show ... --query "{GB:diskSizeGb}"     ->  "GB": null
az vm show   ... --query "{GB:diskSizeGb}"     ->  "GB": 30
Enter fullscreen mode Exit fullscreen mode

The raw JSON explained it:

az disk show -g $RG -n datacenter-disk -o json | grep -i disksize
  "diskSizeBytes": 32212254720,
  "diskSizeGB": 30,
Enter fullscreen mode Exit fullscreen mode

az disk show printed diskSizeGB. az vm show printed diskSizeGb. JMESPath is case-sensitive, and a key that does not exist returns null instead of an error, so my query quietly found nothing.

My notes concluded that Azure spells the property differently on the two resource types. It does not. The REST API uses diskSizeGB in both places. The lowercase b came from how the Azure CLI formatted az vm show output, and the CLI's source code suggests that changed in version 2.84.0, which I have not tested. So the same query can work on one CLI version and return null on the next.

The fix does not depend on which version you have. When a value you know exists comes back null, suspect the query before the resource, and look at the raw object with -o json | grep -i.

What the docs recommend

Do not reference the disk as /dev/sdc. Microsoft: "Device paths in Linux aren't guaranteed to be consistent across restarts." Use the filesystem UUID, or the link the Azure Linux Agent creates for the LUN:

/dev/disk/azure/scsi1/lun0
Enter fullscreen mode Exit fullscreen mode

Choose caching when you attach. The disk came up with caching None, and the REST reference explains that default: "None for Standard storage. ReadOnly for Premium storage." My notes had it as a rule about data disks; it is a rule about the storage type. Microsoft's guidance is None for write-only and write-heavy disks, and ReadOnly for read-only and read-write ones. Changing the setting later detaches and reattaches the disk, so it is better set once, with --caching on the attach command.

Before detaching, unmount the disk in the guest and remove its /etc/fstab entry. And remember that detaching does not delete. Microsoft's warning is that a detached disk "is not automatically deleted", and you keep paying for it:

az disk list --query "[?diskState=='Unattached'].{Name:name,RG:resourceGroup,Sku:sku.name}" -o table
Enter fullscreen mode Exit fullscreen mode

What I learned

Every status today was true and incomplete. 1/1 meant a container had started. Provisioning succeeded meant Azure had finished its half. Attached meant a blank device had appeared. The checks that settled things were the ones that used the result: the login page, the agent's Ready, and lsblk from inside the VM.

So here is the Day 58 question. Which of your checks proves that something was created, and which proves that it works?

Day 58 down. Forty-two to go.

Top comments (0)