DEV Community

Oleksandr Kuryzhev
Oleksandr Kuryzhev

Posted on Originally published at kuryzhev.cloud

Build an Interactive Bash CLI for DevOps Reports

Originally published on kuryzhev.cloud


The scenario

A platform team keeps a folder of one-off shell snippets for checking cluster health, and every on-call engineer pastes them in a slightly different order. What they want is one interactive Bash CLI that shows a menu during an incident, prints a clean table on demand, and still runs unattended in a pipeline. It should not need a web dashboard, a database, or a language runtime beyond what is already on the jump host.

Bash is a reasonable fit when the data already comes from command-line tools that emit JSON. The script's job is to fetch, filter, format, and decide an exit code. It is a poor fit once you need persistent history, charts, or multi-user access, and at that point Grafana or a similar tool is the right answer.

This guide builds a small tool called opsreport. It reports unhealthy Kubernetes pods and nodes that are not Ready. It has three modes: an interactive menu, a live refresh view, and a non-interactive mode that emits TSV and returns a meaningful exit code. The same pattern works for AWS CLI, Docker, or Terraform output. Swap the collector functions and keep the rest.

Prerequisites

You need Bash 4.4 or newer. Most current Linux distributions ship Bash 5.x, but the macOS system Bash is still 3.2, so install a newer one with Homebrew and point the shebang at /usr/bin/env bash. Check with bash --version.

Tools used by the collectors:

  • kubectl with a working kubeconfig and read access to pods and nodes.
  • jq 1.6 or newer for JSON filtering.
  • column from util-linux or bsdextrautils (formerly part of bsdmainutils on Debian and Ubuntu) for aligned tables.
  • shellcheck for linting. It is optional but strongly recommended.

The only design rule to settle up front is that the script must behave differently when no terminal is attached. Menus, colors, and screen clearing are for humans. CI logs and pipes need plain, stable output. The next step builds that switch in first, because retrofitting it later usually means touching every function.

Step 1: Script skeleton with terminal detection

Start with strict mode, terminal detection, and optional color. Bash's -t test checks whether a file descriptor is attached to a terminal. The NO_COLOR environment variable is a widely followed convention for disabling ANSI color.

#!/usr/bin/env bash
set -Eeuo pipefail   # exit on errors and unset vars; a pipeline fails if any stage fails

MODE="menu"          # menu | watch | report
REPORT="pods"
INTERVAL=5

# Interactive only if both stdin and stdout are terminals
INTERACTIVE=0
[[ -t 0 && -t 1 ]] && INTERACTIVE=1

# Color only on a terminal, and respect NO_COLOR
if [[ -t 1 && -z "${NO_COLOR:-}" ]]; then
  RED=$'\e[31m'; GRN=$'\e[32m'; RST=$'\e[0m'
else
  RED=""; GRN=""; RST=""
fi

usage() {
  cat <<'EOF'
Usage: opsreport [--watch] [--report pods|nodes] [--interval SECONDS]
  no flags      interactive menu (requires a terminal)
  --report X    print report X (TSV when piped); exit 0 clean,
                1 problems found, 2 usage error, 3 collector failure
  --watch       refresh the pods report until 'q' is pressed
EOF
}

need_arg() { [[ $# -ge 2 ]] || { echo "missing value for $1" >&2; usage >&2; exit 2; }; }

while [[ $# -gt 0 ]]; do
  case "$1" in
    --report)   need_arg "$@"; MODE="report"; REPORT="$2"; shift 2 ;;
    --watch)    MODE="watch"; shift ;;
    --interval) need_arg "$@"; INTERVAL="$2"; shift 2 ;;
    -h|--help)  usage; exit 0 ;;
    *) echo "unknown option: $1" >&2; usage >&2; exit 2 ;;
  esac
done

if [[ ! $INTERVAL =~ ^[1-9][0-9]*$ ]]; then
  echo "--interval must be a positive integer" >&2; exit 2
fi

# Menu and watch modes make no sense without a terminal
if [[ $MODE != "report" && $INTERACTIVE -eq 0 ]]; then
  echo "no terminal detected, falling back to --report pods" >&2
  MODE="report"
fi

Watch out for set -e inside conditionals and pipelines. It is suppressed in some contexts, such as the left side of && or ||. It is also suppressed inside any function called from those contexts. Do not assume every failure aborts the script. Explicit checks after critical commands are still worth writing.

Step 2: Collector functions that return data, not decoration

Keep data gathering separate from presentation. Each collector prints tab-separated rows and nothing else, which makes it easy to test, pipe, or format later. Both queries use kubectl get -o json and jq. See the official kubectl reference for output options and the exact field names in the pod status.

unhealthy_pods() {
  kubectl get pods --all-namespaces -o json |
    jq -r '
      .items[]
      | select(.status.phase != "Succeeded")            # ignore completed Jobs
      | select(
          .status.phase != "Running"
          or ((.status.containerStatuses // []) | any(.ready == false))
        )                                               # Running but not Ready still counts
      | [.metadata.namespace, .metadata.name, .status.phase] | @tsv'
}

notready_nodes() {
  kubectl get nodes -o json |
    jq -r '
      .items[]
      | { name: .metadata.name,
          ready: (([.status.conditions[]? | select(.type == "Ready") | .status] | .[0]) // "Unknown") }
      | select(.ready != "True")                        # only NotReady or Unknown nodes
      | [.name, .ready] | @tsv'
}

# Presentation: header + rows, aligned and colored only when a human is looking
render() {
  local header="$1"; shift
  local rows
  rows="$("$@")" || return 3        # collector failed: not the same as "clean"
  if [[ -z "$rows" ]]; then
    if [[ $INTERACTIVE -eq 1 ]]; then printf '%sNo problems found%s\n' "$GRN" "$RST"; fi
    return 0
  fi
  if [[ $INTERACTIVE -eq 1 ]]; then
    { printf '%b\n' "$header"; printf '%s\n' "$rows"; } | column -t -s $'\t'
    printf '%sProblems found%s\n' "$RED" "$RST"
  else
    printf '%s\n' "$rows"
  fi
  return 1                          # findings
}

Two gotchas here. First, a pod can report phase Running while a container is failing its readiness probe. That is why the filter checks containerStatuses and not just the phase.

Second, render returns 1 for findings and 3 when the collector itself fails, for example when kubectl loses access. Keeping those apart stops an RBAC error from looking like a clean cluster. The check is explicit because set -e is disabled when render runs under || true. Under set -e, any nonzero return aborts the caller unless you handle it with || true or capture the status explicitly.

Step 3: Menu, live view, and exit codes

Bash's built-in select gives you a numbered menu without extra dependencies. It writes the prompt (from PS3) to standard error and reads the choice from standard input. The live view uses read -t with a timeout, so the refresh loop doubles as a keypress listener.

show() {
  case "$1" in
    pods)  render "NAMESPACE\tPOD\tPHASE" unhealthy_pods ;;
    nodes) render "NODE\tREADY" notready_nodes ;;
    *) echo "unknown report: $1" >&2; return 2 ;;
  esac
}

run_menu() {
  local PS3="Select a report: "
  select _ in "Unhealthy pods" "NotReady nodes" "Live pod view" "Quit"; do
    case "$REPLY" in
      1) show pods  || true ;;   # findings are not a script error here
      2) show nodes || true ;;
      3) run_watch ;;
      4) break ;;
      *) echo "Enter 1-4" ;;
    esac
  done
}

run_watch() {
  local key=""
  tput civis 2>/dev/null || true                # hide cursor
  trap 'tput cnorm 2>/dev/null || true' EXIT    # always restore it
  trap 'exit 130' INT TERM                      # Ctrl-C still quits; EXIT trap runs
  while true; do
    clear
    echo "Unhealthy pods (every ${INTERVAL}s, q to quit) - $(date +%T)"
    show pods || true
    key=""
    read -r -s -t "$INTERVAL" -n1 key || true   # timeout is not a failure
    [[ "$key" == "q" ]] && break
  done
  tput cnorm 2>/dev/null || true
}

case "$MODE" in
  menu)   run_menu ;;
  watch)  run_watch ;;
  report) rc=0; show "$REPORT" || rc=$?; exit "$rc" ;;   # 0 / 1 / 2 / 3 to CI
esac

The exit-code contract matters more than the menu. Return 0 for clean, 1 for findings, 2 for usage errors, and 3 when data could not be collected. A pipeline can then gate on the report without parsing text. The contract is documented in --help so nobody has to read the source.

Verify and test

Start with static checks. Run shellcheck opsreport and fix every warning about unquoted variables. Unquoted variables are the usual source of surprises when names contain unexpected characters. Kubernetes names are restricted, but the habit of quoting is cheap.

Then test the three behaviors separately, replacing the real cluster with a stub where you can. A simple way is to put a fake kubectl script earlier in PATH that prints a saved JSON fixture. That lets you test edge cases such as a completed Job, a CrashLoopBackOff pod, a Running-but-unready pod, and a NotReady node without breaking anything.

  • Non-interactive: run ./opsreport --report pods | cat. Because stdout is a pipe, you should see plain TSV rows with no header, no color codes, and no menu.
  • Exit codes: run ./opsreport --report pods; echo $? against a fixture with problems (expect 1) and a clean fixture (expect 0). Make the stub exit nonzero and expect 3. Run ./opsreport --report bogus; echo $? and expect 2.
  • Interactive: run ./opsreport in a terminal, pick each menu entry, then start the live view and press q. Confirm the cursor comes back, including after Ctrl-C. If it does not, the trap is not firing.
  • Color: run NO_COLOR=1 ./opsreport --report pods in a terminal and confirm the status line has no color escape sequences.

If you want repeatable regression tests, the Bats framework is a common choice for shell scripts. Check its current documentation for installation, since packaging differs by platform.

Also test with your real RBAC identity. A read-only role that cannot list pods across all namespaces will make kubectl fail. The report then exits 3 instead of 0 or 1, which is the correct behavior but worth seeing once.

Where to take it next

The pattern scales by adding collectors and keeping the split between data, presentation, and exit codes. Candidates include failed AWS CodeDeploy deployments, Docker containers restarting repeatedly, or Terraform workspaces with drift. If you outgrow a numbered menu, fzf or gum can add fuzzy selection. They add an install dependency, so keep a plain select fallback for bastion hosts. Keep the script small enough to read in one sitting, and move to a Go or Python tool when the logic outgrows it. For more Bash automation patterns, browse the rest of kuryzhev.cloud.

Related

Top comments (0)