DEV Community

Cover image for Ditching Web UIs: Building a Terminal-Native Kubernetes Monitor with Local AI
Joydeep Banik
Joydeep Banik

Posted on

Ditching Web UIs: Building a Terminal-Native Kubernetes Monitor with Local AI

Hacktoberfest Weekend Challenge: Build for a Friend Submission 🤝

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built KubeSense, a high-density, terminal-based Kubernetes monitoring dashboard.
I built this for my junior batchmate, who is an aspiring DevOps engineer. During one of his internships, he expressed how constant pod monitoring is cumbersome and he wanted something that is terminal native so that he does not have to switch windows while working unlike heavy web-based UIs like Lens or Rancher. He wished something lightweight that stays in the CLI but is smart enough to highlight exactly what matters and filters out the unnecessary jargon. KubeSense solves this by bringing real-time metrics and AI-driven log anomaly detection directly to their terminal window.

Demo

KubeSense Terminal Dashboard Interface

Code

The project is publicly available at: https://github.com/Joydeep2005Banik/KubeSense

How I Built It

KubeSense is built entirely in Python, utilizing a decoupled architecture driven by a central PodMonitor engine. Here is a high-level look at the architecture:

graph TD
    subgraph UI Layer
        A[Terminal UI]
    end
    subgraph Core Logic
        B[Monitor Engine]
    end
    subgraph Data Sources
        C[K8s API]
        D[SSH Fallback]
    end
    subgraph AI Processing
        E[Log Optimizer]
        F[LLM Providers]
    end

    B <-->|State Updates| A
    B -->|Fetch Metrics| C
    B -->|Fallback| D
    B -->|Raw Logs| E
    E -->|Clean Context| F
    F -->|Alerts| A

Here is a breakdown of the core systems:

1. Reactive Terminal UI (Textual)

I used the Textual framework to build the dashboard. Instead of relying on heavy "while-true" redraw loops, the UI uses Textual's event-driven reactive properties. When the Monitor Engine pushes new metrics, only the changed widgets redraw, ensuring a lightning-fast, zero-overhead experience without screen tearing.

flowchart LR
    A[Monitor Engine] -->|Sends updated<br/>Metrics Object| B(Reactive Variable)
    B -->|Triggers Auto-Update| C[UI Node / Widget]
    C -->|Only redraws changed region| D[Terminal Display]

2. Resilient Metric Collection & SSH Fallback

The tool primarily connects to the cluster using the official kubernetes Python client to fetch live status and log streams. However, if the cluster's K8s Metrics Server goes down, KubeSense intelligently engages a graceful fallback mechanism to connect directly to the underlying cluster nodes via asyncssh.

sequenceDiagram
    participant M as Monitor
    participant K as K8s Metrics API
    participant S as SSH Fallback
    participant N as Node Shell

    M->>K: Fetch Pod CPU/Mem
    alt Metrics Server Unavailable
        K-->>M: HTTP 404 / Error
        M->>S: Engage Fallback (asyncssh)
        S->>N: Connect to Node via SSH
        N-->>S: Execute `top` & parse usage
        S-->>M: Return hardware metrics
    else Success
        K-->>M: Standard API Metrics
    end

3. Log Optimization Pipeline

Logs can be incredibly noisy and expensive to process. Before any logs touch the AI, they pass through a custom LogOptimizer:

flowchart LR
    A[Raw Log Stream] --> B{Severity Filter}
    B -- Keep WARNING+ --> C[Deduplication]
    C -->|Collapse repeated| D[PII Scrubbing]
    D -->|Mask IPs/UUIDs| E[Optimized Payload]

This ensures the context window remains small, protecting privacy and saving tokens.

4. Multi-Provider AI Engine (Local-First)

For the AI core, I implemented an abstract factory pattern to handle the anomaly detection. While it supports BYOK for cloud providers like Groq or OpenAI, the default engine runs entirely locally via Ollama.

flowchart TD
    A[Optimized Payload] --> B{AI Factory Config}
    B -->|groq| C[GroqProvider<br/>Cloud API]
    B -->|openai| D[OpenAIProvider<br/>Cloud API]
    B -->|ollama| E[OllamaProvider<br/>Local/Air-gapped]
    B -->|mock| F[MockProvider<br/>Heuristics Fallback]

    C & D & E & F --> G[JSON Enforcement & Parser]
    G --> H[Anomaly Alerts to UI]

Key Features

  • Smart Log Filtering & Deduplication: Aggressively filters noise (INFO, DEBUG), deduplicates repeated errors, and scrubs PII before any AI processing.
  • BYOK (Bring Your Own Key): While local open-weight models via Ollama are the default, you can easily plug in an API key for Groq or OpenAI via config.yaml or environment variables if you prefer cloud inference.
  • SSH Fallback: Automatically drops to a direct SSH connection to scrape node hardware via top if the Kubernetes Metrics Server isn't available.

As KubeSense streams the optimized container logs, it passes these batches to an open-weight model running locally. The model analyzes the logs in the background, strictly returning structured JSON, and alerts to warnings, hidden errors, or unusual patterns before they escalate without adding any overhead to the cluster itself.

Why Does Open Innovation Matter?

Deploying open-weight models locally resolves critical architectural and operational limitations associated with closed APIs:

1. Data Sovereignty & Air-Gapped Execution

Kubernetes logs frequently leak sensitive infrastructure data, internal IP schemas, and environment variables. Streaming these logs to a third-party server via a closed API introduces unacceptable security risks and compliance violations (e.g., SOC2/HIPAA). By running inference locally via Ollama, KubeSense ensures complete data sovereignty. The tool can operate entirely offline, making AI-driven log analysis viable for strict, air-gapped environments.

2. O(1) Inference Costs

Streaming high-volume log data through a token-metered API creates unbounded operational costs. Local inference with open-weight models reduces the marginal cost of analysis to zero, making continuous 24/7 log processing financially viable.

3. Hardware-Specific Model Tuning

Closed ecosystems abstract away model selection. An open architecture allows users to define the exact AI engine based on their hardware constraints. Systems with limited VRAM can deploy highly efficient models like Google's Gemma (gemma:2b), while robust systems can load higher-parameter models (e.g., llama3.1:8b) for complex log reasoning.

4. Deterministic Availability

Cloud AI APIs introduce network latency and are subject to HTTP 429 Too Many Requests rate limits. When triaging a production cluster outage, network dependency and API throttling are unacceptable. Local open-weight models guarantee deterministic availability and consistent latency independent of external network conditions.

Prize Categories

  • Best Use of Gemma (Using Google's Gemma open-weight model via Ollama for local log anomaly detection)

Let's Chat! 💬

If you have any questions about how the architecture works, why I chose Textual, or how to set up Ollama for your own cluster monitoring, please drop a comment below! I'd love to discuss why I think lightweight, terminal-native tools are often far more effective than heavy dashboards, and hear your thoughts on local AI for DevOps.

Top comments (0)