DEV Community

Tanay Jain
Tanay Jain

Posted on AI-assisted

Moving Beyond the Happy Path: Failure Engineering, EndpointSlices, and Kubernetes Debugging

Most Kubernetes tutorials stop when the Pods become Running or Ready.

I wanted to understand what happens when those signals look healthy but the application path is still broken.

So I built and debugged a containerized Flask + PostgreSQL application on a local multi-node kind Kubernetes cluster, then deliberately reproduced failures across configuration, Service routing, dependencies, probes, storage, scheduling, RBAC, and infrastructure.

The goal was not to create a production platform.

The goal was to make Kubernetes failures observable, diagnosable, recoverable, and reproducible.


1. Architecture: What I Built

Architecture Diagram

The capstone separates the Flask workload from PostgreSQL while using Kubernetes primitives for service discovery, configuration, health checks, identity, and storage.

                         ┌──────────────────────┐
                         │    Test Client Pod   │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │    flask-app-svc     │
                         │      ClusterIP       │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │      flask-app       │
                         │ Deployment: 2        │
                         │ replicas             │
                         └──────────┬───────────┘
                                    │
                              DB connection
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │    postgres-svc      │
                         │      ClusterIP       │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │      postgres        │
                         │ Deployment: 1        │
                         │ replica              │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │    postgres-pvc      │
                         │ Persistent storage   │
                         └──────────────────────┘
Enter fullscreen mode Exit fullscreen mode

The application request/dependency path is:

Client
→ flask-app-svc
→ Flask application
→ postgres-svc
→ PostgreSQL
Enter fullscreen mode Exit fullscreen mode

PostgreSQL persistence is provided separately through:

PostgreSQL
→ postgres-pvc
Enter fullscreen mode Exit fullscreen mode

Core resources

Resource Purpose
flask-app Flask Deployment with 2 replicas
flask-app-svc Internal ClusterIP Service for Flask
flask-app-config ConfigMap for non-sensitive configuration
flask-app-secret Kubernetes Secret for database credentials
postgres PostgreSQL Deployment with 1 replica
postgres-svc Internal ClusterIP Service for PostgreSQL
postgres-pvc PersistentVolumeClaim for PostgreSQL storage
flask-app-sa Dedicated ServiceAccount

The Flask workload uses flask-app-sa, with automountServiceAccountToken: false configured on the Pod template because the application does not require Kubernetes API access.

For the actual Flask workload, no Role or RoleBinding is granted because there is no Kubernetes API permission requirement.


2. Health Checks: Running Is Not the Same as Healthy

The Flask workload uses three probe mechanisms:

  • Startup probe: /health
  • Readiness probe: /health
  • Liveness probe: /

The important distinction is what each one is supposed to answer.

The readiness check is dependency-aware. If PostgreSQL becomes unavailable, the Flask workload can remain running while becoming unready for normal Service traffic.

The liveness check is independent of PostgreSQL. That means a database dependency failure does not automatically become a Flask process-restart condition.

The startup probe provides startup protection before normal readiness and liveness behavior takes over.

This separation became important during failure testing.


3. The Mindset Shift: From Assumption to Evidence

One of the most useful moments in the project came when the PostgreSQL Pod was healthy, but the Flask application still could not reliably reach PostgreSQL.

The easy assumption would have been:

“The database Pod is Running, so the database must be fine.”

That turned out to be incomplete.

The better question became:

What evidence proves where the failure actually is?

From there, I repeatedly used the same diagnostic loop:

Symptom
→ Observation
→ Hypothesis
→ Evidence
→ Decisive Evidence
→ Root Cause
→ Fix
→ Verification
Enter fullscreen mode Exit fullscreen mode

The point was not to memorize more kubectl commands.

It was to understand what each piece of evidence actually proves.


4. The Important Lesson: A Healthy Pod Does Not Prove a Healthy Service Path

A dependency Pod being Running does not automatically prove that an application can reach that dependency through its Kubernetes Service path.

For the Flask → PostgreSQL path, I learned to reason through the layers:

Application
→ Service
→ EndpointSlice
→ Target Pod
Enter fullscreen mode Exit fullscreen mode

Each layer answers a different question.

The application tells me what the client actually experiences.

The Service defines the selection and forwarding behavior.

The EndpointSlice shows the backend endpoints Kubernetes currently associates with the Service.

The target Pod tells me whether the selected workload is actually running.

That distinction became especially useful during Service-routing failures.

A Service can exist and have a ClusterIP while still having no usable backends.

A Service can also have populated endpoints while traffic still fails because the forwarding configuration is wrong.

EndpointSlice state is therefore strong evidence about Kubernetes backend selection, but it is not blanket proof that every part of the underlying network path is healthy.

That is why the debugging process also used Pod events, node state, direct connectivity tests where appropriate, and application-level evidence.


5. Failure Engineering Across Kubernetes Layers

Instead of validating only the happy path, I deliberately reproduced failures and then recovered them.

The debugging history covers multiple Kubernetes layers.

Configuration

I reproduced configuration failures including incorrect ConfigMap values and incorrect key references.

The diagnostic path was straightforward:

Pod problem
→ inspect Pod Events
→ identify configuration failure
→ compare expected vs actual key/value
→ fix
→ verify rollout
Enter fullscreen mode Exit fullscreen mode

This reinforced an important habit: inspect the evidence from the failing object before guessing at the fix.

Service routing

I reproduced Service selector mismatches and targetPort misconfigurations.

This produced two different failure signatures.

With a selector mismatch, the Service had no usable backends.

With a targetPort mismatch, the endpoints could be populated while application traffic still failed.

That distinction is easy to miss if the investigation stops at:

“The Service exists.”

PostgreSQL dependency

I deliberately disrupted the PostgreSQL dependency and observed how the Flask workload responded.

The important behavior was that Flask could remain running while its readiness state changed.

That meant a dependency failure could affect Service traffic without unnecessarily turning into a process-restart problem.

Authentication

I reproduced database authentication failures and used application-level evidence to distinguish authentication problems from routing or DNS problems.

This was another example of why the same high-level symptom can have very different root causes.

Probes

I tested readiness failures, liveness failures, and startup protection.

The practical distinction became:

  • Readiness controls whether the workload is considered ready for normal Service traffic.
  • Liveness can cause the container to restart.
  • Startup probing protects legitimate startup time before normal probe behavior takes over.

Resources and scheduling

The debugging history also covers resource and scheduling exercises including:

  • memory enforcement and OOMKilled
  • CPU throttling
  • ResourceQuota rejection
  • Pods remaining Pending because of taints
  • node-related scheduling failures

These exercises reinforced another principle:

The same high-level symptom can originate at different Kubernetes layers.

A Pending Pod is not automatically a resource problem.

A Running Pod is not automatically an application-health guarantee.

A Service with a ClusterIP is not automatically a working traffic path.

Storage

I reproduced storage-related behavior including node-local hostPath limitations and PVCs remaining Pending because of incorrect StorageClass configuration.

That clarified an important distinction:

Persistent storage semantics are not the same thing as disaster recovery.

A PVC can preserve data across Pod replacement without providing a backup or restore strategy.

RBAC and identity

The repository also contains dedicated RBAC exercises.

For the actual Flask workload, the final design intentionally grants no Kubernetes API permissions because the application does not require them.

That boundary was verified with kubectl auth can-i.

Infrastructure-level failure

The independent engineering challenge went below the application layer.

The recorded incident involved:

  • new Pods remaining Pending
  • FailedScheduling evidence
  • the worker node becoming NotReady
  • an unreachable-node taint
  • the underlying worker container exiting
  • a resulting internal DNS failure

The recovery required restoring the worker/runtime state and restarting CoreDNS, followed by verification of node health and application readiness.

The lesson was important:

A Kubernetes application symptom can originate below the application layer entirely.


6. A Representative Debugging Pattern

Across these failures, the tools changed less than the reasoning.

A simplified diagnostic hierarchy became:

What is the symptom?
        ↓
Which object currently shows it?
        ↓
What do Events say?
        ↓
What does the relevant controller / Service / EndpointSlice say?
        ↓
What does the application say?
        ↓
What evidence eliminates competing hypotheses?
        ↓
What is the smallest correct fix?
        ↓
Can the healthy state be verified again?
Enter fullscreen mode Exit fullscreen mode

This was the part of the project I expect to carry forward into future Kubernetes work.


7. Reproducibility: Verifying the Repository Instead of Trusting the README

A repository is not reproducible merely because its README says it is.

So I re-ran the documented reproduction procedure in an isolated namespace:

sep28-repro
Enter fullscreen mode Exit fullscreen mode

The documented reproduction verified:

  1. PostgreSQL Secret creation
  2. PostgreSQL PVC binding
  3. PostgreSQL rollout
  4. PostgreSQL EndpointSlice population
  5. Flask Secret creation
  6. Flask ConfigMap creation
  7. Flask ServiceAccount creation
  8. Flask rollout
  9. Flask EndpointSlice population
  10. Application /health verification

The actual recorded environment included two Ready kind nodes, a bound postgres-pvc, successful PostgreSQL and Flask rollouts, and populated EndpointSlices.

The final application response was:

{"database":"connected","status":"healthy"}
Enter fullscreen mode Exit fullscreen mode

The recorded reproduction result was:

PASS
Enter fullscreen mode Exit fullscreen mode

The reproduction environment was then cleaned up by deleting the isolated namespace and removing the temporary local secrets.

This was an important milestone because the project was no longer only “working on the current cluster.” The documented procedure had been executed again in an isolated namespace.


8. Security and Least-Privilege Reasoning

The project also forced me to separate several concepts that are easy to blur together.

ServiceAccount

The Flask workload uses:

automountServiceAccountToken: false
Enter fullscreen mode Exit fullscreen mode

because it does not need an in-cluster Kubernetes API credential.

RBAC

The actual application workload has no Role or RoleBinding because it has no Kubernetes API requirement.

The principle is simple:

Do not grant Kubernetes API permissions to a workload that does not need them.

Secrets

Database credentials are supplied through a Kubernetes Secret.

Runtime credentials are generated locally rather than being committed as repository configuration.

And:

Base64 encoding is not encryption.

The encoded representation of a Kubernetes Secret should not be confused with confidentiality.

The production-readiness review therefore identifies stronger secret management and encryption-at-rest controls as future production requirements.


9. Known Limitations

This project is explicitly a:

Demonstration / learning project. NOT production-deployed.

That distinction matters.

PostgreSQL availability

PostgreSQL runs as a single replica.

There is no database failover or stateful high-availability mechanism.

Backup and restore

The PVC provides persistence semantics, but backup/restore has not been tested.

A production system would need an actual backup strategy and a tested restore procedure.

Resource tuning

The CPU and memory requests/limits are reasoned baseline values for the local learning environment.

They are not load-tested or production-tuned measurements.

Storage durability

The project uses the local-path storage behavior of the kind environment.

That is useful for learning Kubernetes storage mechanics, but it is not equivalent to production cloud storage with appropriate disaster-recovery characteristics.

Observability

The project demonstrates evidence-driven troubleshooting using native Kubernetes primitives.

It does not include centralized logging, distributed tracing, or automated alerting.

Supply chain

The Flask image reference is pinned rather than using latest, but the project does not yet implement image signing, provenance attestations, SBOM enforcement, or admission controls.


10. What I Would Improve for Production

The next stage would not simply be:

“Deploy the same YAML to the cloud.”

The production gaps are architectural.

I would move toward:

  • managed or highly available PostgreSQL
  • tested backup and restore
  • measured resource tuning and load testing
  • stronger external secrets management
  • stronger image supply-chain verification
  • production networking and policy controls
  • centralized observability
  • Helm packaging where it provides a clear operational benefit
  • eventually, an EKS-based deployment

These are future improvements, not completed features of this project.


11. What This Project Actually Taught Me

The most valuable lesson was not a particular kubectl command.

It was learning to stop treating Kubernetes status as the complete story.

A Pod can be:

Running
Enter fullscreen mode Exit fullscreen mode

while the application is still broken.

A Service can exist while its traffic path is wrong.

A dependency can be healthy while the path to that dependency is not.

A node can fail while the application symptoms appear much higher in the stack.

The more useful question became:

What evidence proves where the observed behavior stops matching the intended state?

That shift changed how I approach Kubernetes debugging.


12. Conclusion

This project started as a Flask + PostgreSQL deployment exercise.

It became a broader Kubernetes debugging laboratory.

I learned by deliberately creating failures, identifying the layer where behavior diverged from the intended state, collecting evidence, applying the smallest correct fix, and then verifying recovery.

The result is not a production platform.

It is something more appropriate for this stage of learning:

a reproducible environment for practicing how Kubernetes actually behaves when things go wrong.

The complete repository, manifests, reproduction procedure, debugging log, and production-readiness review are available here:

flask-devops-app - Kubernetes Deployment

Project Overview

A Kubernetes-deployed Flask application integrated with PostgreSQL, using Kubernetes Deployments, Services, ConfigMap/Secret configuration, health probes, a PersistentVolumeClaim, and a dedicated ServiceAccount. The project is a hands-on Kubernetes/DevOps learning system focused on deployment, dependency-aware health, failure engineering, diagnosis, recovery, and reproducibility.

Project Purpose

This project exists as a hands-on Kubernetes engineering and debugging exercise. The work focuses on understanding Kubernetes reconciliation, service discovery, storage, probes, configuration, least-privilege identity, controlled failure injection, evidence-based diagnosis, recovery, and reproducibility.

Architecture

                         ┌──────────────────────┐
                         │   Test Client Pod    │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │    flask-app-svc     │
                         │       ClusterIP      │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │     flask-app        │
                         │     Deployment       │
                         │      2 replicas      │
                         └──────────┬───────────┘
                                    │
                       DB connection│
                                    ▼
                         ┌──────────────────────┐
                         │    postgres-svc      │
                         │       ClusterIP      │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │      postgres        │
                         │     Deployment       │
                         │      1 replica       │
                         └──────────┬───────────┘
                                    │
                                    ▼
                         ┌──────────────────────┐
                         │    postgres-pvc      │
                         │   persistent data    │
                         └──────────────────────┘
…

Top comments (0)