DEV Community

Aisalkyn Aidarova
Aisalkyn Aidarova

Posted on

FINAL KUBERNETES PRODUCTION CAPSTONE Dev Staging Production | Helm | Argo CD | Security | Storage | 10 Production Incidents

“I built and operated a production-style Kubernetes platform with separate dev, staging, and production environments. I deployed applications with Helm, implemented GitOps with Argo CD, configured resources, probes, autoscaling, availability, RBAC, networking, persistent storage, and practiced production incident response.”

Your role

You have joined a financial company as a DevOps Engineer.

The company is building an online banking platform:

Customers
    ↓
Banking Application
    │
    ├── frontend
    ├── account-service
    ├── payment-service
    └── transaction-service
             ↓
          Database
Enter fullscreen mode Exit fullscreen mode

The company wants Kubernetes environments built using production practices.

You are responsible for designing, deploying, securing, operating, troubleshooting, and explaining the platform.

This is not a copy/paste homework.

Every student must record a video showing the environment and explaining every major decision.


PART 1 — Architecture Design

Before creating anything, draw your architecture.

Your recording must explain:

                         INTERNET
                            │
                            ▼
                    Load Balancer / ALB
                            │
                            ▼
                         Ingress
                            │
                 ┌──────────┼──────────┐
                 ▼          ▼          ▼
             frontend    account    payment
                            │          │
                            └────┬─────┘
                                 ▼
                             Database
Enter fullscreen mode Exit fullscreen mode

Then explain the Kubernetes architecture underneath:

                    KUBERNETES / EKS
                           │
              ┌────────────┴────────────┐
              │                         │
        CONTROL PLANE              WORKER NODES
              │                         │
        API Server                     kubelet
        Scheduler                      runtime
        Controllers                    CNI
        etcd                           kube-proxy
                                        │
                                       Pods
Enter fullscreen mode Exit fullscreen mode

In your recording explain:

What happens from kubectl apply until the container is actually running?

You should be able to explain approximately:

kubectl
   ↓
API Server
   ↓
desired state stored
   ↓
Controllers
   ↓
Scheduler
   ↓
Worker Node selected
   ↓
kubelet
   ↓
CRI
   ↓
container runtime
   ↓
Container starts
   ↓
CNI networking
   ↓
Readiness
   ↓
Service can route traffic
Enter fullscreen mode Exit fullscreen mode

Do not just read component definitions.

Explain the process.


PART 2 — Build Three Environments

The company requires:

DEV
STAGING
PRODUCTION
Enter fullscreen mode Exit fullscreen mode

For this homework, you can implement them as namespaces:

kubectl create namespace dev
kubectl create namespace staging
kubectl create namespace production
Enter fullscreen mode Exit fullscreen mode

Verify:

kubectl get namespaces
Enter fullscreen mode Exit fullscreen mode

But explain an important production point

In your recording answer:

“Would a bank necessarily run dev, staging and production merely as three namespaces in one cluster?”

Expected explanation:

Not necessarily.

For stronger isolation, especially in regulated/security-sensitive organizations, production may use separate clusters and often separate AWS accounts/VPC boundaries.

For example:

AWS Organization
│
├── Development Account
│      └── Dev EKS
│
├── Staging / Pre-Prod Account
│      └── Staging EKS
│
└── Production Account
       └── Production EKS
Enter fullscreen mode Exit fullscreen mode

Namespaces are acceptable for our training environment because creating several full EKS environments costs additional money and operational effort.

But students must understand the difference between:

Lab architecture and real production architecture.


PART 3 — Build the Application

Create a Deployment for the application.

Requirements:

Application: banking-app
Production replicas: 3
Container: nginx or your own application
Container port: 80
Enter fullscreen mode Exit fullscreen mode

Your Deployment must include:

resources:
  requests:
    cpu:
    memory:
  limits:
    cpu:
    memory:
Enter fullscreen mode Exit fullscreen mode

and:

readinessProbe:
Enter fullscreen mode Exit fullscreen mode

and:

livenessProbe:
Enter fullscreen mode Exit fullscreen mode

and a production rolling strategy:

strategy:
  type: RollingUpdate
  rollingUpdate:
    maxUnavailable: 0
    maxSurge: 1
Enter fullscreen mode Exit fullscreen mode

Students must decide appropriate values and explain them.


PART 4 — Explain Requests and Limits

In the recording answer:

Why do we use requests?

Why do we use limits?

What happens when memory exceeds its limit?

What can happen when CPU reaches/exceeds its limit?

Why should we not blindly give every Pod huge requests?

Explain:

Requests
   ↓
Scheduling / reserved expectation

Limits
   ↓
Maximum resource boundary
Enter fullscreen mode Exit fullscreen mode

Also explain why bad resource configuration can create:

Pending Pods
CPU throttling
OOMKilled
Wasted cluster capacity
Higher AWS cost
Enter fullscreen mode Exit fullscreen mode

PART 5 — Probes

Implement:

Startup Probe (optional if appropriate)
Readiness Probe
Liveness Probe
Enter fullscreen mode Exit fullscreen mode

Explain:

Readiness:
Should this Pod receive traffic?

Liveness:
Should Kubernetes restart the container?

Startup:
Has this slow-starting application successfully started?
Enter fullscreen mode Exit fullscreen mode

Demonstrate what happens when readiness fails.


PART 6 — Service

Create:

ClusterIP Service
Enter fullscreen mode Exit fullscreen mode

Explain:

Why shouldn't another application communicate directly with a Pod IP?

Explain:

Service
   ↓
Selector
   ↓
Pod labels
   ↓
EndpointSlices
   ↓
Ready Pods
Enter fullscreen mode Exit fullscreen mode

Show:

kubectl get svc -n production
kubectl get endpointslice -n production
kubectl get pods -n production --show-labels
Enter fullscreen mode Exit fullscreen mode

PART 7 — Ingress

Create an Ingress for the application.

Conceptually:

Customer
   ↓
DNS
   ↓
Load Balancer
   ↓
Ingress Controller
   ↓
Ingress Rules
   ↓
Service
   ↓
Pods
Enter fullscreen mode Exit fullscreen mode

In your recording explain:

Ingress vs Ingress Controller

Ingress vs Service

Where HTTPS/TLS could terminate

How traffic reaches the application

For AWS/EKS, explain how an AWS load-balancing solution can integrate with Kubernetes rather than claiming the Ingress resource itself magically creates everything.


PART 8 — ConfigMap and Secret

Create a ConfigMap containing something such as:

APP_ENV=production
LOG_LEVEL=info
Enter fullscreen mode Exit fullscreen mode

Create a Secret containing a dummy training credential.

Application must consume configuration through environment variables or mounted files.

Explain:

ConfigMap vs Secret

And answer:

“Is base64 encryption?”

Correct answer:

No.

For the banking scenario, discuss why real secrets should be managed carefully and why cloud secret-management/encryption solutions may be integrated rather than committing credentials to Git.

Never commit real secrets to the repository.


PART 9 — RBAC / Least Privilege

Create:

ServiceAccount
Role
RoleBinding
Enter fullscreen mode Exit fullscreen mode

Create a developer ServiceAccount.

It should be allowed to:

get Pods
list Pods
watch Pods
Enter fullscreen mode Exit fullscreen mode

It should not be allowed to:

delete Pods
delete Deployments
delete Secrets
Enter fullscreen mode Exit fullscreen mode

Test:

kubectl auth can-i list pods \
--as=system:serviceaccount:production:developer \
-n production
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl auth can-i delete pods \
--as=system:serviceaccount:production:developer \
-n production
Enter fullscreen mode Exit fullscreen mode

Explain:

Authentication
=
Who are you?

Authorization
=
What are you allowed to do?
Enter fullscreen mode Exit fullscreen mode

Then explain why least privilege is especially important for a banking application.


PART 10 — NetworkPolicy

Create a production NetworkPolicy.

Architecture:

frontend
   │
   ▼
payment-service
   │
   ▼
database
Enter fullscreen mode Exit fullscreen mode

Allowed:

frontend → payment ✓
payment → database ✓
Enter fullscreen mode Exit fullscreen mode

Not allowed:

frontend → database ✗
random Pod → database ✗
Enter fullscreen mode Exit fullscreen mode

Students must explain why RBAC cannot solve this problem.

Expected:

RBAC
→ Kubernetes API authorization

NetworkPolicy
→ workload network communication
Enter fullscreen mode Exit fullscreen mode

Also explain that NetworkPolicy requires a network implementation that enforces it.


PART 11 — HPA

Configure HPA.

Example requirements:

minimum replicas: 3
maximum replicas: 10
CPU target: 60%
Enter fullscreen mode Exit fullscreen mode

Students must show:

kubectl get hpa -n production
kubectl describe hpa -n production
Enter fullscreen mode Exit fullscreen mode

Explain:

What happens when traffic increases?

And:

What happens if HPA reaches maxReplicas but traffic continues increasing?

Also answer:

What does CPU-based HPA need?

Metrics and appropriate resource requests.


PART 12 — PDB

Create a PodDisruptionBudget.

For example:

3 replicas

minAvailable: 2
Enter fullscreen mode Exit fullscreen mode

Explain why this matters during voluntary disruption such as node maintenance.

Also explain:

Does PDB protect against every failure?

No.


PART 13 — Scheduling

Demonstrate or explain:

nodeSelector
node affinity
pod anti-affinity
taints
tolerations
Enter fullscreen mode Exit fullscreen mode

For production, explain why we might avoid placing every replica on the same node.

Bad:

Node 1
├── payment-1
├── payment-2
└── payment-3

Node 1 fails
→ payment completely unavailable
Enter fullscreen mode Exit fullscreen mode

Better:

Node 1 → payment-1
Node 2 → payment-2
Node 3 → payment-3
Enter fullscreen mode Exit fullscreen mode

Discuss topology/AZ resilience where appropriate.


PART 14 — StatefulSet + Persistent Storage

Create a StatefulSet.

It must demonstrate:

Stable Pod identity
PVC
StorageClass
PV
Persistent data
Enter fullscreen mode Exit fullscreen mode

Show:

kubectl get statefulset
kubectl get pvc
kubectl get pv
kubectl get storageclass
Enter fullscreen mode Exit fullscreen mode

Write data.

Delete the Pod.

Allow StatefulSet to recreate it.

Prove the data still exists.

Then explain:

Stable identity and persistent storage are related but different concepts.


PART 15 — AWS EBS CSI

Explain:

StatefulSet
    ↓
PVC
    ↓
StorageClass
    ↓
EBS CSI Driver
    ↓
AWS EBS
    ↓
PV
    ↓
PVC Bound
    ↓
Pod Running
Enter fullscreen mode Exit fullscreen mode

Students must explain the storage incident we previously encountered:

PVC Pending
↓
StorageClass existed
↓
ebs.csi.aws.com expected
↓
EBS CSI add-on/components missing
↓
No dynamic EBS volume
↓
Pod Pending
Enter fullscreen mode Exit fullscreen mode

And how to investigate it.


PART 16 — Helm

Convert the application into a Helm chart.

Suggested structure:

banking-app/
│
├── Chart.yaml
├── values.yaml
├── values-dev.yaml
├── values-staging.yaml
├── values-prod.yaml
│
└── templates/
    ├── deployment.yaml
    ├── service.yaml
    ├── ingress.yaml
    ├── configmap.yaml
    ├── serviceaccount.yaml
    ├── hpa.yaml
    ├── pdb.yaml
    └── networkpolicy.yaml
Enter fullscreen mode Exit fullscreen mode

Different environment values should demonstrate meaningful differences.

For example:

DEV
replicas: 1
smaller resources

STAGING
replicas: 2
production-like configuration

PRODUCTION
replicas: 3+
HPA
PDB
stronger availability settings
Enter fullscreen mode Exit fullscreen mode

Demonstrate:

helm lint
helm template
helm upgrade --install
helm list
helm history
Enter fullscreen mode Exit fullscreen mode

Explain why Helm is useful in a company with many applications and environments.


PART 17 — Argo CD / GitOps

Put Kubernetes/Helm configuration in Git.

Architecture:

Developer
   ↓
Pull Request
   ↓
Review
   ↓
Merge
   ↓
Git
   ↓
Argo CD
   ↓
EKS
Enter fullscreen mode Exit fullscreen mode

Students must create/configure an Argo CD Application and demonstrate:

Synced
Healthy
Enter fullscreen mode Exit fullscreen mode

Then manually modify something in Kubernetes.

Observe:

OutOfSync
Enter fullscreen mode Exit fullscreen mode

Explain drift.

Then restore/reconcile through GitOps.

Critical production question:

“If Argo CD manages production, should engineers routinely modify production using kubectl edit?”

Expected:

Generally no. Desired changes should normally go through the controlled Git/change process, with emergency procedures handled according to company policy.


PART 18 — Banking Production Requirements

This section is mandatory in the recording.

Imagine this is a banking system handling financial transactions and sensitive data.

Before saying “production ready,” discuss:

High availability
Least privilege
Encryption
Secret management
Network segmentation
Auditability
Backups
Disaster recovery
Monitoring
Alerting
Logging
Change control
Image scanning
Vulnerability management
Patch management
Capacity planning
Rollback strategy
Business continuity
Compliance requirements
Enter fullscreen mode Exit fullscreen mode

Do not simply say:

“We have three replicas, therefore the bank is production-ready.”

Production readiness is much larger than Kubernetes manifests.


PART 19 — What Costs Money in AWS?

Students must be able to answer this because engineers should understand cost.

Discuss potential charges for things such as:

EKS cluster/control-plane service
EC2 worker nodes
EBS volumes/snapshots
Load balancers
NAT Gateway
Data transfer
CloudWatch logs/metrics
ECR storage/data transfer
Public IPv4 where applicable
Route 53
Backup/storage services
Other managed AWS services
Enter fullscreen mode Exit fullscreen mode

Do not memorize exact prices because pricing changes.

Instead explain:

“Which architectural decisions create cost?”

For example:

More nodes → more compute cost

More/larger EBS → more storage cost

Multiple ALBs → load-balancer cost

Large log ingestion → observability cost

NAT traffic → networking cost
Enter fullscreen mode Exit fullscreen mode

PART 20 — Production Monitoring

Your application must have an observability plan.

Explain what you would monitor:

CPU
Memory
Pod restarts
Ready replicas
Node health
Request rate
Error rate
Latency
HTTP 5xx
HPA behavior
PVC/storage
Database health
Dependency health
Enter fullscreen mode Exit fullscreen mode

And explain:

Metrics
Logs
Traces
Enter fullscreen mode Exit fullscreen mode

A strong student should mention service-level signals such as:

Traffic
Errors
Latency
Saturation
Enter fullscreen mode Exit fullscreen mode

not only Kubernetes CPU.


PART 21 — TEN REQUIRED PRODUCTION INCIDENTS

This is the most important part of the homework.

Each student must break the environment intentionally, investigate it without jumping directly to the answer, fix/mitigate it, and verify recovery on video.

Incident 1 — Bad Deployment / ImagePullBackOff

Break:

Deploy an invalid/nonexistent image tag.

Student must discover:

New rollout
↓
New Pods unhealthy
↓
ImagePullBackOff
↓
describe Pod
↓
Events
↓
Image cannot be pulled
Enter fullscreen mode Exit fullscreen mode

Mitigate:

kubectl rollout history deployment/<name> -n production
kubectl rollout undo deployment/<name> -n production
kubectl rollout status deployment/<name> -n production
Enter fullscreen mode Exit fullscreen mode

Interview explanation:

“A new release introduced an invalid image reference. I correlated the incident with the rollout, confirmed the image pull failure from Pod events, rolled back to the known-good revision, and verified service recovery.”


Incident 2 — Production Slow / HPA at Maximum

Generate load.

Students investigate:

Application slow
↓
Pods Running
↓
CPU/resource pressure
↓
HPA target exceeded
↓
maxReplicas reached
↓
capacity insufficient for current load
Enter fullscreen mode Exit fullscreen mode

They must not say “traffic” without evidence.


Incident 3 — Service Selector Mismatch

Break Service selector:

Service:
app=wrong-app

Pods:
app=banking-app
Enter fullscreen mode Exit fullscreen mode

Symptoms:

Pods 1/1 Running
Service exists
Application unreachable
No expected Service backend
Enter fullscreen mode Exit fullscreen mode

Student must compare:

kubectl describe svc
kubectl get pods --show-labels
kubectl get endpointslice
Enter fullscreen mode Exit fullscreen mode

Incident 4 — Readiness Probe Failure

Break:

readinessProbe:
  httpGet:
    path: /does-not-exist
Enter fullscreen mode Exit fullscreen mode

Expected:

Running 0/1
Enter fullscreen mode Exit fullscreen mode

Student must explain:

Running ≠ Ready.

Use describe to find probe failure.

Fix/rollback and verify readiness and Service routing.


Incident 5 — OOMKilled

Configure a workload with insufficient memory or a controlled memory-hungry training application.

Student must find:

RESTARTS ↑
↓
CrashLoopBackOff
↓
describe
↓
Last State
↓
OOMKilled
Enter fullscreen mode Exit fullscreen mode

Then:

kubectl logs <pod> --previous
Enter fullscreen mode Exit fullscreen mode

Explain why simply increasing memory may only be mitigation if the application has unbounded memory growth.


Incident 6 — Dependency Failure

Application logs:

Cannot connect to database/Kafka
Enter fullscreen mode Exit fullscreen mode

Student is not allowed to say:

“Kafka is down.”

until proven.

Investigate:

DNS
Service
EndpointSlice
Port
NetworkPolicy
Authentication
Dependency health
Application configuration
Enter fullscreen mode Exit fullscreen mode

Interview statement:

“The log proves the application cannot connect to the dependency; it does not by itself prove the dependency is down.”


Incident 7 — PVC Pending

Break storage provisioning.

Expected:

Pod Pending
↓
PVC Pending
↓
Events
↓
StorageClass
↓
CSI
↓
Cloud permissions/storage
Enter fullscreen mode Exit fullscreen mode

Students must walk the entire storage dependency chain.


Incident 8 — RBAC Forbidden

Make a ServiceAccount attempt:

kubectl delete pod ...
Enter fullscreen mode Exit fullscreen mode

Expected:

Forbidden
Enter fullscreen mode Exit fullscreen mode

Student investigates:

kubectl auth can-i ...
Enter fullscreen mode Exit fullscreen mode

Then Role/RoleBinding.

They must explain:

Authentication succeeded
but
Authorization denied
Enter fullscreen mode Exit fullscreen mode

A Forbidden response is different from a network failure.


Incident 9 — NetworkPolicy Blocks Traffic

Create:

frontend → payment ✗
Enter fullscreen mode Exit fullscreen mode

because NetworkPolicy blocks the connection.

Students must prove:

Pods Running
Service exists
Endpoints exist
DNS may work
BUT network connection blocked
Enter fullscreen mode Exit fullscreen mode

Then inspect:

kubectl get networkpolicy
kubectl describe networkpolicy
Enter fullscreen mode Exit fullscreen mode

They must explain why restarting the application doesn't fix the policy.


Incident 10 — Argo CD OutOfSync / Configuration Drift

Git says:

replicas: 3
Enter fullscreen mode Exit fullscreen mode

Manually change live cluster:

kubectl scale deployment banking-app \
--replicas=1 \
-n production
Enter fullscreen mode Exit fullscreen mode

Observe Argo CD.

Students explain:

Git desired state
      ≠
Live state

→ OutOfSync
Enter fullscreen mode Exit fullscreen mode

Restore desired state through the GitOps workflow/reconciliation.

Explain why uncontrolled manual production changes are dangerous.


PART 22 — For Every Incident, Use the Same Professional Method

Students may not begin with random commands.

Every recording must follow:

1. ALERT / TICKET

2. CUSTOMER IMPACT
   What is broken?

3. SCOPE
   Which application/service/environment?

4. WHAT CHANGED?
   Deployment?
   Config?
   Traffic?
   Infrastructure?

5. OBSERVE
   Metrics
   Logs
   Events
   Traces

6. HYPOTHESIS
   What might explain the evidence?

7. PROVE / DISPROVE
   Collect evidence.

8. ROOT CAUSE
   What actually caused the problem?

9. MITIGATION / FIX
   Restore service safely.

10. VERIFY
    Prove service is healthy.

11. RCA
    Why did this happen?
    How do we prevent recurrence?
Enter fullscreen mode Exit fullscreen mode

The sentence I want every student to learn is:

“I don't want to guess the root cause. I form a hypothesis and validate it with evidence.”


PART 23 — Production Verification

After every fix, students must prove recovery.

They cannot say:

“The command succeeded, so production is fixed.”

They must check the relevant layers:

Deployment healthy
       ↓
Pods Ready
       ↓
Restart count stable
       ↓
Service backend healthy
       ↓
Application responds
       ↓
Errors normal
       ↓
Latency normal
       ↓
Critical customer flow works
Enter fullscreen mode Exit fullscreen mode

PART 24 — Git Repository Requirements

Repository should look professional.

For example:

banking-platform/
│
├── README.md
│
├── docs/
│   ├── architecture.md
│   ├── troubleshooting.md
│   └── incidents.md
│
├── helm/
│   └── banking-app/
│       ├── Chart.yaml
│       ├── values.yaml
│       ├── values-dev.yaml
│       ├── values-staging.yaml
│       ├── values-prod.yaml
│       └── templates/
│
├── argocd/
│   ├── dev.yaml
│   ├── staging.yaml
│   └── production.yaml
│
└── kubernetes/
    ├── rbac/
    ├── network-policy/
    ├── storage/
    └── troubleshooting/
Enter fullscreen mode Exit fullscreen mode

README should explain architecture, environments, deployment process, security controls, availability, scaling, storage, GitOps and troubleshooting.


PART 25 — Required Video Recording

I would require approximately 45–60 minutes, but quality matters more than exact duration.

Tell students:

Do not read definitions from Google or from your notes. Share your screen and explain your own environment as though I am the interviewer.

The recording must show and explain the full lifecycle:

Architecture
    ↓
AWS / EKS
    ↓
Dev / Staging / Production
    ↓
Docker Image / ECR
    ↓
Deployment / ReplicaSet / Pods
    ↓
Resources / Probes
    ↓
Service / EndpointSlice
    ↓
Ingress
    ↓
ConfigMap / Secret
    ↓
RBAC
    ↓
NetworkPolicy
    ↓
HPA / PDB
    ↓
Scheduling
    ↓
StatefulSet / PV / PVC / StorageClass / CSI
    ↓
Helm
    ↓
Argo CD / GitOps
    ↓
Observability
    ↓
10 Production Incidents
    ↓
RCA / Verification
Enter fullscreen mode Exit fullscreen mode

At random points in the recording they should be able to answer:

Why did you configure it this way?

That is more important than merely showing YAML.


Top comments (0)