“I built and operated a production-style Kubernetes platform with separate dev, staging, and production environments. I deployed applications with Helm, implemented GitOps with Argo CD, configured resources, probes, autoscaling, availability, RBAC, networking, persistent storage, and practiced production incident response.”
Your role
You have joined a financial company as a DevOps Engineer.
The company is building an online banking platform:
Customers
↓
Banking Application
│
├── frontend
├── account-service
├── payment-service
└── transaction-service
↓
Database
The company wants Kubernetes environments built using production practices.
You are responsible for designing, deploying, securing, operating, troubleshooting, and explaining the platform.
This is not a copy/paste homework.
Every student must record a video showing the environment and explaining every major decision.
PART 1 — Architecture Design
Before creating anything, draw your architecture.
Your recording must explain:
INTERNET
│
▼
Load Balancer / ALB
│
▼
Ingress
│
┌──────────┼──────────┐
▼ ▼ ▼
frontend account payment
│ │
└────┬─────┘
▼
Database
Then explain the Kubernetes architecture underneath:
KUBERNETES / EKS
│
┌────────────┴────────────┐
│ │
CONTROL PLANE WORKER NODES
│ │
API Server kubelet
Scheduler runtime
Controllers CNI
etcd kube-proxy
│
Pods
In your recording explain:
What happens from kubectl apply until the container is actually running?
You should be able to explain approximately:
kubectl
↓
API Server
↓
desired state stored
↓
Controllers
↓
Scheduler
↓
Worker Node selected
↓
kubelet
↓
CRI
↓
container runtime
↓
Container starts
↓
CNI networking
↓
Readiness
↓
Service can route traffic
Do not just read component definitions.
Explain the process.
PART 2 — Build Three Environments
The company requires:
DEV
STAGING
PRODUCTION
For this homework, you can implement them as namespaces:
kubectl create namespace dev
kubectl create namespace staging
kubectl create namespace production
Verify:
kubectl get namespaces
But explain an important production point
In your recording answer:
“Would a bank necessarily run dev, staging and production merely as three namespaces in one cluster?”
Expected explanation:
Not necessarily.
For stronger isolation, especially in regulated/security-sensitive organizations, production may use separate clusters and often separate AWS accounts/VPC boundaries.
For example:
AWS Organization
│
├── Development Account
│ └── Dev EKS
│
├── Staging / Pre-Prod Account
│ └── Staging EKS
│
└── Production Account
└── Production EKS
Namespaces are acceptable for our training environment because creating several full EKS environments costs additional money and operational effort.
But students must understand the difference between:
Lab architecture and real production architecture.
PART 3 — Build the Application
Create a Deployment for the application.
Requirements:
Application: banking-app
Production replicas: 3
Container: nginx or your own application
Container port: 80
Your Deployment must include:
resources:
requests:
cpu:
memory:
limits:
cpu:
memory:
and:
readinessProbe:
and:
livenessProbe:
and a production rolling strategy:
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
Students must decide appropriate values and explain them.
PART 4 — Explain Requests and Limits
In the recording answer:
Why do we use requests?
Why do we use limits?
What happens when memory exceeds its limit?
What can happen when CPU reaches/exceeds its limit?
Why should we not blindly give every Pod huge requests?
Explain:
Requests
↓
Scheduling / reserved expectation
Limits
↓
Maximum resource boundary
Also explain why bad resource configuration can create:
Pending Pods
CPU throttling
OOMKilled
Wasted cluster capacity
Higher AWS cost
PART 5 — Probes
Implement:
Startup Probe (optional if appropriate)
Readiness Probe
Liveness Probe
Explain:
Readiness:
Should this Pod receive traffic?
Liveness:
Should Kubernetes restart the container?
Startup:
Has this slow-starting application successfully started?
Demonstrate what happens when readiness fails.
PART 6 — Service
Create:
ClusterIP Service
Explain:
Why shouldn't another application communicate directly with a Pod IP?
Explain:
Service
↓
Selector
↓
Pod labels
↓
EndpointSlices
↓
Ready Pods
Show:
kubectl get svc -n production
kubectl get endpointslice -n production
kubectl get pods -n production --show-labels
PART 7 — Ingress
Create an Ingress for the application.
Conceptually:
Customer
↓
DNS
↓
Load Balancer
↓
Ingress Controller
↓
Ingress Rules
↓
Service
↓
Pods
In your recording explain:
Ingress vs Ingress Controller
Ingress vs Service
Where HTTPS/TLS could terminate
How traffic reaches the application
For AWS/EKS, explain how an AWS load-balancing solution can integrate with Kubernetes rather than claiming the Ingress resource itself magically creates everything.
PART 8 — ConfigMap and Secret
Create a ConfigMap containing something such as:
APP_ENV=production
LOG_LEVEL=info
Create a Secret containing a dummy training credential.
Application must consume configuration through environment variables or mounted files.
Explain:
ConfigMap vs Secret
And answer:
“Is base64 encryption?”
Correct answer:
No.
For the banking scenario, discuss why real secrets should be managed carefully and why cloud secret-management/encryption solutions may be integrated rather than committing credentials to Git.
Never commit real secrets to the repository.
PART 9 — RBAC / Least Privilege
Create:
ServiceAccount
Role
RoleBinding
Create a developer ServiceAccount.
It should be allowed to:
get Pods
list Pods
watch Pods
It should not be allowed to:
delete Pods
delete Deployments
delete Secrets
Test:
kubectl auth can-i list pods \
--as=system:serviceaccount:production:developer \
-n production
Then:
kubectl auth can-i delete pods \
--as=system:serviceaccount:production:developer \
-n production
Explain:
Authentication
=
Who are you?
Authorization
=
What are you allowed to do?
Then explain why least privilege is especially important for a banking application.
PART 10 — NetworkPolicy
Create a production NetworkPolicy.
Architecture:
frontend
│
▼
payment-service
│
▼
database
Allowed:
frontend → payment ✓
payment → database ✓
Not allowed:
frontend → database ✗
random Pod → database ✗
Students must explain why RBAC cannot solve this problem.
Expected:
RBAC
→ Kubernetes API authorization
NetworkPolicy
→ workload network communication
Also explain that NetworkPolicy requires a network implementation that enforces it.
PART 11 — HPA
Configure HPA.
Example requirements:
minimum replicas: 3
maximum replicas: 10
CPU target: 60%
Students must show:
kubectl get hpa -n production
kubectl describe hpa -n production
Explain:
What happens when traffic increases?
And:
What happens if HPA reaches
maxReplicasbut traffic continues increasing?
Also answer:
What does CPU-based HPA need?
Metrics and appropriate resource requests.
PART 12 — PDB
Create a PodDisruptionBudget.
For example:
3 replicas
minAvailable: 2
Explain why this matters during voluntary disruption such as node maintenance.
Also explain:
Does PDB protect against every failure?
No.
PART 13 — Scheduling
Demonstrate or explain:
nodeSelector
node affinity
pod anti-affinity
taints
tolerations
For production, explain why we might avoid placing every replica on the same node.
Bad:
Node 1
├── payment-1
├── payment-2
└── payment-3
Node 1 fails
→ payment completely unavailable
Better:
Node 1 → payment-1
Node 2 → payment-2
Node 3 → payment-3
Discuss topology/AZ resilience where appropriate.
PART 14 — StatefulSet + Persistent Storage
Create a StatefulSet.
It must demonstrate:
Stable Pod identity
PVC
StorageClass
PV
Persistent data
Show:
kubectl get statefulset
kubectl get pvc
kubectl get pv
kubectl get storageclass
Write data.
Delete the Pod.
Allow StatefulSet to recreate it.
Prove the data still exists.
Then explain:
Stable identity and persistent storage are related but different concepts.
PART 15 — AWS EBS CSI
Explain:
StatefulSet
↓
PVC
↓
StorageClass
↓
EBS CSI Driver
↓
AWS EBS
↓
PV
↓
PVC Bound
↓
Pod Running
Students must explain the storage incident we previously encountered:
PVC Pending
↓
StorageClass existed
↓
ebs.csi.aws.com expected
↓
EBS CSI add-on/components missing
↓
No dynamic EBS volume
↓
Pod Pending
And how to investigate it.
PART 16 — Helm
Convert the application into a Helm chart.
Suggested structure:
banking-app/
│
├── Chart.yaml
├── values.yaml
├── values-dev.yaml
├── values-staging.yaml
├── values-prod.yaml
│
└── templates/
├── deployment.yaml
├── service.yaml
├── ingress.yaml
├── configmap.yaml
├── serviceaccount.yaml
├── hpa.yaml
├── pdb.yaml
└── networkpolicy.yaml
Different environment values should demonstrate meaningful differences.
For example:
DEV
replicas: 1
smaller resources
STAGING
replicas: 2
production-like configuration
PRODUCTION
replicas: 3+
HPA
PDB
stronger availability settings
Demonstrate:
helm lint
helm template
helm upgrade --install
helm list
helm history
Explain why Helm is useful in a company with many applications and environments.
PART 17 — Argo CD / GitOps
Put Kubernetes/Helm configuration in Git.
Architecture:
Developer
↓
Pull Request
↓
Review
↓
Merge
↓
Git
↓
Argo CD
↓
EKS
Students must create/configure an Argo CD Application and demonstrate:
Synced
Healthy
Then manually modify something in Kubernetes.
Observe:
OutOfSync
Explain drift.
Then restore/reconcile through GitOps.
Critical production question:
“If Argo CD manages production, should engineers routinely modify production using
kubectl edit?”
Expected:
Generally no. Desired changes should normally go through the controlled Git/change process, with emergency procedures handled according to company policy.
PART 18 — Banking Production Requirements
This section is mandatory in the recording.
Imagine this is a banking system handling financial transactions and sensitive data.
Before saying “production ready,” discuss:
High availability
Least privilege
Encryption
Secret management
Network segmentation
Auditability
Backups
Disaster recovery
Monitoring
Alerting
Logging
Change control
Image scanning
Vulnerability management
Patch management
Capacity planning
Rollback strategy
Business continuity
Compliance requirements
Do not simply say:
“We have three replicas, therefore the bank is production-ready.”
Production readiness is much larger than Kubernetes manifests.
PART 19 — What Costs Money in AWS?
Students must be able to answer this because engineers should understand cost.
Discuss potential charges for things such as:
EKS cluster/control-plane service
EC2 worker nodes
EBS volumes/snapshots
Load balancers
NAT Gateway
Data transfer
CloudWatch logs/metrics
ECR storage/data transfer
Public IPv4 where applicable
Route 53
Backup/storage services
Other managed AWS services
Do not memorize exact prices because pricing changes.
Instead explain:
“Which architectural decisions create cost?”
For example:
More nodes → more compute cost
More/larger EBS → more storage cost
Multiple ALBs → load-balancer cost
Large log ingestion → observability cost
NAT traffic → networking cost
PART 20 — Production Monitoring
Your application must have an observability plan.
Explain what you would monitor:
CPU
Memory
Pod restarts
Ready replicas
Node health
Request rate
Error rate
Latency
HTTP 5xx
HPA behavior
PVC/storage
Database health
Dependency health
And explain:
Metrics
Logs
Traces
A strong student should mention service-level signals such as:
Traffic
Errors
Latency
Saturation
not only Kubernetes CPU.
PART 21 — TEN REQUIRED PRODUCTION INCIDENTS
This is the most important part of the homework.
Each student must break the environment intentionally, investigate it without jumping directly to the answer, fix/mitigate it, and verify recovery on video.
Incident 1 — Bad Deployment / ImagePullBackOff
Break:
Deploy an invalid/nonexistent image tag.
Student must discover:
New rollout
↓
New Pods unhealthy
↓
ImagePullBackOff
↓
describe Pod
↓
Events
↓
Image cannot be pulled
Mitigate:
kubectl rollout history deployment/<name> -n production
kubectl rollout undo deployment/<name> -n production
kubectl rollout status deployment/<name> -n production
Interview explanation:
“A new release introduced an invalid image reference. I correlated the incident with the rollout, confirmed the image pull failure from Pod events, rolled back to the known-good revision, and verified service recovery.”
Incident 2 — Production Slow / HPA at Maximum
Generate load.
Students investigate:
Application slow
↓
Pods Running
↓
CPU/resource pressure
↓
HPA target exceeded
↓
maxReplicas reached
↓
capacity insufficient for current load
They must not say “traffic” without evidence.
Incident 3 — Service Selector Mismatch
Break Service selector:
Service:
app=wrong-app
Pods:
app=banking-app
Symptoms:
Pods 1/1 Running
Service exists
Application unreachable
No expected Service backend
Student must compare:
kubectl describe svc
kubectl get pods --show-labels
kubectl get endpointslice
Incident 4 — Readiness Probe Failure
Break:
readinessProbe:
httpGet:
path: /does-not-exist
Expected:
Running 0/1
Student must explain:
Running ≠ Ready.
Use describe to find probe failure.
Fix/rollback and verify readiness and Service routing.
Incident 5 — OOMKilled
Configure a workload with insufficient memory or a controlled memory-hungry training application.
Student must find:
RESTARTS ↑
↓
CrashLoopBackOff
↓
describe
↓
Last State
↓
OOMKilled
Then:
kubectl logs <pod> --previous
Explain why simply increasing memory may only be mitigation if the application has unbounded memory growth.
Incident 6 — Dependency Failure
Application logs:
Cannot connect to database/Kafka
Student is not allowed to say:
“Kafka is down.”
until proven.
Investigate:
DNS
Service
EndpointSlice
Port
NetworkPolicy
Authentication
Dependency health
Application configuration
Interview statement:
“The log proves the application cannot connect to the dependency; it does not by itself prove the dependency is down.”
Incident 7 — PVC Pending
Break storage provisioning.
Expected:
Pod Pending
↓
PVC Pending
↓
Events
↓
StorageClass
↓
CSI
↓
Cloud permissions/storage
Students must walk the entire storage dependency chain.
Incident 8 — RBAC Forbidden
Make a ServiceAccount attempt:
kubectl delete pod ...
Expected:
Forbidden
Student investigates:
kubectl auth can-i ...
Then Role/RoleBinding.
They must explain:
Authentication succeeded
but
Authorization denied
A Forbidden response is different from a network failure.
Incident 9 — NetworkPolicy Blocks Traffic
Create:
frontend → payment ✗
because NetworkPolicy blocks the connection.
Students must prove:
Pods Running
Service exists
Endpoints exist
DNS may work
BUT network connection blocked
Then inspect:
kubectl get networkpolicy
kubectl describe networkpolicy
They must explain why restarting the application doesn't fix the policy.
Incident 10 — Argo CD OutOfSync / Configuration Drift
Git says:
replicas: 3
Manually change live cluster:
kubectl scale deployment banking-app \
--replicas=1 \
-n production
Observe Argo CD.
Students explain:
Git desired state
≠
Live state
→ OutOfSync
Restore desired state through the GitOps workflow/reconciliation.
Explain why uncontrolled manual production changes are dangerous.
PART 22 — For Every Incident, Use the Same Professional Method
Students may not begin with random commands.
Every recording must follow:
1. ALERT / TICKET
2. CUSTOMER IMPACT
What is broken?
3. SCOPE
Which application/service/environment?
4. WHAT CHANGED?
Deployment?
Config?
Traffic?
Infrastructure?
5. OBSERVE
Metrics
Logs
Events
Traces
6. HYPOTHESIS
What might explain the evidence?
7. PROVE / DISPROVE
Collect evidence.
8. ROOT CAUSE
What actually caused the problem?
9. MITIGATION / FIX
Restore service safely.
10. VERIFY
Prove service is healthy.
11. RCA
Why did this happen?
How do we prevent recurrence?
The sentence I want every student to learn is:
“I don't want to guess the root cause. I form a hypothesis and validate it with evidence.”
PART 23 — Production Verification
After every fix, students must prove recovery.
They cannot say:
“The command succeeded, so production is fixed.”
They must check the relevant layers:
Deployment healthy
↓
Pods Ready
↓
Restart count stable
↓
Service backend healthy
↓
Application responds
↓
Errors normal
↓
Latency normal
↓
Critical customer flow works
PART 24 — Git Repository Requirements
Repository should look professional.
For example:
banking-platform/
│
├── README.md
│
├── docs/
│ ├── architecture.md
│ ├── troubleshooting.md
│ └── incidents.md
│
├── helm/
│ └── banking-app/
│ ├── Chart.yaml
│ ├── values.yaml
│ ├── values-dev.yaml
│ ├── values-staging.yaml
│ ├── values-prod.yaml
│ └── templates/
│
├── argocd/
│ ├── dev.yaml
│ ├── staging.yaml
│ └── production.yaml
│
└── kubernetes/
├── rbac/
├── network-policy/
├── storage/
└── troubleshooting/
README should explain architecture, environments, deployment process, security controls, availability, scaling, storage, GitOps and troubleshooting.
PART 25 — Required Video Recording
I would require approximately 45–60 minutes, but quality matters more than exact duration.
Tell students:
Do not read definitions from Google or from your notes. Share your screen and explain your own environment as though I am the interviewer.
The recording must show and explain the full lifecycle:
Architecture
↓
AWS / EKS
↓
Dev / Staging / Production
↓
Docker Image / ECR
↓
Deployment / ReplicaSet / Pods
↓
Resources / Probes
↓
Service / EndpointSlice
↓
Ingress
↓
ConfigMap / Secret
↓
RBAC
↓
NetworkPolicy
↓
HPA / PDB
↓
Scheduling
↓
StatefulSet / PV / PVC / StorageClass / CSI
↓
Helm
↓
Argo CD / GitOps
↓
Observability
↓
10 Production Incidents
↓
RCA / Verification
At random points in the recording they should be able to answer:
Why did you configure it this way?
That is more important than merely showing YAML.
Top comments (0)