DEV Community

Aisalkyn Aidarova
Aisalkyn Aidarova

Posted on

DevOps / SRE Interview Master Guide

The memory method

For most technical questions, remember D-E-R:

D — Definition: What is it?

E — Example: Where did you use it?

R — Real production detail: How did you troubleshoot/manage it?

For troubleshooting questions, remember C-L-I-F:

C — Check symptoms/events

L — Logs and metrics

I — Identify root cause

F — Fix + verify

For behavioral questions, remember STAR:

Situation → Task → Action → Result


PART 1 — RECRUITER / HR SCREEN

1. Tell me about yourself.

Easy memory: Present → Tools → Responsibilities → Goal.

Answer:

“I’m a DevOps Engineer with several years of experience working with AWS and cloud-native infrastructure.

My main technologies are AWS, Kubernetes/EKS, Docker, Terraform, GitHub Actions, Jenkins, Argo CD, Helm, Linux, and monitoring tools such as Prometheus, Grafana, Datadog, and CloudWatch.

In my current environment, I automate infrastructure with Terraform, build CI/CD pipelines, manage Kubernetes deployments, implement GitOps with Argo CD, and troubleshoot production issues.

I work closely with developers and other engineering teams to make deployments reliable, repeatable, secure, and automated.

I'm looking for a DevOps/SRE opportunity where I can continue working with cloud infrastructure, Kubernetes, automation, and reliability.”


2. What do you do in your current role?

Memory: Infrastructure → CI/CD → Kubernetes → Monitoring → Support.

Answer:

“My responsibilities can be divided into five areas.

First, infrastructure: I provision AWS resources using Terraform.

Second, CI/CD: I maintain pipelines using GitHub Actions and Jenkins.

Third, Kubernetes: I deploy and support applications running on EKS.

Fourth, monitoring: I use tools such as CloudWatch, Prometheus, Grafana, and Datadog.

Finally, production support: I troubleshoot failed deployments, application issues, infrastructure problems, and alerts.”


3. Why are you looking for another opportunity?

“I've learned a lot in my current environment, but I'm looking for an opportunity where I can continue growing in cloud infrastructure, Kubernetes, automation, and SRE practices while taking on larger technical responsibilities.”

Avoid complaining about managers or coworkers.


4. Why DevOps?

“I enjoy automation and solving infrastructure problems. DevOps allows me to work across development, infrastructure, cloud, security, CI/CD, and operations. I especially enjoy taking a manual process and turning it into a reliable automated workflow.”


5. What are your strongest technologies?

“My strongest areas are AWS, Kubernetes/EKS, Terraform, Docker, CI/CD, Linux, and GitOps.

I've used GitHub Actions and Jenkins for CI/CD, Terraform for infrastructure automation, Argo CD and Helm for Kubernetes deployments, and CloudWatch, Prometheus, Grafana, and Datadog for observability.”


6. What is your biggest strength?

“One of my strongest skills is troubleshooting. When an incident happens, I try not to randomly change things. I first understand the symptoms, check logs, events and metrics, isolate the affected layer, identify the root cause, implement the fix, and verify the system is stable.”


7. What is your weakness?

“I used to spend too much time trying to solve every problem independently. I've learned that in production environments, escalation and collaboration are important. I now investigate first, document what I've found, and involve the appropriate team when necessary.”


8. Are you comfortable with on-call?

“Yes. I understand that SRE and DevOps positions can include production support and on-call responsibilities. During an incident my priority is restoring service safely, communicating clearly, and then completing root-cause analysis afterward.”


9. What salary are you looking for?

“I’m open to discussing compensation based on the responsibilities, level, overall package, and expectations for the position. Could you share the budgeted range for the role?”


10. Why should we hire you?

“My experience aligns well with cloud-native DevOps environments. I have hands-on experience with AWS, Terraform, Kubernetes, Docker, CI/CD, GitOps, monitoring, and production troubleshooting.

I can work across the entire delivery lifecycle—from infrastructure provisioning to application deployment and production support.”


PART 2 — DEVOPS FUNDAMENTALS

11. What is DevOps?

Remember: People + Process + Automation.

“DevOps is a combination of culture, practices, and automation that improves collaboration between development and operations teams.

The goal is to deliver software faster and more reliably through automation, CI/CD, Infrastructure as Code, monitoring, and shared responsibility.”


12. What does a DevOps engineer do?

“A DevOps engineer automates infrastructure and software delivery.

Typical responsibilities include cloud infrastructure, CI/CD pipelines, containers, Kubernetes, Infrastructure as Code, monitoring, security automation, and production troubleshooting.”


13. What is CI/CD?

CI = test. CD = deliver/deploy.

“Continuous Integration means developers frequently integrate code and automated pipelines validate it through builds, linting, tests, quality checks, and security scans.

Continuous Delivery or Deployment handles getting validated software into environments such as staging and production.”


14. Describe your complete CI/CD process.

MEMORIZE THIS FLOW:

Developer → Git → CI → Test → Scan → Build → ECR → GitOps → Argo CD → EKS → Monitor

Answer:

“A developer pushes code or opens a pull request.

That triggers GitHub Actions.

The pipeline checks out the repository, installs dependencies, runs linting and tests, performs SonarQube analysis and Trivy security scanning.

After the checks pass, we build a Docker image.

We tag the image using an immutable identifier such as the commit SHA and push it to Amazon ECR.

For deployment we use GitOps. Deployment configuration is updated in Git, and Argo CD detects the desired-state change.

Argo CD synchronizes the manifests with EKS.

Kubernetes performs the rollout and readiness probes determine when new pods can receive traffic.

After deployment we monitor logs, metrics, application health, and alerts.”

This answer alone can lead to 20+ follow-up questions.


PART 3 — LINUX

15. What Linux commands do you use every day?

“I commonly use:

ls, cd, pwd, cat, less, grep, find, tail, head, ps, top, df, du, free, curl, ss, chmod, chown, systemctl, journalctl, ssh, scp, kill and tar.”


16. How do you check disk space?

df -h
Enter fullscreen mode Exit fullscreen mode

For directories:

du -sh *
Enter fullscreen mode Exit fullscreen mode

17. How do you check memory?

free -h
Enter fullscreen mode Exit fullscreen mode

And:

top
Enter fullscreen mode Exit fullscreen mode

18. How do you check processes?

ps aux
top
Enter fullscreen mode Exit fullscreen mode

Find specific process:

ps aux | grep nginx
Enter fullscreen mode Exit fullscreen mode

19. How do you check ports?

ss -tulpn
Enter fullscreen mode Exit fullscreen mode

20. How do you check service status?

systemctl status nginx
Enter fullscreen mode Exit fullscreen mode

Logs:

journalctl -u nginx
Enter fullscreen mode Exit fullscreen mode

21. How do you find errors in a log?

grep -i error application.log
Enter fullscreen mode Exit fullscreen mode

Live logs:

tail -f application.log
Enter fullscreen mode Exit fullscreen mode

22. Difference between kill and kill -9?

“Normal kill sends SIGTERM and gives the process an opportunity to shut down gracefully.

kill -9 sends SIGKILL and immediately terminates the process.

I prefer graceful termination first and use SIGKILL only when necessary.”


23. What is chmod 755?

Owner:

7 = read + write + execute

Group:

5 = read + execute

Others:

5 = read + execute


24. Hard link vs symbolic link?

“A hard link references the same inode/data.

A symbolic link stores a path pointing to another file and can cross filesystems.”


25. Server CPU is 100%. What do you do?

CLIF.

“I first confirm the CPU utilization and identify the processes consuming CPU using top or ps.

Then I check application and system logs, recent deployments, traffic changes and monitoring metrics.

I determine whether the problem is application code, abnormal traffic, a runaway process, or insufficient capacity.

Then I mitigate the immediate impact and investigate the root cause.”


PART 4 — GIT / GITHUB

26. What is Git?

“Git is a distributed version-control system used to track source-code changes and collaborate safely.”


27. Git vs GitHub?

“Git is the version-control technology.

GitHub is a platform that hosts Git repositories and provides features such as pull requests, Actions, reviews, branch protection, and collaboration.”


28. Describe your Git workflow.

“Normally I pull the latest main branch, create a feature branch, make changes, commit them, push the branch, and open a pull request.

The PR triggers CI checks and code review.

Once approvals and required checks pass, the PR is merged according to the team's branching policy.”


29. git fetch vs git pull?

“git fetch downloads remote changes without modifying my working branch.

git pull fetches changes and then integrates them into my current branch.”


30. Merge vs rebase?

“Merge combines histories and normally creates a merge commit.

Rebase moves commits onto another base and creates a cleaner linear history.

I follow the team's Git policy, particularly because rebasing already-shared history can cause problems.”


31. What is a merge conflict?

“A merge conflict happens when Git cannot automatically determine how to combine competing changes.

I inspect the conflicting files, determine the intended final content, resolve the conflict, test the result, stage it and continue the merge or rebase.”


32. git revert vs git reset?

“git revert creates a new commit that reverses a previous commit and is safer for shared branches.

git reset moves the branch pointer and can rewrite history, so I use it carefully.”


PART 5 — DOCKER

33. What is Docker?

“Docker packages an application and its dependencies into a portable container image so it can run consistently across environments.”


34. Image vs container?

Image = blueprint. Container = running instance.

“An image is an immutable package containing the application and dependencies.

A container is a running instance of an image.”


35. What is Dockerfile?

“A Dockerfile contains instructions for building a Docker image.”

Example:

FROM nginx:alpine
COPY dist /usr/share/nginx/html
EXPOSE 80
Enter fullscreen mode Exit fullscreen mode

36. Build an image.

docker build -t application:v1 .
Enter fullscreen mode Exit fullscreen mode

37. Run container.

docker run -d -p 8080:80 application:v1
Enter fullscreen mode Exit fullscreen mode

38. CMD vs ENTRYPOINT?

“Both define container startup behavior.

ENTRYPOINT normally defines the executable.

CMD commonly provides the default command or arguments that can be overridden.”


39. COPY vs ADD?

“COPY simply copies files into the image and is usually preferred.

ADD has additional functionality such as automatic local tar extraction.”


40. Why multi-stage builds?

“To reduce final image size and attack surface by separating build dependencies from runtime dependencies.”


41. Container keeps crashing. What do you check?

docker ps -a
docker logs <container>
docker inspect <container>
Enter fullscreen mode Exit fullscreen mode

“I check exit code, logs, environment variables, command/entrypoint, resource constraints, configuration and application dependencies.”


PART 6 — KUBERNETES

42. What is Kubernetes?

“Kubernetes is a container orchestration platform that manages deployment, scaling, networking, availability, and lifecycle of containerized applications.”


43. Explain Kubernetes architecture.

Control Plane:

API Server

Scheduler

Controller Manager

etcd

Worker:

kubelet

container runtime

kube-proxy/networking components

Pods


44. What is a Pod?

“A Pod is Kubernetes' smallest deployable unit. It contains one or more containers that share networking and can share storage.”


45. Pod vs Deployment?

“A Pod runs containers.

A Deployment manages replicated application Pods through ReplicaSets and supports declarative updates, scaling and rollouts.”


46. Deployment vs StatefulSet?

“Deployment is generally used for stateless applications where pod identity isn't important.

StatefulSet is designed for stateful workloads requiring stable pod identities, ordered behavior, and persistent storage relationships.”


47. Deployment vs DaemonSet?

“A Deployment controls a desired number of application replicas.

A DaemonSet normally ensures a pod runs on each eligible node, making it useful for node-level agents such as logging or monitoring components.”


48. What is a Service?

“A Service provides a stable network endpoint for a changing set of Pods selected by labels.”


49. Service types?

ClusterIP — internal access.

NodePort — exposes a port through nodes.

LoadBalancer — integrates with an external load balancer where supported.

ExternalName — DNS mapping to an external hostname.


50. What is Ingress?

“Ingress defines HTTP/HTTPS routing rules into services.

An Ingress Controller actually implements those rules.

In AWS environments, the AWS Load Balancer Controller can create and manage ALB resources based on Kubernetes resources.”


51. ConfigMap vs Secret?

“ConfigMap stores non-sensitive configuration.

Secret is intended for sensitive values.

For production secrets, I prefer integration with an appropriate secret-management system rather than treating base64-encoded Kubernetes Secrets alone as sufficient protection.”


52. What is Namespace?

“A namespace provides logical isolation and organization of Kubernetes resources.”


53. What is HPA?

“HPA automatically changes replica counts based on observed metrics such as CPU, memory, or supported custom/external metrics.”


54. What is PDB?

“A PodDisruptionBudget limits how many replicas can be voluntarily disrupted simultaneously, helping preserve application availability during operations such as node maintenance.”


55. HPA vs PDB?

Easy memory:

HPA = how many Pods.

PDB = how many can safely disappear during voluntary disruption.


56. Readiness vs liveness?

Readiness = traffic.

Liveness = restart.

“Readiness determines whether a pod should receive traffic.

Liveness determines whether Kubernetes should restart an unhealthy container.”


57. What is a startup probe?

“It protects slow-starting applications by delaying liveness/readiness behavior until startup has successfully completed.”


58. Pod is Pending. What do you do?

kubectl get pods
kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

“I look at Events.

Common causes include insufficient CPU/memory, scheduling constraints, taints/tolerations, affinity rules, unbound PVCs, or unavailable nodes.”


59. What is CrashLoopBackOff?

“It means the container repeatedly starts, fails, and Kubernetes backs off before restarting it again.”

Check:

kubectl logs <pod>
kubectl logs <pod> --previous
kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

60. ImagePullBackOff?

“Kubernetes cannot successfully pull the requested container image.”

Check:

kubectl describe pod <pod>
Enter fullscreen mode Exit fullscreen mode

Then investigate image/tag, registry authentication, IAM, network connectivity, and registry availability.


61. Pod is Running but application doesn't work.

Remember the layers:

Pod → Application → Readiness → Service → Endpoints → Ingress → LB/DNS.

Check:

kubectl logs <pod>
kubectl describe pod <pod>
kubectl get svc
kubectl get endpoints
kubectl get ingress
Enter fullscreen mode Exit fullscreen mode

62. What happens if a Pod dies?

“If it is managed by a controller such as a Deployment, the controller works to restore the desired replica count by creating a replacement Pod.”


63. What if a Kubernetes node dies?

“Pods on that node become unavailable. Kubernetes eventually reschedules controller-managed workloads onto healthy nodes when capacity and scheduling constraints permit.”


64. What is etcd?

“etcd is Kubernetes' distributed key-value store containing cluster state used by the control plane.”


65. What is kubelet?

“Kubelet is the node agent responsible for ensuring containers described by assigned PodSpecs are running on that node.”


66. What does Scheduler do?

“The scheduler selects appropriate nodes for unscheduled Pods based on resources, constraints, affinity, taints/tolerations and other scheduling factors.”


67. What are requests and limits?

“Requests influence scheduling and reserve the resource amount considered necessary.

Limits constrain resource usage.

For CPU, exceeding the limit can cause throttling. Exceeding a memory limit can result in an OOM kill.”


68. What is OOMKilled?

“The container exceeded its available memory limit or encountered a memory-related termination condition.”

I investigate:

kubectl describe pod <pod>
kubectl top pod
Enter fullscreen mode Exit fullscreen mode

Then check memory usage patterns and application behavior before simply increasing resources.


69. What are labels and selectors?

“Labels are metadata attached to Kubernetes resources.

Selectors allow other resources such as Services and Deployments to identify matching resources.”


70. Service has no endpoints. Why?

“First I check whether the Service selector matches the labels on the Pods and whether the Pods are ready.”

kubectl get svc
kubectl get pods --show-labels
kubectl get endpoints
Enter fullscreen mode Exit fullscreen mode

71. How do you rollback Kubernetes deployment?

kubectl rollout history deployment/app
kubectl rollout undo deployment/app
Enter fullscreen mode Exit fullscreen mode

PART 7 — AWS

72. What AWS services have you used?

“EC2, VPC, IAM, EKS, ECR, S3, RDS, Route 53, CloudWatch, ALB, Auto Scaling, Secrets Manager, KMS and related networking/security services.”


73. What is EC2?

“EC2 provides virtual compute instances in AWS.”


74. What is VPC?

“A VPC is a logically isolated network in AWS where we define IP ranges, subnets, routing and connectivity.”


75. Public vs private subnet?

“A public subnet has a route that enables internet access through an Internet Gateway for appropriately configured resources.

Private subnets don't directly route outbound internet traffic through an Internet Gateway. When outbound internet access is required, they can use mechanisms such as a NAT Gateway.”


76. Security Group vs NACL?

SG = instance/resource level + stateful.

NACL = subnet level + stateless.

“Security Groups are stateful and apply to supported network interfaces/resources.

Network ACLs operate at the subnet boundary and are stateless.”


77. What is IAM?

“IAM controls authentication and authorization for AWS resources through identities, roles and policies.”


78. User vs role?

“An IAM user represents a long-lived IAM identity.

A role is assumed to obtain temporary credentials and is preferred for workloads and service-to-service access.”


79. What is least privilege?

“Grant only the permissions necessary to perform the required task and no more.”


80. What is EKS?

“Amazon EKS is AWS's managed Kubernetes service. AWS manages the Kubernetes control plane while customers manage their workloads and, depending on the configuration, compute and associated infrastructure.”


81. ECR?

“Amazon ECR is AWS's managed container registry used to store and distribute container images.”


82. S3?

“S3 is AWS object storage used for objects such as artifacts, logs, backups and Terraform state.”


83. ALB vs NLB?

“ALB operates at Layer 7 and provides HTTP/HTTPS-aware routing capabilities.

NLB operates primarily at Layer 4 and is designed for high-performance TCP/UDP/TLS traffic.”


84. What is Route 53?

“AWS's DNS service. It manages DNS records and provides routing and health-check capabilities.”


85. What is Auto Scaling?

“It dynamically adjusts compute capacity based on configured policies, metrics and desired capacity.”


PART 8 — TERRAFORM

86. What is Terraform?

“Terraform is an Infrastructure as Code tool that allows us to declaratively define and manage infrastructure.”


87. Main Terraform workflow?

Write → Init → Plan → Apply.

terraform init
terraform fmt
terraform validate
terraform plan
terraform apply
Enter fullscreen mode Exit fullscreen mode

88. What is Terraform state?

“Terraform state maps resources in configuration to the infrastructure Terraform manages and stores important resource metadata.”


89. Why remote state?

“For team environments, remote state provides centralized storage and enables safer collaboration.

For AWS, S3 is commonly used for remote state, with appropriate state-locking mechanisms and versioning/protection.”


90. What is a Terraform module?

“A reusable collection of Terraform configuration representing infrastructure components.”

Example:

modules/
  vpc/
  eks/
  rds/
Enter fullscreen mode Exit fullscreen mode

91. Variables vs outputs?

“Variables provide inputs to configuration.

Outputs expose useful values from configuration or modules.”


92. What is terraform plan?

“It calculates and displays the proposed changes without intentionally applying them.”


93. What is Terraform drift?

“Drift occurs when actual infrastructure differs from the configuration/state expectations, often because something was changed outside the normal IaC workflow.”


94. Someone manually changed AWS. What do you do?

“I first determine what changed and why.

I run terraform plan to understand the difference.

Then we decide whether the manual change should be represented in code or reverted so the environment again matches the desired configuration.

I avoid blindly applying Terraform without reviewing the plan.”


95. Terraform state is lost. What do you do?

“I don't immediately run terraform apply.

First I stop changes to the environment.

I check the remote backend, S3 version history/backups and locking configuration.

If a valid state version exists, I restore it carefully.

If state cannot be recovered, existing resources may need to be imported into Terraform.

Then I run terraform plan and verify that Terraform isn't planning to recreate or destroy production infrastructure unexpectedly.”


96. What is terraform import?

“It associates an existing infrastructure resource with a Terraform resource address/state so Terraform can manage it. Configuration still needs to accurately describe the imported resource.”


97. count vs for_each?

“Both create multiple resource instances.

count uses numeric indexes.

for_each uses keys from a map or set and can provide more stable resource identity when managing named objects.”


98. What is depends_on?

“It explicitly defines a dependency when Terraform cannot infer the dependency from resource references.”


PART 9 — JENKINS / GITHUB ACTIONS

99. What is Jenkins?

“Jenkins is an automation server commonly used to implement CI/CD pipelines.”


100. What is GitHub Actions?

“GitHub Actions is GitHub's workflow automation platform. Workflows are defined in YAML and triggered by events such as pushes and pull requests.”


101. Workflow vs job vs step?

Workflow → Jobs → Steps.

“A workflow is the overall automation.

It contains one or more jobs.

Each job contains steps.”


102. Why run jobs in parallel?

“To reduce pipeline duration when jobs don't depend on each other.”

For example:

Lint | Tests | SonarQube | Trivy

can run concurrently when appropriate.


103. How do you store CI secrets?

“I don't hardcode credentials in the repository.

I use the CI platform's secrets mechanism, OIDC/workload identity where supported, and external secret-management solutions depending on the environment.”


104. Pipeline fails. What do you do?

“I identify the failing job and exact step, inspect its logs, determine whether the failure is code, tests, dependencies, credentials, infrastructure or the pipeline itself, reproduce locally when practical, fix the root cause, rerun the affected workflow and verify the result.”


PART 10 — SONARQUBE / TRIVY / SECURITY

105. What is SonarQube?

“SonarQube performs static code analysis and helps identify bugs, vulnerabilities, security hotspots, code smells and maintainability issues.”


106. What is a Quality Gate?

“A Quality Gate defines conditions that code analysis must satisfy before we consider the code acceptable according to the organization's policy.”


107. What is Trivy?

“Trivy is a security scanner that can detect vulnerabilities and certain misconfigurations across container images, filesystems, repositories and IaC depending on how it's used.”


108. SonarQube vs Trivy?

Easy memory:

SonarQube → primarily code quality/static analysis.

Trivy → vulnerabilities/configuration/container/IaC scanning.

“They overlap somewhat in security-related capabilities, but they solve different parts of the software-security workflow.”


109. Trivy finds CRITICAL vulnerability. What happens?

“I don't ignore it just to make CI green.

I identify the affected package and available remediation, assess exploitability and organizational policy, upgrade or replace the vulnerable dependency/base image when possible, rebuild and scan again.

Any exception should follow the organization's documented risk-acceptance process.”


PART 11 — ARGO CD / GITOPS

110. What is GitOps?

“GitOps uses Git as the version-controlled source of desired configuration. Changes go through Git workflows, and an automated controller reconciles the environment toward that desired state.”


111. What is Argo CD?

“Argo CD is a GitOps continuous-delivery controller for Kubernetes. It compares the desired state in Git with the live cluster and can synchronize differences.”


112. Why Argo CD if we already have GitHub Actions?

MEMORIZE:

GitHub Actions builds. Argo CD deploys/reconciles.

“A common separation is:

GitHub Actions handles CI—testing, scanning, building and publishing artifacts.

Argo CD handles Kubernetes CD by reconciling cluster configuration with Git.

That separation keeps deployment state declarative and auditable.”


113. What is OutOfSync?

“It means the live Kubernetes state differs from the desired state Argo CD sees in Git.”


114. What is self-healing?

“When configured, Argo CD can detect certain live-state drift and reconcile the cluster back toward the desired Git state.”


PART 12 — HELM

115. What is Helm?

“Helm is a Kubernetes package manager and templating system used to package, configure, install and upgrade Kubernetes applications.”


116. What is a Helm chart?

“A chart is a package containing Kubernetes templates, default values and chart metadata.”

Typical:

Chart.yaml
values.yaml
templates/
Enter fullscreen mode Exit fullscreen mode

117. Why Helm?

“It reduces duplicated YAML and makes Kubernetes configuration reusable and configurable across environments.”


118. Helm vs Argo CD?

“Helm renders/packages Kubernetes configuration.

Argo CD manages GitOps reconciliation and deployment.

They can work together: Argo CD can deploy an application represented by a Helm chart.”


PART 13 — NETWORKING

119. What is DNS?

“DNS translates names such as app.company.com into addresses/resources clients can connect to.”


120. TCP vs UDP?

“TCP is connection-oriented and provides ordered, reliable delivery.

UDP is connectionless with lower protocol overhead but doesn't provide TCP's delivery guarantees.”


121. HTTP vs HTTPS?

“HTTPS is HTTP protected using TLS, providing encryption and server authentication, and potentially client authentication depending on configuration.”


122. What happens when you type a URL?

Excellent interview question.

Remember: DNS → TCP/TLS → HTTP → LB → App → Response.

“First the client resolves the hostname through DNS.

It establishes network connectivity to the destination.

For HTTPS, a TLS handshake establishes the secure connection.

The client sends an HTTP request.

The request may pass through a load balancer or ingress to the application.

The application processes it, potentially communicating with databases or other services, and returns an HTTP response.”


123. What is a load balancer?

“It distributes traffic across multiple healthy application targets to improve availability and scalability.”


124. 502 vs 503 vs 504?

Easy memory:

502 → bad upstream response.

503 → service unavailable.

504 → upstream timeout.

“But the exact source depends on which proxy/load balancer generated the status code, so I check the relevant logs and metrics rather than assuming the cause.”


PART 14 — MONITORING / OBSERVABILITY

125. What is monitoring?

“Monitoring collects and evaluates system/application signals so teams can understand health and detect problems.”


126. Metrics vs logs vs traces?

Metrics = numbers.

Logs = events/details.

Traces = request journey.


127. What do you monitor in Kubernetes?

“Node CPU/memory/disk, pod health, resource utilization, restart counts, request latency, error rates, application throughput, deployment health and important business/application metrics.”


128. What is Prometheus?

“Prometheus is a monitoring and time-series system commonly used for collecting metrics and evaluating alerting rules.”


129. What is Grafana?

“Grafana provides visualization and dashboards over data sources such as Prometheus and others.”


130. What is CloudWatch?

“AWS's monitoring and observability service for metrics, logs, alarms, dashboards and related AWS telemetry.”


131. What is Datadog?

“Datadog is an observability platform providing capabilities around infrastructure monitoring, metrics, logs, traces, dashboards and alerting.”


PART 15 — SRE QUESTIONS

132. What is SRE?

“Site Reliability Engineering applies software-engineering approaches to operations and reliability.”


133. What is SLI?

“SLI is the actual measured reliability indicator.”

Example:

99.95% successful requests.


134. What is SLO?

“SLO is the reliability target for an SLI.”

Example:

99.9% successful requests over the defined period.


135. What is SLA?

“An SLA is a formal service agreement that can define service commitments and consequences when commitments aren't met.”


136. What is error budget?

“An error budget represents the amount of unreliability allowed by an SLO.

For a 99.9% availability objective, the remaining 0.1% represents the basic availability error budget over the measurement period.”


137. What is toil?

“Toil is repetitive operational work that is manual, automatable, tactical and scales poorly.

An SRE objective is to reduce unnecessary toil through engineering and automation.”


PART 16 — PRODUCTION INCIDENTS

138. Production is down. What do you do?

THIS IS CRITICAL.

Remember:

Confirm → Impact → Changes → Logs/Metrics → Mitigate → Fix → Verify → RCA

Answer:

“First I confirm the incident and determine customer impact.

I check monitoring, alerts, logs and metrics.

I look for recent deployments or infrastructure/configuration changes.

My first operational priority is restoring service safely. If a recent deployment caused the incident, rollback may be the fastest mitigation.

Once service is stable, we identify and fix the root cause, verify recovery, communicate status and document the incident.

Afterward we conduct a blameless postmortem and implement preventive actions.”


139. Deployment caused 500 errors.

“I check application metrics and logs and correlate the start of errors with the deployment.

If customer impact is significant and the previous release is known-good, I prioritize safe rollback.

After restoring service, I investigate the new release offline and identify the root cause.”


140. Deployment caused 504 errors.

“I investigate the complete request path:

Load balancer → Ingress → Service → Pod → downstream dependencies.

I check target health, readiness, endpoints, application logs, response latency, database/network dependencies and recent configuration changes.

If the release introduced the problem, rollback may be appropriate while we investigate.”


141. Database connection suddenly fails.

“I check application errors, database availability, network connectivity, DNS, security rules, credentials/secrets, connection pool behavior and recent changes.

I avoid assuming the database itself is down until I identify which layer is failing.”


142. Application suddenly becomes slow.

“First I establish whether latency affects all requests or a specific endpoint.

Then I correlate latency with CPU, memory, traffic, error rate, database latency and downstream-service metrics.

Distributed tracing can help identify which component is consuming the time.

Then I mitigate the bottleneck and investigate its root cause.”


143. Disk is full.

df -h
du -sh /*
Enter fullscreen mode Exit fullscreen mode

“I identify which filesystem and directories are consuming space, check logs and temporary/application files, and determine why growth occurred.

I don't blindly delete files from production.

After safely recovering capacity, I fix retention, log rotation, storage sizing, or the application behavior that caused it.”


PART 17 — SCENARIO QUESTIONS

144. Developer says: “It works on my machine.”

“I compare environmental differences: image version, environment variables, dependencies, configuration, secrets, networking and runtime versions.

Containers and reproducible build processes help reduce these inconsistencies.”


145. CI passed but production failed. Why?

“CI success doesn't guarantee production success.

Potential differences include production configuration, secrets, traffic patterns, infrastructure, databases, network policies, resource limits and external dependencies.

I investigate production telemetry and compare the deployed artifact and configuration against what CI validated.”


146. Application works directly through Pod IP but not Service.

“I check Service selectors, Pod labels, endpoints, ports, targetPort, readiness and network policies.”


147. Service works but domain doesn't.

“I move outward:

Service → Ingress → Load Balancer → DNS → TLS.

I verify ingress configuration, load-balancer health/listeners, DNS records and certificate/TLS configuration.”


148. One Pod works, another doesn't.

“I compare the failing Pod with a healthy Pod:

image digest/tag, node, environment, mounted configuration/secrets, logs, events, resource usage and network connectivity.

That comparison usually narrows the problem quickly.”


PART 18 — SECURITY QUESTIONS

149. How do you secure a CI/CD pipeline?

“I use least-privilege permissions, protected branches, code reviews, short-lived credentials/OIDC when available, secret management, dependency and image scanning, SAST, artifact integrity controls, restricted runners and auditable deployment processes.”


150. How do you secure Kubernetes?

“RBAC and least privilege, workload identity, namespace and network isolation, Pod Security controls, image scanning, secret management, resource limits, admission policies where appropriate, audit logging and regular patching/upgrades.”


151. How do you manage secrets?

“I never intentionally store plaintext production secrets in Git.

Depending on the platform, I use solutions such as AWS Secrets Manager and appropriate Kubernetes integrations, with IAM/RBAC controls and rotation policies.”


152. Secret was committed to Git. What do you do?

“First I treat it as compromised.

I revoke or rotate the credential immediately.

Then I determine its exposure and review relevant audit logs.

We remove it from the repository/history as appropriate, but deleting the Git line alone isn't enough because the credential may already have been copied.

Finally, I add controls to prevent recurrence.”


PART 19 — BEHAVIORAL QUESTIONS

Students should use STAR.

153. Tell me about a production incident you handled.

Situation: A production application started returning errors after deployment.

Task: Restore service quickly and determine the cause.

Action:

“I checked monitoring and application logs and correlated the incident with the new deployment. We rolled back to the previous stable release, which restored service. Further investigation identified an incorrect environment configuration affecting the application's database connectivity.”

Result:

“Service was restored, we corrected the configuration, improved readiness validation and documented the incident in a postmortem.”


154. Tell me about a disagreement with a developer.

“I focus on technical evidence rather than personalities.

For example, if a developer believes infrastructure caused an issue, I compare logs, metrics, deployment history and configuration with them.

The goal isn't to prove who is wrong. The goal is to identify the failing layer and restore the service.”


155. Tell me about a mistake you made.

“I explain the mistake clearly, take responsibility for my part, explain how I corrected it and focus heavily on what process or automation I introduced so the same type of mistake is less likely to happen again.”

Students should prepare a genuine example rather than invent one.


156. How do you prioritize multiple incidents?

“I prioritize based on customer/business impact, severity, scope, security risk and service dependencies.

A production outage affecting customers takes priority over a lower-severity development issue.”


157. How do you work under pressure?

“I rely on structured troubleshooting. I confirm impact, use logs and metrics, make one controlled change at a time, communicate clearly and prioritize restoring service over trying experimental fixes in production.”


PART 20 — SENIOR DEVOPS QUESTIONS

158. How would you design a highly available application in AWS?

“I would distribute workloads across multiple Availability Zones.

A typical architecture could use Route 53, an ALB, application workloads across multiple AZs, autoscaling, a highly available database architecture, caching where appropriate, durable object storage, monitoring and automated recovery.

For Kubernetes, EKS worker capacity should span multiple AZs, applications should have multiple replicas, topology-aware scheduling where appropriate, health probes, PDBs and autoscaling.

The exact architecture depends on the application's availability and recovery requirements.”


159. How do you achieve zero/minimal-downtime deployments?

“Multiple replicas, rolling updates, readiness probes, appropriate termination handling, graceful shutdown, capacity during rollout, PDBs where appropriate and backward-compatible application/database changes.

For higher-risk releases we can also use strategies such as canary or blue/green deployments.”


160. Blue/green vs canary?

Blue/Green: two environments/releases; switch traffic.

Canary: gradually expose a new version to a subset of traffic/users.


161. How would you reduce AWS costs?

“I start with measurements rather than arbitrary downsizing.

I review utilization and right-size compute, use appropriate autoscaling, identify idle resources, optimize storage, lifecycle old data, evaluate Savings Plans/Reserved capacity for predictable workloads, consider Spot for suitable fault-tolerant workloads, and analyze network/data-transfer costs.

Then I verify that cost optimizations don't violate reliability requirements.”


162. How do you design disaster recovery?

“I start with business requirements, particularly RTO and RPO.

Then I design backups, replication, infrastructure recovery through IaC, cross-region or multi-region strategies where justified, documented recovery procedures and—very importantly—regular recovery testing.”


163. RTO vs RPO?

RTO = How quickly must we recover?

RPO = How much data can we afford to lose?


PART 21 — RAPID-FIRE QUESTIONS

Students should answer these in one or two sentences.

164. What is YAML?

A human-readable data serialization format heavily used for configuration.

165. What is JSON?

A structured data-interchange format using objects, arrays and primitive values.

166. What is API?

A defined interface that allows software components to communicate.

167. What is REST?

An architectural style for networked APIs commonly using HTTP resources and methods.

168. GET vs POST?

GET normally retrieves a resource. POST commonly submits data or requests creation/processing.

169. What is SSH?

A secure protocol commonly used for remote shell access and tunneling.

170. What is TLS?

A protocol providing encrypted and authenticated communications.

171. What is port 22?

Common default SSH port.

172. Port 80?

HTTP.

173. Port 443?

HTTPS.

174. What is cron?

A Unix/Linux facility for scheduling recurring commands/jobs.

175. What is Bash?

A Unix shell and scripting language commonly used for automation.

176. What is a reverse proxy?

A server/proxy that accepts client requests and forwards them to backend services.

177. Horizontal vs vertical scaling?

Horizontal = add more instances.

Vertical = increase resources of an instance.

178. Authentication vs authorization?

Authentication = who are you?

Authorization = what can you do?

179. Encryption at rest vs in transit?

At rest protects stored data.

In transit protects data moving across networks.


PART 22 — INTERVIEWER WILL DIG DEEPER

This is where students often get caught.

If student says:

“I use Kubernetes.”

Interviewer asks:

“What happens internally when you run kubectl apply?”

Strong answer:

“kubectl sends the desired resource definition to the Kubernetes API server.

The API server authenticates and authorizes the request, validates it and persists the desired state through the control plane.

Controllers observe the desired state and reconcile resources.

If new Pods need scheduling, the scheduler selects nodes.

Kubelets on those nodes observe their assigned workloads and work with the container runtime to start containers.

Other controllers and networking components establish related resources as required.”


If student says:

“I use Terraform.”

Interviewer:

“Explain what happens when you run terraform apply.”

“Terraform loads configuration and provider information, evaluates dependencies and current state, refreshes/reads necessary infrastructure information, determines required changes and invokes provider APIs to create, update or delete resources according to the approved plan.”


If student says:

“I use Docker.”

Interviewer:

“Why not just deploy the source code?”

“Containers package the application with its runtime dependencies into a versioned artifact. That improves consistency between CI, testing and production and makes releases easier to reproduce and roll back.”


PART 23 — THE MOST IMPORTANT TROUBLESHOOTING TREE

Have students memorize this.

Website is down

Start from the outside and move inward:

DNS

↓

Load Balancer

↓

Ingress

↓

Service

↓

Endpoints

↓

Pod

↓

Container

↓

Application

↓

Database / dependencies

Useful commands:

nslookup app.company.com

curl -v https://app.company.com

kubectl get ingress

kubectl describe ingress <name>

kubectl get svc

kubectl get endpoints

kubectl get pods

kubectl describe pod <pod>

kubectl logs <pod>

kubectl logs <pod> --previous

kubectl top pod

kubectl get events --sort-by=.metadata.creationTimestamp
Enter fullscreen mode Exit fullscreen mode

This is much stronger than randomly restarting Pods.


PART 24 — QUESTIONS STUDENTS SHOULD ASK THE INTERVIEWER

Students should always have questions.

Good examples:

“What does your current CI/CD architecture look like?”

“How is Kubernetes used within the organization?”

“How do you manage Infrastructure as Code?”

“What monitoring and observability platforms does the team use?”

“How does the team handle production incidents and on-call?”

“What would you expect the person in this role to accomplish during the first 90 days?”

“What are the biggest reliability or infrastructure challenges the team is currently working on?”


PART 25 — THE 10 ANSWERS EVERY STUDENT MUST KNOW

If they cannot memorize everything immediately, start here:

1. Tell me about yourself

2. Explain your current project

3. Explain your CI/CD pipeline

4. Explain Kubernetes architecture

5. Troubleshoot a failing Pod

6. Explain Terraform and state

7. Explain AWS VPC/networking

8. Explain Docker

9. Explain a production incident

10. Explain monitoring and troubleshooting

Top comments (0)