the misunderstanding
pod autoscaling is different from node autoscaling. they both need different configurations and they solve different problems.
the imagination
to walkthrough this scenario, imagine a restaurant. for the restaurant to operate, we have 3 components. we have customers, chefs, and the restaurant itself.
each customer orders 1 food, and each chef can handle 5 orders. at a given moment, a restaurant can fit a maximum of 3 chefs because the kitchen only has 3 workstations.
the number of orders from customers refers to how much traffic we have, a chef is a pod, and the restaurant is a node.
now imagine we start with only 1 chef.
customers start coming in and the chef becomes busier. instead of waiting until the chef reaches the maximum 5 orders, the business owner sets an alarm. when the chef reaches around 3 out of 5 orders, the alarm goes off and the business owner hires another chef to help with the orders.
this is similar to HPA. when the pod reaches the CPU utilization target, HPA will increase the number of pods.
as more customers come in, the same thing happens again. eventually, we have 3 chefs working inside the restaurant.
now we have another problem.
the restaurant only has 3 workstations, and all 3 are already occupied. if more customers come in and the alarm says we need another chef, the business owner can hire the chef, but there is nowhere for the new chef to work.
this is where node autoscaling comes in.
because the new chef cannot find an available workstation, the business owner opens another restaurant with another 3 workstations. the new chef can then work inside the new restaurant.
so the important part is, more customers do not directly create another restaurant.
more customers increase the workload, the workload causes HPA to create more pods, and when there is no more space to run those pods, node autoscaling creates another node.
that is how pod and node autoscaling work together.

Top comments (0)