DEV Community

Cover image for Who Owns Kubernetes on Day 2? Defining the Operational Boundary
Kubernetes with Naveen
Kubernetes with Naveen

Posted on

Who Owns Kubernetes on Day 2? Defining the Operational Boundary

Day 1 of a Kubernetes deployment feels like a victory. Day 2 is where the organizational cracks show up.

The cluster is provisioned. Nodes are healthy. Ingress is routing traffic. CoreDNS is responding. Monitoring is green. The application team gets a namespace, deploys its workloads, and everyone moves on to the next project.

Then reality arrives.

Spotify

A deployment starts getting OOMKilled. A rollout hangs because of a restrictive PodDisruptionBudget. Someone changes a production manifest directly with kubectl. Three weeks later, the Git repository says one thing while the cluster is running something else. A node pool needs an upgrade, but nobody wants to touch it because six teams have workloads running there.

Then comes the inevitable question: Who owns this?

That question sounds simple, but it exposes one of the most important design decisions in Kubernetes: the operational boundary between the platform team and the application team. If that boundary is too loose, Kubernetes becomes the Wild West. Developers get excessive privileges, production state changes outside Git, and platform engineers spend their days cleaning up application mistakes.

If that boundary is too restrictive, Kubernetes becomes Ticket-Ops. Developers cannot change a ConfigMap without filing a Jira ticket, cannot restart their own workload, and cannot inspect production logs without asking the platform team.

Neither model scales.

The goal is not to decide who owns Kubernetes as a single system. The goal is to decide who owns each operational responsibility after the cluster becomes a production platform.

Twitter


The Day 2 Reality Check: Why Healthy Clusters Still Host Broken Applications

A Kubernetes cluster can be completely healthy while the application running inside it is completely broken. That distinction sounds obvious, but many organizations blur it.

Suppose every node reports Ready, the Kubernetes API server is responsive, the CNI is functioning, CoreDNS is healthy, and the ingress controller has no errors. From a platform perspective, the cluster looks excellent. Now imagine an application deployment has a memory limit of 256Mi, but its actual working set regularly reaches 400Mi.

The container gets terminated with:

OOMKilled
Enter fullscreen mode Exit fullscreen mode

Kubernetes is doing exactly what it was configured to do. The platform team did not break the application. The application team configured a workload that cannot operate within its declared resource boundary.

The same distinction appears in networking. The cluster network can be healthy while an application has an incorrect Service selector. The ingress controller can be perfectly operational while an application has configured the wrong backend port. The scheduler can be functioning correctly while a workload cannot be placed because its resource requests are unrealistic. This is why Day 2 operations need explicit ownership boundaries.

A useful mental model is:

The platform team owns the environment in which workloads run. Application teams own the workloads themselves.

That sounds simple until you start defining the edges.

  • Who owns RBAC?
  • Who owns secrets?
  • Who owns resource quotas?
  • Who owns node upgrades?
  • Who owns a failed deployment?
  • Who owns configuration drift?

The answers need to be designed before the first serious incident, not negotiated during one.


The Core Philosophy: Platform as a Product vs. Infrastructure Monopoly

The strongest Kubernetes operating models treat the internal platform as a product. The platform team is not simply the group that manages Kubernetes. It provides a product consumed by engineering teams.

That product includes the Kubernetes API, namespaces, networking capabilities, identity integration, observability, deployment mechanisms, security policies, storage primitives, upgrade processes, and paved roads for deploying applications.

The platform team owns the API and the paved road. Application teams own what they run on top of that road.

This distinction is important because platform engineering is not about taking operational responsibility away from developers. It is about removing unnecessary infrastructure complexity while preserving application ownership.

A good platform might provide an application team with:

Application Repository
        |
        v
GitOps Configuration
        |
        v
Deployment Pipeline
        |
        v
Kubernetes Namespace
        |
        +---- Service
        +---- Deployment
        +---- Ingress
        +---- ConfigMap
        +---- HPA
        +---- Resource Requests/Limits
Enter fullscreen mode Exit fullscreen mode

The platform team controls the guardrails around this system. The application team controls the application-specific configuration within those guardrails. That is Platform as a Product.

The platform should answer:

"How do we safely run workloads?"

The application team should answer:

"What workload should we run, and how should it behave?"

When those two questions become mixed together, operational ownership becomes unclear.


The Responsibility Breakdown: The Operational Matrix

There is no universal ownership matrix that fits every company, but the following model works well for organizations operating shared Kubernetes platforms.

Operational Domain Platform / SRE Application Team Shared Boundary
Control plane Own Consume Platform health
Node pools Own Validate workloads Upgrade coordination
Kubernetes upgrades Own Test applications Compatibility
Namespace creation Own / automate Request/consume Standard templates
RBAC platform roles Own Consume Least privilege
Application RBAC Guardrails Own Policy
Application deployment Provide platform Own GitOps
ConfigMaps Guardrails Own GitOps
Secrets Secret platform Own application secrets Security policy
Resource quotas Set boundaries Operate within them Capacity planning
CPU/memory requests Validate/policy Own Platform defaults
Application health Infrastructure signals Own Observability
Configuration drift GitOps platform Own desired state Reconciliation
Cluster incidents Own Support Incident coordination
Application incidents Support platform Own Escalation
Security policies Own guardrails Comply Exceptions
Cost optimization Infrastructure efficiency Workload efficiency Shared

The important part is not the exact table. It is the principle behind it:

Platform owns the boundaries. Application teams own the workload behavior inside those boundaries.


Cluster Infrastructure & Upgrades: Platform/SRE Owns the Floor

The Kubernetes control plane and worker infrastructure should normally belong to the platform or SRE organization.

That includes:

  • Kubernetes version upgrades
  • Control plane lifecycle
  • Node pools
  • Container runtime configuration
  • CNI
  • CSI infrastructure
  • CoreDNS
  • Ingress infrastructure
  • Cluster-wide observability
  • Cluster-wide security controls
  • Autoscaling infrastructure
  • Node replacement
  • Cluster capacity

Application developers should not need to understand how the control plane is upgraded in order to deploy an application. That is exactly what the platform exists to abstract. But abstraction does not mean isolation. A cluster upgrade can still break an application.

Imagine the platform team wants to replace a node pool. They cordon nodes and begin draining them. Kubernetes attempts to evict workloads. Then the drain gets stuck.

The reason?

An application has a single replica and an overly restrictive PodDisruptionBudget:

spec:
  minAvailable: 1
Enter fullscreen mode Exit fullscreen mode

The platform team owns the node lifecycle. The application team owns the workload availability configuration. This is where operational boundaries become shared interfaces rather than rigid walls.

Pod Disruption Budgets Are a Contract

A PodDisruptionBudget tells the platform:

"This is how much voluntary disruption this workload can tolerate."

The platform team must respect it. The application team must configure it correctly. Neither side can treat it as someone else’s problem. The same applies to resource requests and limits.

If a workload declares:

resources:
  requests:
    cpu: "500m"
    memory: "512Mi"
  limits:
    cpu: "1"
    memory: "1Gi"
Enter fullscreen mode Exit fullscreen mode

the scheduler and kubelet use those values to make placement and enforcement decisions. The platform provides the scheduling environment. The application team provides realistic workload requirements.

A production node upgrade therefore becomes a coordination problem, not a ticket handoff.


Workload Lifecycle & Health: Application Teams Own What They Deploy

Once the platform provides a namespace and the required primitives, the application team should own the lifecycle of its workload.

That includes:

  • Deployment manifests
  • StatefulSets
  • Jobs
  • CronJobs
  • Services
  • Ingress configuration
  • ConfigMaps
  • Application-level RBAC
  • Resource requests and limits
  • Horizontal Pod Autoscalers
  • Probes
  • PodDisruptionBudgets
  • Application dependencies
  • Deployment strategies
  • Application-level dashboards and alerts

This does not mean developers need unrestricted access to the cluster. They should have enough access to operate their applications without requiring the platform team to perform routine actions for them.

A developer should generally be able to answer:

Is my deployment running?
Are my pods ready?
Why did my pod restart?
What does the application log say?
What image version is deployed?
Is my rollout progressing?
Are my resources being throttled?
Enter fullscreen mode Exit fullscreen mode

If answering these questions requires opening a Jira ticket, the platform has created unnecessary friction. The solution is not cluster-admin. The solution is better scoped access.


Access Control, Secrets & RBAC Boundaries

One of the easiest ways to create a dangerous Kubernetes environment is to give developers cluster-admin because it makes everything easier. It also eliminates the need to design access properly. That is precisely why it should be avoided.

A developer working on an application normally does not need permission to:

delete nodes
modify cluster-wide RBAC
change admission policies
modify CRDs
inspect unrelated namespaces
change networking infrastructure
delete persistent volumes belonging to another team
Enter fullscreen mode Exit fullscreen mode

Their access should normally be scoped around their namespace and the resources they actually operate.

For example:

Cluster
|
+-- platform-system
|
+-- monitoring
|
+-- team-a
|   +-- Deployment
|   +-- Service
|   +-- ConfigMap
|   +-- Secret
|   +-- Pods
|
+-- team-b
    +-- Deployment
    +-- Service
    +-- ConfigMap
    +-- Secret
    +-- Pods
Enter fullscreen mode Exit fullscreen mode

Team A should not need administrative access to Team B’s namespace. This is where Kubernetes Role and RoleBinding become operational boundaries rather than just security objects.

Production kubectl Access Is a Design Smell

Direct production access through kubectl is not automatically forbidden. Emergency debugging sometimes requires it. The problem is treating unrestricted interactive access as the normal deployment mechanism.

If engineers routinely run:

kubectl edit deployment production-api
Enter fullscreen mode Exit fullscreen mode

or:

kubectl set image deployment/api api=myimage:latest
Enter fullscreen mode Exit fullscreen mode

you have created a second configuration system. Git says one thing. The cluster says another. The next GitOps reconciliation may overwrite the manual change. Or worse, the manual change becomes an undocumented production configuration that nobody remembers making.

A healthier model is:

Git defines desired state. Kubernetes executes that state. GitOps continuously reconciles the two.


Managing Configuration Drift via GitOps

Configuration drift is one of the biggest Day 2 problems because it is often invisible until something fails.

Consider this sequence:

Monday:
Git says replicas = 3

Tuesday:
Developer changes replicas to 5 with kubectl

Wednesday:
Platform engineer upgrades the cluster

Thursday:
GitOps reconciliation changes replicas back to 3

Friday:
Traffic increases and the application behaves differently
Enter fullscreen mode Exit fullscreen mode

Nobody necessarily made a mistake on Friday. The system simply exposed a configuration ownership problem that started on Tuesday. GitOps tools such as Argo CD and Flux solve part of this problem by turning Git into the declared source of truth.

A simplified model looks like this:

                Git Repository
                     |
                     v
             Desired State
                     |
                     v
              GitOps Controller
                     |
                     v
              Kubernetes API
                     |
                     v
             Running Workload
                     |
                     |
              Drift Detection
                     |
                     +------> Reconcile
Enter fullscreen mode Exit fullscreen mode

The critical organizational decision is not merely adopting Argo CD or Flux. It is deciding which repository owns which resources.

For example:

platform-config/
├── ingress-controller/
├── cert-manager/
├── monitoring/
└── policies/

application-config/
├── payments/
├── checkout/
└── catalog/
Enter fullscreen mode Exit fullscreen mode

The platform repository owns platform components. The application repository owns application configuration. This creates an explicit ownership contract. If a developer wants to change an application deployment, they change the application repository. If the platform team wants to change the ingress controller, it changes the platform repository. Neither team needs to manually edit live resources.

Git Becomes More Than Version Control

In a mature GitOps environment, Git becomes an operational contract.

It answers:

  • Who owns this resource?
  • Who approved the change?
  • What changed?
  • When did it change?
  • What should the cluster look like?
  • Who can modify the configuration?
  • Can we reproduce the environment?

That is far more useful than simply saying, "Everything is managed through Git." The real value is clear ownership of desired state.


Access Should Be Self-Service, Not Permissionless

There is an important distinction between self-service and unrestricted access. Self-service means a developer can perform an approved operation without opening a ticket. It does not mean the developer can bypass every control.

For example, a platform can provide an application template:

service:
  name: payments

resources:
  requests:
    cpu: 250m
    memory: 256Mi

autoscaling:
  enabled: true
  minReplicas: 2
  maxReplicas: 10
Enter fullscreen mode Exit fullscreen mode

The developer controls application-specific values.

The platform enforces constraints such as:

Maximum CPU request
Maximum memory request
Required securityContext
Required probes
Allowed container registries
Required labels
Allowed ingress classes
Approved storage classes
Enter fullscreen mode Exit fullscreen mode

This is where admission policies, resource quotas, namespace templates, and policy engines become useful. The platform team defines the guardrails. The developer gets the steering wheel.


Resource Management & Cost: Who Owns the OOMKilled Pod?

This is where ownership models often become emotional.

A pod gets:

OOMKilled
Enter fullscreen mode Exit fullscreen mode

The application team says:

"Kubernetes killed our application."

The platform team says:

"You configured a 256Mi memory limit."

Both statements describe part of the situation. But operational ownership should still be clear. The application team owns the workload’s resource requirements. The platform team owns the cluster’s capacity and enforcement mechanisms.

A useful distinction is:

Application Team Owns

  • CPU requests
  • Memory requests
  • CPU limits where appropriate
  • Memory limits
  • Autoscaling configuration
  • Application efficiency
  • Capacity requirements
  • Workload behavior under load

Platform Team Owns

  • Node capacity
  • Cluster autoscaling
  • ResourceQuota enforcement
  • LimitRange defaults
  • Scheduling infrastructure
  • Node sizing
  • Cluster-level capacity planning
  • Cost visibility at the infrastructure layer

If a Java application genuinely requires 2Gi but is configured with a 512Mi memory limit, the platform team should not be expected to permanently compensate by adding larger nodes.

That is hiding an application configuration problem inside infrastructure spending. At the same time, if dozens of workloads are correctly requesting resources and the cluster consistently runs at capacity, the platform team cannot tell developers to "optimize their pods" forever.

That becomes a platform capacity problem.

Resource ownership is therefore shared at the boundary, but accountability remains specific.


Failure Modes of Bad Ownership: Ticket-Ops vs. The Wild West

There are two predictable extremes. Both are bad.

Failure Mode #1: Ticket-Ops

The organization tries to maintain control by putting the platform team in the middle of everything. A developer needs to change an environment variable. They create a Jira ticket. They need to increase replicas. Another ticket. They need to restart a deployment. Another ticket. They need logs from production. Another ticket. They need a namespace. Another ticket.

Eventually the platform team becomes an infrastructure help desk.

The consequences are predictable:

Developer
   |
   v
Jira Ticket
   |
   v
Platform Engineer
   |
   v
Manual kubectl
   |
   v
Production
Enter fullscreen mode Exit fullscreen mode

This does not create governance. It creates a bottleneck. Worse, it encourages developers to find workarounds.

Failure Mode #2: The Wild West

The opposite model gives every engineering team broad production access.

Everyone can run:

kubectl delete pod
kubectl edit deployment
kubectl apply -f production.yaml
Enter fullscreen mode Exit fullscreen mode

It feels productive at first. Then someone changes a production deployment manually. Another engineer deletes a resource they thought was unused. A developer modifies a cluster-wide object while troubleshooting. Nobody knows who owns the change. Auditability disappears. The platform team eventually responds by introducing even more controls. The organization swings from Wild West to Ticket-Ops.

This cycle happens repeatedly when teams treat permissions as the ownership model. They are not the same thing.


The Better Model: Guardrails, Not Gates

A mature platform does not ask:

"How can we stop developers from changing things?"

It asks:

"How can developers safely change the things they own without being able to damage the things they do not?"

That leads to a much better architecture.

                 Platform Team
                       |
       +---------------+----------------+
       |               |                |
   Guardrails       Paved Roads      Platform APIs
       |               |                |
       +---------------+----------------+
                       |
                 Kubernetes
                       |
        +--------------+--------------+
        |              |              |
     Team A          Team B         Team C
        |              |              |
     Workloads      Workloads      Workloads
        |              |              |
      GitOps         GitOps         GitOps
Enter fullscreen mode Exit fullscreen mode

The platform team does not become the operator of every workload.

Instead, it provides a safe operating environment.

That means:

  • Namespace-level RBAC
  • Standard workload templates
  • GitOps
  • Policy enforcement
  • Resource quotas
  • Observability
  • Self-service deployment
  • Automated namespace provisioning
  • Standard secrets management
  • Documented escalation paths
  • Automated cluster upgrades
  • Clear ownership metadata

The result is a platform where developers can move quickly without requiring administrative access to the entire cluster.


Pragmatic Recommendations for Mid-to-Senior Engineers

The ownership model does not need to be perfect on day one. It needs to be explicit. Start by documenting the operational boundary.

1. Write the Ownership Contract

For every major Kubernetes resource, answer:

Who creates it?
Who modifies it?
Who approves changes?
Who monitors it?
Who gets paged?
Who can delete it?
Enter fullscreen mode Exit fullscreen mode

If the answer is "everyone," the boundary is probably not defined well enough.


2. Keep Cluster-Admin Extremely Small

cluster-admin should be an exceptional privilege. Platform engineers who maintain the cluster may need it. Most application developers do not. Build namespace-scoped roles that provide exactly what teams need. The objective is not zero access. The objective is appropriate access.


3. Make Git the Default Change Path

If production configuration is represented in Git, developers should be able to submit pull requests rather than tickets.

A good flow looks like:

Developer
   |
   v
Pull Request
   |
   v
Validation / Policy Checks
   |
   v
Review
   |
   v
GitOps
   |
   v
Kubernetes
Enter fullscreen mode Exit fullscreen mode

This gives developers autonomy while preserving review, history, and rollback.


4. Automate the Boring Requests

If developers frequently ask the platform team for:

  • Namespace creation
  • Standard RBAC
  • Service accounts
  • Basic monitoring
  • Resource quotas
  • Application templates
  • Standard ingress
  • Deployment scaffolding

then those are candidates for platform automation. A platform engineer should not manually create the same namespace structure twenty times.


5. Define Upgrade Contracts

Before upgrading Kubernetes, application teams should know:

  • Which Kubernetes version is being targeted
  • What APIs are deprecated
  • When workloads will be disrupted
  • What PDB behavior is expected
  • Which applications need compatibility testing
  • How rollback or remediation works

The platform team owns the upgrade. Application teams own application compatibility.


6. Make Resource Requests a First-Class Engineering Concern

Do not treat CPU and memory values as YAML decoration.

They affect:

  • Scheduling
  • Cluster capacity
  • Autoscaling
  • Cost
  • Reliability
  • Eviction behavior
  • Upgrade safety

Every production workload should have a reasonable resource profile. If the platform automatically rejects workloads without resource requests, that is not bureaucracy. It is preventing undefined scheduling behavior from becoming a production incident.


7. Give Developers Production Visibility Without Production Control

This is one of the most useful boundaries an organization can establish.

Developers should usually be able to see:

Logs
Metrics
Traces
Events
Deployment status
Pod status
Resource consumption
Rollout history
Enter fullscreen mode Exit fullscreen mode

without automatically receiving permission to modify:

Nodes
Cluster RBAC
Admission policies
Other teams’ namespaces
Cluster networking
Control-plane resources
Enter fullscreen mode Exit fullscreen mode

Read access and write access do not need to be coupled.

This single distinction can eliminate a surprising amount of unnecessary cluster-admin usage.


8. Build Incident Ownership Around the Failure Domain

When something breaks, do not ask:

"Which team owns Kubernetes?"

Ask:

"Where is the failure?"

If the API server is unavailable, platform owns the incident. If nodes are failing, platform owns the incident. If an application’s readiness probe is broken, the application team owns the incident. If the ingress controller is failing globally, platform owns it. If one application’s ingress rule is incorrect, the application team owns it.

If both sides contributed to the failure, establish a shared incident channel and resolve the immediate problem first. Conduct the ownership discussion during the follow-up, not while production traffic is burning.

That is what mature operations looks like.


A Practical Day 2 Operating Model

A healthy Kubernetes organization should eventually look something like this:

                    PLATFORM TEAM
                         |
       +-----------------+------------------+
       |                 |                  |
   Kubernetes         Guardrails        Paved Roads
   Infrastructure       |                  |
       |            RBAC / Policy       GitOps
       |            Quotas / Security   Templates
       |                                  |
       +----------------+-----------------+
                        |
                   Kubernetes API
                        |
       +----------------+----------------+
       |                |                |
    Team A            Team B           Team C
       |                |                |
   Application       Application      Application
   Ownership         Ownership        Ownership
       |                |                |
      Git              Git              Git
       |                |                |
     GitOps           GitOps           GitOps
Enter fullscreen mode Exit fullscreen mode

This architecture creates a useful separation:

Platform engineering is responsible for making the platform safe and usable.

Application engineering is responsible for making the application reliable.

Those responsibilities overlap, but they should not be confused.


3 Key Takeaways

1. Platform owns the platform. Application teams own the workloads.

The platform team should own Kubernetes infrastructure, upgrades, node lifecycle, cluster-wide security, networking primitives, and platform reliability.

Application teams should own deployments, configuration, resource requirements, application health, and application-level operational behavior. The boundary should be explicit.

2. Self-service beats both Ticket-Ops and unrestricted access.

Developers should not need a Jira ticket to make routine changes to workloads they own. They also should not need cluster-admin to do it.

Use namespace-scoped RBAC, GitOps, policy enforcement, templates, and automated workflows to give developers autonomy inside well-defined boundaries. The goal is not to remove control.

The goal is to move control into automation and policy instead of human gatekeeping.

3. Day 2 ownership is an architectural decision, not an org-chart decision.

Kubernetes does not automatically create operational boundaries. Your platform architecture does. Git repositories, RBAC roles, namespaces, resource quotas, admission policies, GitOps controllers, observability systems, and incident processes should all reinforce the same ownership model.

When they do, the question changes from:

"Who owns Kubernetes?"

to the much more useful question:

"Who owns this layer of the platform, and what does the other team need from them?"

That is the boundary worth designing. Because a Kubernetes cluster is not the product. The platform is the product.

And on Day 2, the quality of that platform is measured less by whether the cluster is running and more by whether engineers can safely operate what is running on it.

Top comments (0)