<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Kenta Yamaguchi</title>
    <description>The latest articles on DEV Community by Kenta Yamaguchi (@key60228).</description>
    <link>https://dev.to/key60228</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4134978%2Fe79fff1f-d00f-435a-8752-52d19d663cc0.jpg</url>
      <title>DEV Community: Kenta Yamaguchi</title>
      <link>https://dev.to/key60228</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/key60228"/>
    <language>en</language>
    <item>
      <title>Gotchas When Setting Up Open Cluster Management (OCM) on GKE</title>
      <dc:creator>Kenta Yamaguchi</dc:creator>
      <pubDate>Sun, 04 Oct 2026 14:31:28 +0000</pubDate>
      <link>https://dev.to/key60228/gotchas-when-setting-up-open-cluster-management-ocm-on-gke-5f69</link>
      <guid>https://dev.to/key60228/gotchas-when-setting-up-open-cluster-management-ocm-on-gke-5f69</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;This is an English translation of my article originally published in Japanese on Zenn: &lt;a href="https://zenn.dev/aishift/articles/050045247128e9" rel="noopener noreferrer"&gt;GKE に Open Cluster Management (OCM) を導入するのにハマったことメモ&lt;/a&gt;.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Hi, I'm &lt;a href="https://twitter.com/key60228" rel="noopener noreferrer"&gt;@key60228&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;In the Kubernetes and Cloud Native world, multi-cluster operations have been getting more and more attention lately, alongside AI workloads.&lt;/p&gt;

&lt;p&gt;There were several sessions on the topic at KubeCon + CloudNativeCon Japan 2026 as well. (The first one below is essentially a "don't go multi-cluster lightly" talk.)&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/-0gJNQogilQ" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/G8yOpne_T04" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;In this post I'll go over the issues I ran into while rolling out &lt;a href="https://open-cluster-management.io/" rel="noopener noreferrer"&gt;Open Cluster Management (OCM)&lt;/a&gt;, which the second session mentions, together with its add-on &lt;a href="https://open-cluster-management.io/docs/getting-started/integration/fleetconfig-controller/" rel="noopener noreferrer"&gt;FleetConfig Controller&lt;/a&gt;, on GKE.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Open Cluster Management (OCM)?
&lt;/h2&gt;

&lt;p&gt;I'll skip the details, but in short, it's a project for managing multiple Kubernetes clusters from one place.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/open-cluster-management-io" rel="noopener noreferrer"&gt;https://github.com/open-cluster-management-io&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It was originally started by Red Hat, donated to the CNCF in 2021, and is currently a Sandbox project.&lt;sup id="fnref1"&gt;1&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;It uses a hub-spoke architecture. The hub cluster holds the desired state for each spoke cluster, and an agent on the spoke called the Klusterlet pulls that state and applies it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1cevtst0z1nypuarqfuq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1cevtst0z1nypuarqfuq.png" alt="OCM architecture"&gt;&lt;/a&gt;&lt;br&gt;
&lt;em&gt;Source: &lt;a href="https://open-cluster-management.io/docs/concepts/architecture/" rel="noopener noreferrer"&gt;https://open-cluster-management.io/docs/concepts/architecture/&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;The desired state is defined in a &lt;code&gt;ManifestWork&lt;/code&gt; resource, and each spoke cluster is represented by a &lt;code&gt;ManagedCluster&lt;/code&gt; resource.&lt;/p&gt;

&lt;p&gt;The Klusterlet treats every &lt;code&gt;ManifestWork&lt;/code&gt; in the &lt;code&gt;Namespace&lt;/code&gt; that shares its &lt;code&gt;ManagedCluster&lt;/code&gt;'s name as its own desired state and applies it.&lt;/p&gt;
&lt;h2&gt;
  
  
  What is FleetConfig Controller?
&lt;/h2&gt;

&lt;p&gt;OCM has an &lt;a href="https://open-cluster-management.io/docs/concepts/add-on-extensibility/addon/" rel="noopener noreferrer"&gt;add-on&lt;/a&gt; mechanism for extending its functionality, and FleetConfig Controller is one of the officially provided add-ons.&lt;/p&gt;


&lt;div class="ltag-github-readme-tag"&gt;
  &lt;div class="readme-overview"&gt;
    &lt;h2&gt;
      &lt;img src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg" alt="GitHub logo"&gt;
      &lt;a href="https://github.com/open-cluster-management-io" rel="noopener noreferrer"&gt;
        open-cluster-management-io
      &lt;/a&gt; / &lt;a href="https://github.com/open-cluster-management-io/lab" rel="noopener noreferrer"&gt;
        lab
      &lt;/a&gt;
    &lt;/h2&gt;
    &lt;h3&gt;
      Lab projects for Open Cluster Management
    &lt;/h3&gt;
  &lt;/div&gt;
  &lt;div class="ltag-github-body"&gt;
    
&lt;div id="readme" class="md"&gt;&lt;p&gt;&lt;a rel="noopener noreferrer" href="https://github.com/open-cluster-management-io/lab/assets/ocm-lab-logo.jpg"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fraw.githubusercontent.com%2Fopen-cluster-management-io%2Flab%2FHEAD%2Fassets%2Focm-lab-logo.jpg" alt="image"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://opensource.org/licenses/Apache-2.0" rel="nofollow noopener noreferrer"&gt;&lt;img src="https://camo.githubusercontent.com/5b60841bea9e11d9d0b0950d690c9bc554e06385634056a7d5d62a15d1a4eabe/68747470733a2f2f696d672e736869656c64732e696f2f62616467652f4c6963656e73652d4170616368655f322e302d626c75652e737667" alt="License"&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Welcome to the lab repo for &lt;a href="https://open-cluster-management.io/" rel="nofollow noopener noreferrer"&gt;Open Cluster Management (OCM)&lt;/a&gt;
This repo hosts experimental projects that anyone in the community can try out, provide feedback on, and contribute to
Feel free to link to these projects from your own websites or repos, gauge interest, and help us improve as we iterate.&lt;/p&gt;
&lt;p&gt;Unlike a dedicated labs GitHub org (ie: &lt;a href="https://github.com/argoproj-labs" rel="noopener noreferrer"&gt;argoproj-labs&lt;/a&gt;)
this repo hosts all lab projects in one place.
New projects are onboarded via PR and added as subfolders, each governed by its own &lt;code&gt;OWNERS&lt;/code&gt; file.
Since our community is still small compared to the &lt;a href="https://github.com/argoproj" rel="noopener noreferrer"&gt;arogoproj&lt;/a&gt;,
keeping everything together avoids unnecessary fragmentation.&lt;/p&gt;
&lt;p&gt;For new add-on projects, please use the
&lt;a href="https://github.com/open-cluster-management-io/addon-contrib" rel="noopener noreferrer"&gt;addon-contrib&lt;/a&gt; repo.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Table of Contents&lt;/h2&gt;

&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://github.com/open-cluster-management-io/lab#current-projects" rel="noopener noreferrer"&gt;Current Projects&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/open-cluster-management-io/lab#onboarding-a-new-project" rel="noopener noreferrer"&gt;Onboarding a New Project&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/open-cluster-management-io/lab#governance" rel="noopener noreferrer"&gt;Governance&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/open-cluster-management-io/lab#issues-for-lab-projects" rel="noopener noreferrer"&gt;Issues for Lab Projects&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/open-cluster-management-io/lab#prs-for-lab-projects" rel="noopener noreferrer"&gt;PRs for Lab Projects&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Current Projects&lt;/h2&gt;

&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://github.com/open-cluster-management-io/lab/./ai-assistance/README.md" rel="noopener noreferrer"&gt;ai-assistance&lt;/a&gt;: Vendor-neutral AI assistance content (prompts, guides) for OCM development.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/open-cluster-management-io/lab/./dashboard/README.md" rel="noopener noreferrer"&gt;dashboard&lt;/a&gt;: OCM UI Dashboard.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://github.com/open-cluster-management-io/lab/./fleetconfig-controller/README.md" rel="noopener noreferrer"&gt;fleetconfig-controller&lt;/a&gt;…&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
  &lt;/div&gt;
  &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/open-cluster-management-io/lab" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
&lt;/div&gt;


&lt;p&gt;According to the &lt;a href="https://open-cluster-management.io/docs/getting-started/installation/" rel="noopener noreferrer"&gt;docs&lt;/a&gt;, the de facto standard way to set up OCM is the &lt;a href="https://github.com/open-cluster-management-io/clusteradm" rel="noopener noreferrer"&gt;clusteradm&lt;/a&gt; CLI. FleetConfig Controller wraps those clusteradm operations behind two custom resources, &lt;code&gt;Hub&lt;/code&gt; and &lt;code&gt;Spoke&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;Hub&lt;/code&gt; CR takes over what &lt;code&gt;clusteradm init&lt;/code&gt; does (initializing the hub cluster), and the &lt;code&gt;Spoke&lt;/code&gt; CR takes over what &lt;code&gt;clusteradm join&lt;/code&gt; does (registering a spoke cluster as a &lt;code&gt;ManagedCluster&lt;/code&gt;).&lt;/p&gt;

&lt;p&gt;We already run Argo CD as part of the platform at my company, AI Shift, and wanted to stay as close to GitOps as possible. We also wanted to cut down on the toil of adding spoke clusters and upgrading things like the Klusterlet on existing spokes. So we decided to give it a try.&lt;/p&gt;

&lt;h2&gt;
  
  
  The environment
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Details&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hub cluster&lt;/td&gt;
&lt;td&gt;GKE on Project α (Standard mode, private nodes, public endpoint enabled)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spoke cluster&lt;/td&gt;
&lt;td&gt;GKE on Project β (Standard mode, private nodes, public endpoint enabled)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;FleetConfig Controller and the Hub / Spoke CRs are deployed to the hub cluster as a Helm chart via Argo CD.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setup flow
&lt;/h2&gt;

&lt;p&gt;The overall procedure looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install FleetConfig Controller on the hub cluster&lt;/li&gt;
&lt;li&gt;Create the Hub CR&lt;/li&gt;
&lt;li&gt;Create a bootstrap kubeconfig for the spoke cluster and register it on the hub as a Secret&lt;/li&gt;
&lt;li&gt;Create the Spoke CR&lt;/li&gt;
&lt;li&gt;Once the join completes, delete the bootstrap resources&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The only real difference from the &lt;a href="https://github.com/open-cluster-management-io/lab/tree/main/fleetconfig-controller#%E2%80%8D%EF%B8%8F-quick-start" rel="noopener noreferrer"&gt;kind-based quick start&lt;/a&gt; is that you build the bootstrap kubeconfig yourself. Nothing special beyond that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gotchas
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. clusteradm join fails with i/o timeout
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9oiwlr06t1guj3mmdbuo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9oiwlr06t1guj3mmdbuo.png" alt="case-1"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After creating the Spoke CR, its PHASE stayed at &lt;code&gt;Unhealthy&lt;/code&gt; and the status showed this error:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;clusteradm join command failed for spoke spoke-1: exit status 1, output:
W0821 12:32:43 exec.go:250] Join continues without an external API server URL for the klusterlet because :
Get "https://xxx.xxx.xxx.xxx/api/v1/namespaces/kube-public/configmaps/cluster-info": dial tcp xxx.xxx.xxx.xxx:443: i/o timeout
...
Error: Get "https://xxx.xxx.xxx.xxx/apis/apps/v1/namespaces/open-cluster-management/deployments/klusterlet": dial tcp xxx.xxx.xxx.xxx:443: i/o timeout
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;xxx.xxx.xxx.xxx&lt;/code&gt; is the spoke's public endpoint.&lt;/p&gt;

&lt;p&gt;The &lt;code&gt;clusteradm join&lt;/code&gt; equivalent is run by the controller on the hub, which talks directly to the spoke's kube-apiserver. So something on the path from hub to spoke was blocking the connection.&lt;/p&gt;

&lt;p&gt;The cause was that &lt;code&gt;gcp_public_cidrs_access_enabled&lt;/code&gt; ("Access using Google Cloud public IP addresses") was set to &lt;code&gt;false&lt;/code&gt; on the spoke GKE cluster.&lt;/p&gt;

&lt;p&gt;When this setting is &lt;code&gt;false&lt;/code&gt;, access from Google Cloud public IP ranges is rejected even if you put &lt;code&gt;0.0.0.0/0&lt;/code&gt; in the master authorized networks.&lt;sup id="fnref2"&gt;2&lt;/sup&gt;&lt;/p&gt;

&lt;p&gt;Connections from the hub come from its Cloud NAT IP, which is a Google Cloud public IP, so they were being dropped right there.&lt;/p&gt;
&lt;h3&gt;
  
  
  2. The Hub CR needs spec.apiServer
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ac2agu4rs80qkv07i00.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ac2agu4rs80qkv07i00.png" alt="case-2"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Once the network was fixed, the join itself went through, but the Spoke CR then got stuck at &lt;code&gt;Joining&lt;/code&gt;. The klusterlet registration-agent on the spoke kept logging this error:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Get "https://10.2.0.1:443/apis/cluster.open-cluster-management.io/v1/managedclusters/spoke-1-ab513522":
tls: failed to verify certificate: x509: certificate signed by unknown authority
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;&lt;code&gt;10.2.0.1&lt;/code&gt; is an in-cluster ClusterIP.&lt;/p&gt;

&lt;p&gt;If the kubeconfig setting on the Hub CR is just &lt;code&gt;inCluster: true&lt;/code&gt;, the "hub API server URL" that FleetConfig Controller hands to the spoke cluster also ends up being the in-cluster address.&lt;/p&gt;

&lt;p&gt;From the spoke cluster's point of view, &lt;code&gt;https://10.2.0.1&lt;/code&gt; is its own kube-apiserver, so the certificate can't be verified against the hub's CA and you get a TLS error.&lt;/p&gt;

&lt;p&gt;On kind (kubeadm), the &lt;code&gt;kube-public/cluster-info&lt;/code&gt; ConfigMap exists, so FleetConfig Controller can "helpfully" fill in the hub's endpoint and pass it to the spoke. GKE doesn't have that ConfigMap, so the hub endpoint has to be set explicitly.&lt;/p&gt;

&lt;p&gt;Setting the hub cluster's endpoint in &lt;code&gt;spec.apiServer&lt;/code&gt; on the Hub CR fixed it.&lt;/p&gt;
&lt;h3&gt;
  
  
  3. Only fleetconfig-controller-agent goes into ImagePullBackOff
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2pz8rxdvgaxfafycpef.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2pz8rxdvgaxfafycpef.png" alt="case-3"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After the ManagedCluster reached &lt;code&gt;JOINED=True / AVAILABLE=True&lt;/code&gt; and all the klusterlet Pods were Running, the fleetconfig-controller-agent on the spoke cluster, and only that Pod, went into &lt;code&gt;ImagePullBackOff&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Failed to pull image "asia-northeast1-docker.pkg.dev/&amp;lt;HUB_PROJECT&amp;gt;/remote-quay-io/open-cluster-management/fleetconfig-controller:v0.3.5@sha256:...":
... 403 Forbidden
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;In our setup, the hub cluster pulls quay.io images through an Artifact Registry remote repository.&lt;/p&gt;

&lt;p&gt;For the klusterlet images (operator / registration / work), I had already overridden the references via &lt;code&gt;spec.klusterlet.values.images.overrides&lt;/code&gt; to point directly at quay.io. The fleetconfig-controller-agent image, however, has to be overridden through a different path.&lt;/p&gt;

&lt;p&gt;Because I had missed that override, the hub-side Helm chart's &lt;code&gt;image.repository&lt;/code&gt; was written into the AddOnTemplate. The spoke cluster's nodes then tried to pull from the Artifact Registry in the hub's Google Cloud project, had no permission, and got a 403.&lt;/p&gt;

&lt;p&gt;The fix was to set &lt;code&gt;spec.addOns[].deploymentConfig.registries.source&lt;/code&gt; to the hub's Artifact Registry and &lt;code&gt;spec.addOns[].deploymentConfig.registries.mirror&lt;/code&gt; to the public quay.io repository.&lt;/p&gt;
&lt;h3&gt;
  
  
  4. Unauthorized due to an expired bootstrap token
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7gjjcoso7guhunkalsr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh7gjjcoso7guhunkalsr.png" alt="case-4"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;After spending a while on the issues above, I tried the join again and got yet another error:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;clusteradm join command failed for spoke spoke-1: exit status 1, output:
W0824 01:54:54 exec.go:250] Join continues without an external API server URL for the klusterlet because : Unauthorized
Error: Unauthorized
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;This one was simple: I had given the bootstrap kubeconfig token a short lifetime, and it had just expired.&lt;/p&gt;

&lt;p&gt;Reissuing the token and replacing the Secret fixed it.&lt;/p&gt;
&lt;h3&gt;
  
  
  5. ClusterRoleBinding error from the webhook
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmuegp5booz251fecp0n3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmuegp5booz251fecp0n3.png" alt="case-5"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At some point, the FleetConfig Controller Pod went into &lt;code&gt;CrashLoopBackOff&lt;/code&gt; and kept crashing.&lt;/p&gt;

&lt;p&gt;The fleetconfig-controller-manager logs showed this:&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;2026-08-28T07:28:00Z       ERROR   setup   problem running manager {"error": "failed to create or update global ManagedClusterSetBinding: admission webhook \"managedclustersetbindingvalidators.admission.cluster.open-cluster-management.io\" denied the request: managedclustersets/bind.apps \"global\" is forbidden: user \"system:serviceaccount:fleetconfig-system:fleetconfig-controller-manager\" is not allowed to bind cluster set \"global\""}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;The cause was that I had created the Hub CR once in the &lt;code&gt;γ&lt;/code&gt; namespace, then moved it to a different namespace, &lt;code&gt;δ&lt;/code&gt; (by deleting and recreating it).&lt;/p&gt;

&lt;p&gt;By default, the FleetConfig Controller Helm chart bundles the OCM CRDs and creates a &lt;code&gt;ManagedClusterSet&lt;/code&gt; and &lt;code&gt;ManagedClusterSetBinding&lt;/code&gt; at startup.&lt;/p&gt;

&lt;p&gt;On the other hand, deleting the Hub CR runs &lt;code&gt;clusteradm clean&lt;/code&gt; under the hood, which also deletes the OCM CRDs.&lt;/p&gt;

&lt;p&gt;Deleting a CRD deletes its CRs too,&lt;sup id="fnref3"&gt;3&lt;/sup&gt; so the &lt;code&gt;ManagedClusterSetBinding&lt;/code&gt; that FleetConfig Controller had created was gone.&lt;/p&gt;

&lt;p&gt;When the FleetConfig Controller Pod later restarted, it tried to recreate the &lt;code&gt;ManagedClusterSetBinding&lt;/code&gt;, but because of a missing RBAC rule, the admission webhook's &lt;code&gt;SubjectAccessReview&lt;/code&gt; returned Forbidden.&lt;/p&gt;

&lt;p&gt;On the very first startup, the Hub CR hadn't been initialized yet and the validating webhook didn't exist, so the request skipped the RBAC check and the controller came up fine.&lt;/p&gt;

&lt;p&gt;For now I worked around it by defining the ClusterRole / ClusterRoleBinding explicitly on our side.&lt;br&gt;
&lt;/p&gt;
&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rbac.authorization.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterRole&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fleetconfig-controller-clusterset-bind&lt;/span&gt;
&lt;span class="na"&gt;rules&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;apiGroups&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;cluster.open-cluster-management.io&lt;/span&gt;
    &lt;span class="na"&gt;resources&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;managedclustersets/bind&lt;/span&gt;
    &lt;span class="na"&gt;resourceNames&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;default&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;global&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;spokes&lt;/span&gt;
    &lt;span class="na"&gt;verbs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="s"&gt;create&lt;/span&gt;
&lt;span class="nn"&gt;---&lt;/span&gt;
&lt;span class="na"&gt;apiVersion&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rbac.authorization.k8s.io/v1&lt;/span&gt;
&lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterRoleBinding&lt;/span&gt;
&lt;span class="na"&gt;metadata&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fleetconfig-controller-clusterset-bind&lt;/span&gt;
&lt;span class="na"&gt;roleRef&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;apiGroup&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rbac.authorization.k8s.io&lt;/span&gt;
  &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ClusterRole&lt;/span&gt;
  &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fleetconfig-controller-clusterset-bind&lt;/span&gt;
&lt;span class="na"&gt;subjects&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;kind&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ServiceAccount&lt;/span&gt;
    &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fleetconfig-controller-manager&lt;/span&gt;
    &lt;span class="na"&gt;namespace&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;fleetconfig-system&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;


&lt;p&gt;(I also sent a fix upstream and it has been merged, so this shouldn't happen in future releases.)&lt;/p&gt;


&lt;div class="ltag_github-liquid-tag"&gt;
  &lt;h1&gt;
    &lt;a href="https://github.com/open-cluster-management-io/lab/pull/249" rel="noopener noreferrer"&gt;
      &lt;img class="github-logo" alt="GitHub logo" src="https://assets.dev.to/assets/github-logo-5a155e1f9a670af7944dd5e12375bc76ed542ea80224905ecaf878b9157cdefc.svg"&gt;
      &lt;span class="issue-title"&gt;
        fix: add managedclustersets/bind RBAC for topology resources (fleetconfig-controller)
      &lt;/span&gt;
      &lt;span class="issue-number"&gt;#249&lt;/span&gt;
    &lt;/a&gt;
  &lt;/h1&gt;
  &lt;div class="github-thread"&gt;
    &lt;div class="timeline-comment-header"&gt;
      &lt;a href="https://github.com/KEY60228" rel="noopener noreferrer"&gt;
        &lt;img class="github-liquid-tag-img" src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Favatars.githubusercontent.com%2Fu%2F56732734%3Fv%3D4" alt="KEY60228 avatar"&gt;
      &lt;/a&gt;
      &lt;div class="timeline-comment-header-text"&gt;
        &lt;strong&gt;
          &lt;a href="https://github.com/KEY60228" rel="noopener noreferrer"&gt;KEY60228&lt;/a&gt;
        &lt;/strong&gt; posted on &lt;a href="https://github.com/open-cluster-management-io/lab/pull/249" rel="noopener noreferrer"&gt;&lt;time&gt;Aug 31, 2026&lt;/time&gt;&lt;/a&gt;
      &lt;/div&gt;
    &lt;/div&gt;
    &lt;div class="ltag-github-body"&gt;
      &lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;Summary&lt;/h3&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;p&gt;The manager crashed at startup because OCM's ManagedClusterSetBinding validating webhook denied the topology resource creation: the &lt;code&gt;SubjectAccessReview&lt;/code&gt; on &lt;code&gt;managedclustersets/bind&lt;/code&gt; failed for the &lt;code&gt;fleetconfig-controller-manager&lt;/code&gt; ServiceAccount. Added the missing &lt;code&gt;managedclustersets/bind&lt;/code&gt; (create) rule to the manager ClusterRole, scoped to the &lt;code&gt;default&lt;/code&gt;, &lt;code&gt;global&lt;/code&gt;, and &lt;code&gt;spokes&lt;/code&gt; cluster sets.&lt;/p&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;How to reproduce&lt;/h3&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;ol&gt;
&lt;li&gt;Deploy fleetconfig-controller with &lt;code&gt;fleetConfig.enabled: false&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Create a Hub CR in namespace A&lt;/li&gt;
&lt;li&gt;Recreate the Hub CR in namespace B (deleting the Hub runs &lt;code&gt;clusteradm clean&lt;/code&gt;, which removes the hub CRDs and garbage-collects the ManagedClusterSetBindings; recreating it brings the webhook back, but not the bindings)&lt;/li&gt;
&lt;li&gt;Restart the controller pod → it fails to recreate the bindings through the now-live webhook and enters CrashLoopBackOff&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="markdown-heading"&gt;
&lt;h3 class="heading-element"&gt;Testing&lt;/h3&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;[x] &lt;code&gt;make test-e2e&lt;/code&gt; passes with the new spec included.&lt;/li&gt;
&lt;li&gt;[x] &lt;code&gt;helm template&lt;/code&gt; / &lt;code&gt;helm lint&lt;/code&gt; pass.&lt;/li&gt;
&lt;li&gt;[x] Verified on a real cluster: applying new ClusterRole / ClusterRoleBinding recovered the controller from CrashLoopBackOff to Running.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="markdown-heading"&gt;
&lt;h2 class="heading-element"&gt;Summary by CodeRabbit&lt;/h2&gt;
&lt;span class="octicon octicon-link"&gt;&lt;/span&gt;
&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Bug Fixes&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Improved recovery after a controller restart by recreating required cluster topology bindings.&lt;/li&gt;
&lt;li&gt;Ensured the controller can establish the necessary permissions for cluster set binding operations.&lt;/li&gt;
&lt;li&gt;Improved rollout reliability by confirming updated and available controller replicas before reporting recovery complete.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Tests&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Added end-to-end coverage verifying topology resources are recreated and the controller becomes fully available after restart.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;


    &lt;/div&gt;
    &lt;div class="gh-btn-container"&gt;&lt;a class="gh-btn" href="https://github.com/open-cluster-management-io/lab/pull/249" rel="noopener noreferrer"&gt;View on GitHub&lt;/a&gt;&lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;



&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;That's the list of things that tripped me up while getting OCM / fleetconfig-controller running on GKE.&lt;/p&gt;

&lt;p&gt;I had done a fair amount of testing locally on kind beforehand and figured it would go smoothly, but there were more differences than I expected.&lt;/p&gt;

&lt;p&gt;Hopefully this helps anyone who is stuck, or about to get stuck, with the same setup. (If anyone out there is!)&lt;/p&gt;




&lt;ol&gt;

&lt;li id="fn1"&gt;
&lt;p&gt;&lt;a href="https://developers.redhat.com/articles/2023/01/19/how-distribute-workloads-using-open-cluster-management" rel="noopener noreferrer"&gt;How to distribute workloads using Open Cluster Management - Red Hat Developer Blog&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn2"&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/concepts/network-isolation#how_authorized_networks_work" rel="noopener noreferrer"&gt;About network isolation in GKE - Google Cloud Docs&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;li id="fn3"&gt;
&lt;p&gt;&lt;a href="https://kubernetes.io/docs/tasks/extend-kubernetes/custom-resources/custom-resource-definitions/#delete-a-customresourcedefinition" rel="noopener noreferrer"&gt;Extend the Kubernetes API with CustomResourceDefinitions - Kubernetes Documentation&lt;/a&gt;&amp;nbsp;↩&lt;/p&gt;
&lt;/li&gt;

&lt;/ol&gt;

</description>
      <category>kubernetes</category>
      <category>gke</category>
      <category>ocm</category>
      <category>multicluster</category>
    </item>
  </channel>
</rss>
