DEV Community

NTCTech
NTCTech

Posted on Originally published at rack2cloud.com

KMS Circular Dependency: The Cluster Came Back. The Keys Didn't.

A KMS circular dependency turns an online cluster into an unreadable one: the data needs a key, and the service that holds the key lives on the data. Hosts can be running, membership can be intact, and the storage can still be unable to complete the unlock path.

KMS circular dependency — key service hosted on the encrypted datastore it unlocks, with the unlock path blocked

What "Available" Hides

Availability and readability are different properties. In virtualization architecture, "available" usually describes the coordination layer: hosts joined, quorum held, services running. "Readable" describes something narrower: the cluster can obtain the key material it needs to decrypt the data it is holding. A cluster can have the first and lack the second.

The two-node edge deployment argument separates availability from capacity the same way: quorum says the cluster is up and says nothing about what that state costs. The gap here opens on a different axis. An encrypted cluster can be up and unable to read its own disks.

The scope is narrow, and it sits at the edge of an earlier comparison. Availability vs Authority scopes its Nutanix outcomes to clusters that don't depend on a key manager outside the cluster. This post covers the ones that do. The claim is that storage survivability is configuration-dependent, and that one function, unlock, shows it with vendor primary documentation behind it.

Cluster online versus dataset readable — two states of an encrypted storage cluster after a full restart

Five Storage Functions, One Populated Row

Between power-on and a usable datastore, a storage platform performs more than one function, and each function has its own dependency path. Five are worth separating.

Function Dependency question Status in this post
Unlock Does the key service stay reachable when the storage is cold? Worked below
Mount What must be present before the unlocked volume presents to hosts? Under investigation, not claimed
Discover What must resolve before hosts and VMs find the datastore? Under investigation, not claimed
Replicate What must be reachable for a replication link to authenticate and move data? Under investigation, not claimed
Validate What must be reachable before a restored state can be checked as usable? Under investigation, not claimed

Only the first row carries evidence here: unlock, where the KMS circular dependency lives. The other four are the same question asked of different functions. None is shown to fail the same way, and nothing below implies it.

This sits beside the Storage Survivability Boundary, Framework #109: the point where a storage architecture can no longer absorb component failure, rebuild activity, or locality loss without breaking its performance or availability assumptions. A cluster can sit well inside that boundary and still be unreadable, because the unlock dependency isn't part of the arithmetic the boundary models.

Five storage functions with unlock populated by evidence and mount, discover, replicate and validate marked not claimed

The KMS Circular Dependency: Unlock and Decrypt, Worked

The evidence standard for this section has three rungs: the dependency a vendor documents, the consequence it documents, and the mechanism it documents. Where the vendor stops, this post stops.

vSAN: The Circular Case Is Documented

Broadcom's KB 453198 (vSAN 8.x and 9.x) names the KMS circular dependency directly. It occurs when the KMS virtual machines sit on the same encrypted vSAN datastore they are responsible for unlocking. While the cluster is online and the keys are cached in host memory, Broadcom notes, the configuration appears functional. A full power outage or a host reboot clears that cache. At startup the ESXi hosts cannot mount the encrypted disk groups without a network connection to the KMS. The KMS virtual machines cannot power on, because the datastore underneath them is locked. Each side waits on the other.

The guidance for a KMS circular dependency follows from the mechanism: never locate the KMS on the datastore it unlocks. A new KMS installation that holds none of the original keys will not decrypt existing data, so recovery means restoring the original KMS or its keys, with restoring the KMS to non-encrypted storage as the preferred path. Without external backups of the KMS appliances or exported keys, the data is permanently inaccessible.

⚠ Do not reboot your way out: Broadcom's guidance is explicit: do not reboot hosts while key access is unavailable. Keys can remain cached in memory after a KMS outage, and a reboot destroys that cache and removes the remaining ability to access the data. Do not recreate disk groups or disable encryption to bypass the lock. Recreating disk groups can destroy the remaining possibility of recovering the data.

Two variants appear in Broadcom's vSAN 8.0 administration guide, which documents the baseline the same way: a rebooting host does not mount its disk groups until it receives the key-encryption key from the KMS. From vSAN 7.0 Update 3, key persistence lets hosts retain their keys across a reboot, in a TPM where one is present, so encryption operations can continue while the key server is unavailable. Native Key Provider, supported from vSAN 7.0 Update 2, needs no external KMS: vCenter Server generates the key material and pushes it to the hosts, which then generate the data encryption keys. KB 453198 describes the cached-key case and does not mention either variant. This post does not claim how they interact with the KMS circular dependency.

Nutanix: The Warning Is Documented, the Cycle Isn't

The AOS Security 7.6 Security Guide is the governing source for the Nutanix claims here. The relevant language is unchanged from AOS Security 6.8. The guide separates two encryption modes, and the evidence differs by mode.

For self-encrypting drives using an external KMS, the key service is outside the cluster and the Controller VM uses KMIP to upload and retrieve the drive keys. One KMS is required, several are recommended so the KMS is not a single point of failure, and the guide advises clustering them (p. 149). After a full power-off and power-on with protection enabled, the Controller VM retrieves the drive keys from the KMS to unlock the drives. If it cannot get the correct keys, it cannot access the data (p. 149).

Software-only encryption does not inherit that sentence. The guide gives it an access-condition list instead. The system accesses the external KMS when a cluster starts, when a key is regenerated, when the KMS type is switched, and when the NCC heartbeat checks that the KMS is alive. For software-only encryption specifically, it also accesses the KMS when a service starts or restarts and when AOS is upgraded (pp. 149–150).

For an external KMS in both modes, the guide carries an explicit caution: do not host a key management server VM on the encrypted cluster that is using it, because a problem with that VM could result in complete data loss (pp. 149, 164). That is a topology prohibition and a stated consequence. The guide does not describe the startup sequence that produces the loss.

Native KMS removes the external KMS dependency by design, within conditions. Native KMS (local) is software-only and needs at least three nodes. One- and two-node clusters are not supported for it (pp. 147, 160). Its keys must be backed up (p. 147). Native KMS (remote), meant for one- and two-node remote-office clusters, is available only when the cluster is registered to Prism Central (p. 160), and the guide notes that the root key is retained on Prism Central if the cluster is destroyed while still registered (p. 167). That documents a dependency on Prism Central registration. It does not document how or when keys are retrieved, and this post does not claim it.

What Each Vendor Documents, and Where It Stops

Both vendors warn against placing the key-management function on the encrypted storage it is responsible for unlocking. The documentation differs in how much of the failure it exposes.

Rung Broadcom (vSAN 8.x, 9.x) Nutanix (AOS Security 7.6)
Dependency documented Hosts need the KMS to mount encrypted disk groups at startup The CVM retrieves keys from an external KMS; access conditions listed by mode
Consequence documented Locked datastore; a new KMS without the original keys cannot decrypt; no external backups means permanent inaccessibility SED: the CVM cannot access data without the correct keys. Hosted KMS VM: complete data loss warning
Mechanism documented Yes: hosts cannot mount without the KMS, and the KMS VMs cannot power on from the locked datastore Not for a hosted KMS VM, and not for Native KMS (remote) retrieval

No ranking follows. A vendor that documents less of a failure has not been shown to fail differently, and a vendor that documents the whole cycle has not been shown to be more exposed to it.

The Dependencies Recovery Plans Forget names four dependency objects that recovery plans treat as assumptions: identity, DNS, certificates, and network. A key manager is not on that list, and the circular case is the version where the dependency is not only forgotten. It cannot start until the storage it unlocks is already readable. The System Recovered. Your Recovery Boundary Didn't. argues that recovery boundaries get drawn around what the organization owns instead of what the service depends on, which is a question about dependencies left outside the boundary. This one can sit inside the boundary and still fail, because of ordering.

Configuration Preconditions

Survivability here depends on a configuration state: where the key service runs relative to the storage it unlocks. Eight configurations follow, graded by what the evidence supports. The grade labels the evidence class. It does not mean recommended. One row establishes the external placement condition; another creates the KMS circular dependency outright. The remaining rows set the conditions around it.

Platform Configuration Cold-start / access dependency Evidence / Scope
vSAN External KMS, off the encrypted datastore Hosts need a network connection to the KMS to mount encrypted disk groups at startup PRIMARY — KB 453198 — vSAN 8.x, 9.x
vSAN KMS VM on the encrypted datastore it unlocks Creates the KMS circular dependency: unlock requires the KMS the datastore hosts PRIMARY — KB 453198 — vSAN 8.x, 9.x
vSAN Host key persistence with a TPM Hosts retain keys across a reboot and encryption operations continue while the key server is unavailable (vSAN 7.0 U3+); whether it eliminates or changes the circular failure path is not established here PRIMARY — vSAN 8.0 ADMIN GUIDE
vSAN Native Key Provider No external KMS: vCenter Server generates the key material and pushes it to the hosts (vSAN 7.0 U2+); whether it eliminates or changes the circular failure path is not established here PRIMARY — vSAN 8.0 ADMIN GUIDE
Nutanix External KMS, SED After a full power-off and power-on, the CVM retrieves drive keys from the KMS; without the correct keys it cannot access the data PRIMARY — SED — 7.6
Nutanix External KMS, software-only Guide lists KMS access at cluster start, key regeneration, KMS type switch and NCC heartbeat, plus service start/restart and AOS upgrade PRIMARY — SOFTWARE — 7.6
Nutanix Native KMS, local Software-only; minimum 3 nodes, 1- and 2-node clusters not supported; keys must be backed up PRIMARY — 3-NODE — 7.6
Nutanix Native KMS, remote Available only when the cluster is registered to Prism Central; retrieval mechanism not documented here PRIMARY — PRISM CENTRAL — 7.6

AOS 7.5+ also documents Prism Central-managed external KMIP for software-only encryption. After association, the cluster communicates directly with the KMS and Prism Central is not in the cluster-to-KMS data path (AOS Security 7.6, pp. 103–104). This is distinct from Native KMS (remote) and is not included as a separate configuration row here.

Diagnostic: "If every host in this cluster powered off at once, which service has to answer before the first datastore mounts, and where does that service run?"

What This Post Does Not Claim

  • The SED power-cycle sentence is not generalized to software-only encryption. Software-only is covered by its own access-condition list.
  • The Native KMS (remote) retrieval mechanism is not established. Prism Central registration is the documented dependency.
  • Cloud KMS, which the 7.6 guide documents for Azure Key Vault only (p. 108), is not treated as equivalent to a KMS VM hosted on the cluster, and is not worked here.
  • AOS 7.5+ Prism Central-managed external KMIP is distinct from Native KMS (remote).
  • Whether a KMS circular dependency forms in a software-only Nutanix cluster is not claimed. The guide states the prohibition and the consequence only.
  • vCenter-offline behavior on vSAN is not claimed.
  • Whether TPM key persistence or Native Key Provider eliminates or changes the circular failure path is not claimed.
  • Mount, discover, replicate, and validate are not claimed.
  • No platform is ranked, and nothing here extends beyond the documented versions and configurations: vSAN 8.x and 9.x per Broadcom KB 453198, and Nutanix AOS Security 7.6 per the Nutanix Security Guide. The relevant Nutanix language is unchanged from the 6.8 guide.

Recovery Readiness Assessment — test your own recovery path against this failure mode.

Architect's Verdict

Online and readable are different states, and encrypted storage is where confusing them costs the most. Both vendors' primary documentation says the same thing about the key service: it does not belong on the storage it unlocks.

The failure is not encryption by itself. It is a dependency nobody drew. In the configuration where a KMS circular dependency forms, the cluster that most needs the key service is the one that cannot start it. Broadcom documents that wait step by step. Nutanix documents the prohibition and the consequence and leaves the sequence unstated. A vendor that documents less has not been shown to fail differently.

Configuration decides the outcome. Same platform, same encryption, one placement decision apart, and one cluster restarts while the other waits on itself. The review question is not whether encryption is enabled. It is where the key service runs when everything is off, who can restore it, and from what.

A cluster that has come back has not recovered until it can read what it came back with.

Additional Resources

Originally published at rack2cloud.com

Top comments (0)