DEV Community

Cover image for Terraform Cloud Migration Playbook: A 30/60/90 Day Plan for Enterprise Teams
Yuvraj
Yuvraj

Posted on

Terraform Cloud Migration Playbook: A 30/60/90 Day Plan for Enterprise Teams

TL;DR

A Terraform Cloud migration can go quietly wrong in two ways: state divergence between the old and new backend, and a policy gap the moment Sentinel gets turned off before its OPA/Rego replacement is proven. The fix for both is the same: run old and new in parallel before cutting over, and decide rollback criteria in advance rather than mid-incident. This guide breaks the move into four phases across a 90-day window: assessment, state migration, policy migration, and cutover, each with a built-in retention period so a problem surfaces before it becomes irreversible.

Getting Started: What a Terraform Cloud Migration Actually Involves

A Terraform Cloud migration means moving three things off HCP Terraform or Terraform Enterprise onto a different platform: the state files that track what is actually deployed, the policy enforcement that governs what is allowed to deploy, and the workspaces and pipelines that tie the two together. It is not like migrating an application to a new host, where a bad migration usually just means recoverable downtime. A bad state migration means losing track of what is actually deployed, two backends disagreeing about which resources exist, or a policy engine going dark for the exact window nobody is watching closely enough to notice.

Two things go wrong specifically when this gets rushed. The first is state divergence, where the old and new backend both briefly think they are the source of truth, and a Terraform apply against either one silently overwrites the other's understanding of reality. The second is a policy gap: Sentinel and its OPA/Rego replacement are never functionally identical on day one, so disabling one before the other is proven leaves a window with no enforcement at all. Both risks get handled the same way across every phase below: run old and new in parallel before cutting over, and decide rollback criteria in advance rather than improvising mid-incident.

This playbook assumes a target platform is already picked, whether that is a managed Terraform-only backend like Scalr or Spacelift, a self-hosted GitOps setup like Atlantis, a multi-IaC lifecycle platform like env0, or a governance and inventory layer like Firefly. That decision depends on IaC tool mix, execution model, pricing predictability, and governance scope, which sits outside the scope of this piece. What follows assumes that decision is already made and focuses entirely on executing the move safely.

Phase 0: Assessment, Week 1 to 2

Before any file moves, the team needs an honest inventory of current TFC or TFE usage:

  • How many workspaces exist, categorized by team, environment, and change frequency
  • Where state is stored: TFC-managed backend or remote backends like S3, GCS, or Azure Blob
  • Which Sentinel or OPA policies are active, with each rule's enforcement level and what it blocks documented
  • What VCS integrations are in use (GitHub, GitLab, Azure DevOps)
  • What notification and alerting hooks exist
  • Whether Terraform Cloud Agents are in use for private network access
  • What secrets and variable sets need to be migrated

This step is easy to underestimate. A team that thinks it has about 40 workspaces frequently finds closer to 60 once workspaces created for one-off testing or abandoned proofs of concept get counted, and every one of those needs a decision- migrate, archive, or delete- before Phase 1 starts, not during it.

Once the inventory is done, the target category shapes what comes next. Teams that are Terraform or OpenTofu only and want a managed backend tend toward Scalr or Spacelift; GitOps PR automation with self-hosting points toward Atlantis. Multi-IaC plus lifecycle and FinOps needs point toward env0. Unified inventory, cross-tool governance, and drift detection point toward Firefly.

The workspace inventory above is manual by design; someone has to count and categorize. An independent baseline scan run alongside that manual pass, connected read-only to the same cloud accounts TFC manages, can surface every resource tagged by IaC status regardless of whether a human remembered it exists. That baseline percentage- what share of the environment is actually codified today- is worth capturing before migration starts, since it is the number a team should be able to point to improving by day 90, not just "we moved platforms."

Phase 1: State and Backend Migration, Days 1 to 30

State migration follows a consistent sequence. Pull the current state files for each workspace with terraform state pull > backup.tfstate, decide on a state destination (a managed backend on the new platform, or bring your own via S3 and DynamoDB, Azure Blob and Table, or GCS), and update the backend configuration in the .tf files accordingly. From there, run terraform init -migrate-state or terraform init -reconfigure in each workspace, validate by running terraform plan and confirming zero unexpected changes, enable state locking on the new backend before any team member switches over, and retain the TFC workspace in read-only mode for at least two weeks as a rollback window.

That two-week read-only retention step is the one teams most often skip under time pressure, and it is the one worth protecting above the others. Consider a workspace with a database resource carrying a prevent_destroy lifecycle rule. If the migration subtly drops that lifecycle block during backend reconfiguration, a Terraform plan run immediately after can still come back clean, since the plan comparison only checks against the new state file, not against what the resource actually depended on before. The gap does not surface until someone runs an apply that would have hit the missing safeguard, which might be weeks later.

That Terraform plan validation step has a structural blind spot worth naming directly: it only compares the new state file against the .tf configuration, not against what is actually running in the cloud. If the migration itself introduced drift, a tag lost, a lifecycle rule dropped, an attribute silently changed during backend reconfiguration, a clean Terraform plan will not catch it, because it is not checking live provider state at all. A drift detection scan run immediately after each workspace migrates acts as a second, independent validation layer on top of Terraform plan, catching exactly the class of error a plan, which only checks structurally, cannot see.

Phase 2: Policy Migration, Days 15 to 45

Policy migration runs in parallel with the tail end of Phase 1. It starts with exporting all Sentinel policies and policy sets from TFC, then mapping each Sentinel rule to an equivalent OPA/Rego policy. Each rule needs a placement decision, pre-plan, post-plan, approval gate, or periodic scan, and each translated policy should be tested against a known good plan JSON before it goes anywhere near production, using something like conftest test plan.json.

From there, policies deploy to the new platform in advisory mode first, meaning they warn rather than block. Sentinel stays active on TFC while OPA runs in parallel on the new platform. After two weeks of clean parallel runs, OPA switches to enforcement mode, and Sentinel gets deactivated on TFC.

The advisory mode step exists for a specific reason. A Sentinel rule and its OPA/Rego translation are rarely byte-for-byte equivalent on the first attempt; edge cases in how each engine evaluates a plan JSON can differ in ways that only show up against real traffic, not a handful of test cases. Running the new policy in warn-only mode for two weeks surfaces those mismatches while Sentinel is still the actual enforcement layer underneath.

A Sentinel rule blocking public S3 buckets, translated into an OPA/Rego policy, looks like this:

| package Cx

import data.generic.terraform as tf\_lib

CxPolicy\[result\] {
    resource := input.document\[i\].resource.aws\_s3\_bucket\_public\_access\_block\[name\]
    resource.block\_public\_acls \== false

    result := {
        "documentId": input.document\[i\].id,
        "resourceType": "aws\_s3\_bucket\_public\_access\_block",
        "resourceName": tf\_lib.get\_resource\_name(resource, name),
        "searchKey": sprintf("aws\_s3\_bucket\_public\_access\_block\[%s\].block\_public\_acls", \[name\]),
        "issueType": "IncorrectValue",
        "keyExpectedValue": sprintf("aws\_s3\_bucket\_public\_access\_block\[%s\].block\_public\_acls should be true", \[name\]),
        "keyActualValue": sprintf("aws\_s3\_bucket\_public\_access\_block\[%s\].block\_public\_acls is set to false", \[name\]),
        "remediation": json.marshal({
            "before": "false",
            "after": "true",
        }),
        "remediationType": "replacement",
     |
Enter fullscreen mode Exit fullscreen mode

Phase 3: Cutover, Days 45 to 90

Cutover is where the earlier parallel running pays off. The TFC workspace configuration gets frozen: no new variables, no new policies. Parallel deploys push the same change through TFC and the new platform, comparing plan output directly. A drift validation period runs at least one full scan cycle on the new platform to confirm drift detection matches expectations. CI/CD VCS integrations migrate to point at the new platform, team documentation and runbooks get updated, and TFC workspaces get decommissioned after a 30-day post-migration clean period.

The rollback criteria step deserves to be written down explicitly rather than left as a judgment call. "Revert if something breaks" is not a criterion; it is a description of every incident. A usable version looks more like: revert if drift detection on the new platform flags more than a defined percentage of resources as unexpectedly changed within the first 72 hours, or if a policy enforcement gap lets through a change that Sentinel would have blocked.

How Firefly Fits Into Each Migration Phase

Firefly is not a migration destination in the sense of replacing TFC's execution role, but its inventory, drift detection, and Guardrails engine map onto specific checklist items across all four phases above, and the mapping is mechanical rather than a generic feature list bolted on afterward.

Phase 0: Establishing an Independent Codification Baseline

Firefly's Cloud Asset Inventory can run alongside the manual workspace count and produce an independent baseline. Connected read-only to the same cloud accounts TFC manages, it surfaces every resource tagged by IaC status, Codified, Drifted, or Unmanaged, regardless of whether a human remembered the workspace exists. That baseline codification percentage is the number worth capturing on day one.

Phase 1: Catching Drift That Terraform Plan Cannot See

Firefly's drift detection directly closes the gap left by Terraform plan only checking the new state file against configuration. Running an inventory scan immediately after each workspace migration catches drift introduced during backend reconfiguration that a plan comparison structurally cannot see. Firefly also supports both state destinations from the Phase 1 checklist natively: bring your own S3, Azure Blob, or GCS, or Firefly-managed, so state ownership does not have to change again if the target platform changes later.

A Guardrail check confirming state locking is actually enabled on the new backend, closing that Phase 1 checklist item automatically rather than trusting it got done manually, looks like this:

 package Cx

import data.generic.terraform as tf\_lib

CxPolicy\[result\] {
    resource := input.document\[i\].resource.aws\_dynamodb\_table\[name\]
    not resource.point\_in\_time\_recovery

    result := {
        "documentId": input.document\[i\].id,
        "resourceType": "aws\_dynamodb\_table",
        "resourceName": tf\_lib.get\_resource\_name(resource, name),
        "searchKey": sprintf("aws\_dynamodb\_table\[%s\].point\_in\_time\_recovery", \[name\]),
        "issueType": "MissingAttribute",
        "keyExpectedValue": sprintf("aws\_dynamodb\_table\[%s\] used for Terraform state locking should have point\_in\_time\_recovery enabled", \[name\]),
        "keyActualValue": sprintf("aws\_dynamodb\_table\[%s\] has no point-in-time recovery configured", \[name\]),
        "remediation": json.marshal({
            "before": "point\_in\_time\_recovery undefined",
            "after": "point\_in\_time\_recovery { enabled \= true }",
        }),
        "remediationType": "addition",
    }
Enter fullscreen mode Exit fullscreen mode

In Phase 2, a Sentinel rule can be translated into Firefly's policy engine through four paths depending on complexity: a pre-built Policy Pack for common controls, the no-code rule builder for straightforward attribute or tag checks, hand-written Rego for anything genuinely custom, or AI-assisted generation from a plain English description of what the Sentinel rule was doing, validated in a testing playground before it goes anywhere near production. The advisory mode checklist step maps directly onto a real Guardrail setting rather than a workaround; every policy violation can be configured to block the deployment, alert an administrator without blocking, or allow an override for authorized users in specific circumstances. Setting a newly translated policy to alert only during the two-week parallel run, then flipping it to block once proven clean, is the same mechanism the checklist describes. Firefly's five severity levels, Info, Low, Medium, High, and Critical, give this a practical rollout order too: start Critical severity translated rules in alert-only mode with the tightest review, and let Info or Low severity rules move to enforcement faster, since the blast radius of an early mismatch is smaller.

In Phase 3, the checklist's own rollback criteria example, more than a defined percentage of resources unexpectedly changed within 72 hours, is directly answerable from a single, attributed mutation log rather than assembled by hand from scattered logs. Every mutation during the cutover window, whether it came through the console, CLI, the new platform's pipeline, or a stray script someone ran to unblock something quickly, gets logged with the identity responsible, attributed via cloud audit logs like AWS CloudTrail or Git blame for IaC-driven changes. The 72-hour rollback check becomes a query against one source instead of a manual cross-reference across TFC's logs, the new platform's logs, and CloudTrail separately.

firefly

A Guardrail can also enforce the parallel deploy comparison step directly, blocking a cutover window plan that touches an unusually large blast radius without explicit sign-off:

package Cx

import data.generic.terraform as tf_lib

CxPolicy[result] {
    count(input.document[i].resource_changes) > 15
    not input.document[i].metadata.cutover_approved

    result := {
        "documentId": input.document[i].id,
        "resourceType": "plan",
        "issueType": "IncorrectValue",
        "keyExpectedValue": "Plans touching more than 15 resources during the cutover window require explicit cutover_approved metadata",
        "keyActualValue": sprintf("Plan touches %v resources without cutover approval", [count(input.document[i].resource_changes)]),
        "remediationType": "manual_review",
    }
}
Enter fullscreen mode Exit fullscreen mode

Where Should You Start

This playbook's core argument is that migration risk concentrates in two places: state divergence during backend cutover and policy enforcement gaps during the Sentinel to OPA transition, and both are manageable with the same pattern repeated across every phase: run old and new in parallel, and decide rollback criteria before an incident forces the decision under pressure. The 90-day window and its built-in retention periods- two weeks of read-only state access, two weeks of policy advisory mode, 30 days of post-cutover workspace retention, are not padding. They turn an irreversible mistake into a caught one.

Start with the Phase 0 inventory this week, even if the target platform decision is not fully locked yet. Knowing the real workspace count, the active policy set, and a baseline codification percentage shapes every phase that follows, and it is the one step that gets harder to do accurately the further into migration a team gets. For teams evaluating where Firefly's inventory and drift detection layer fits alongside a chosen execution platform, the Cloud Asset Inventory documentation covers the read-only connection setup, and the Policy & Governance guide walks through the four policy authoring paths referenced above.

Frequently Asked Questions

How long does a Terraform Cloud migration actually take?

For most enterprise environments, budget the full 90-day window this playbook outlines. Smaller environments with fewer workspaces and simpler policy sets can compress this. The parallel operation windows, two weeks minimum for both state validation and policy advisory mode, are safety mechanisms worth keeping even on a faster timeline.

What is the biggest risk during cutover specifically?

State divergence and policy gaps are the two structural risks, but during cutover specifically, the most common practical failure is an undefined rollback trigger. Teams that have not decided in advance what conditions justify reverting to TFC tend to either revert too late, after a real incident, or not at all, talking themselves out of a legitimate rollback because no clear criteria exist.

Can Terraform state be migrated without downtime?

Yes, if state locking is enabled on the new backend before any team member starts applying against it, and if the old backend is frozen to read-only rather than left active in parallel with write access. The risk is not downtime; it is two backends both accepting writes for the same workspace simultaneously.

Do all workspaces need to migrate at the same time?

No. For larger environments, migrating in batches by team or environment is usually safer than a single cutover across every workspace. The four-phase structure in this playbook can run per batch, with earlier batches informing what to adjust before later ones start.

What is the difference between Sentinel and OPA/Rego for policy enforcement?

Sentinel is HashiCorp's proprietary policy-as-code framework, built specifically for Terraform Cloud and Terraform Enterprise. OPA, or Open Policy Agent, is an open-source, general-purpose policy engine using the Rego language, usable across many platforms beyond Terraform. Migrating off TFC typically means translating existing Sentinel rules into Rego equivalents, since most alternative platforms do not support Sentinel natively.

Why does a clean Terraform plan after migration not guarantee nothing broke?

Because terraform plan only compares the new state file against the .tf configuration, it does not check what is actually running in the cloud. If the migration process itself introduced drift, a dropped lifecycle rule, or a changed attribute, the plan comparison has no way to detect it, since it never queries live provider state independently.

Does moving off Terraform Cloud mean losing policy enforcement entirely during the transition?

Not if the parallel run pattern is followed. Keeping Sentinel active on TFC while the OPA equivalent runs in advisory mode on the new platform means enforcement never actually goes dark; the two-week overlap exists specifically to catch translation mismatches before Sentinel gets switched off.

Top comments (0)