DEV Community

Kenta Goto for AWS Heroes

Posted on

4 Ways to Safely Refactor AWS CDK Code Without Resource Replacement

In AWS CDK, refactoring such as extracting resource definitions into a custom construct can cause resources to be replaced at deployment time — that is, deleted and recreated. For stateful resources, this means the data they held is lost.

In this article, I will introduce four ways to refactor safely without triggering this replacement.

NOTE: This article is based on AWS CDK (aws-cdk-lib) v2.263.0 and AWS CDK CLI (aws-cdk) v2.1139.0. To use the cdk refactor and cdk orphan commands introduced here, not only the CDK CLI but also the bootstrap stack must support both commands. If these commands result in an error, update the CDK CLI and then re-run cdk bootstrap to update the bootstrap stack.


Construct Tree and Path

Let's start with the mechanism by which refactoring causes resource replacement. When you define a resource in AWS CDK, you write something like new Queue(this, 'Queue'), passing the parent construct as the first argument and a string called the construct ID as the second. The first argument is usually this, which points to the stack or construct being defined. The construct ID is a string that identifies a construct among those under the same parent. The hierarchy in which constructs are connected as parents and children, with the App at the top, is called the construct tree, and a position in the tree is represented by a path such as MyStack/Queue, formed by joining the construct IDs from the parent down.

Example of a construct tree


Code Before Refactoring

Now let's look at an example where replacement actually occurs. Consider the following code, which defines three resources directly under the stack: an Amazon SQS queue, an Amazon SNS topic, and an AWS Lambda function. The queue holds notification messages waiting to be processed, and deleting it would affect the running system. The topic, on the other hand, is a stateless resource, and we will assume its replacement is acceptable. Note that Code.fromAsset, which specifies the Lambda function code, deploys the contents of the local lambda directory as the function code.

export class MyStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    new Queue(this, 'Queue');

    new Topic(this, 'Topic');

    new Function(this, 'Function', {
      runtime: Runtime.NODEJS_24_X,
      handler: 'index.handler',
      code: Code.fromAsset('lambda'),
    });
  }
}
Enter fullscreen mode Exit fullscreen mode

Code After Refactoring

Of these three resources, we extract the queue and the topic, which make up the notification feature, into a custom construct called Notifications, and have the stack instantiate it. The Lambda function stays directly under the stack. The actual AWS resource configuration to be deployed remains the same.

export class Notifications extends Construct {
  constructor(scope: Construct, id: string) {
    super(scope, id);

    new Queue(this, 'Queue');

    new Topic(this, 'Topic');
  }
}

export class MyStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    new Notifications(this, 'Notifications');

    new Function(this, 'Function', {
      runtime: Runtime.NODEJS_24_X,
      handler: 'index.handler',
      code: Code.fromAsset('lambda'),
    });
  }
}
Enter fullscreen mode Exit fullscreen mode

Logical ID Changes and Resource Replacement

AWS CDK automatically generates the logical IDs in the AWS CloudFormation template based on paths. For example, the Queue before refactoring gets a logical ID like Queue4A7E3555, a path-derived string with a hash appended. In other words, when refactoring changes a construct's position in the tree or its construct ID, the logical ID changes along with the path.

In our example as well, the path of Queue changes from MyStack/Queue to MyStack/Notifications/Queue, and its logical ID changes from Queue4A7E3555 to NotificationsQueue91395D8F. Similarly, Topic changes from TopicBFC7AF6E to NotificationsTopicAE679CBD. Meanwhile, the Lambda function left directly under the stack keeps its path, so its logical ID stays the same.

CloudFormation identifies resources by their logical IDs. Therefore, if you deploy in this state, the queue and the topic, whose logical IDs have changed, are treated as different resources: they are deleted and created anew. We assumed the topic's replacement was acceptable, but the notification messages accumulated in the queue would be lost.

Logical ID changes caused by refactoring


Four Ways to Refactor Safely

To prevent this replacement, I will introduce the following four methods. All of them carry the deployed resources over to the refactored code without recreating them.

  1. The cdk refactor command
  2. The cdk orphan + cdk import commands
  3. Setting the construct ID to Default
  4. Overriding the logical ID

1. The cdk refactor Command

The first method is the cdk refactor command. When you run it, resources whose positions in the construct tree have changed are detected and moved to their new positions without being recreated. It supports moves both within the same stack and to a different stack. It can also apply multiple resource moves at once, as in our example, so not only the queue but also the topic is carried over without replacement.

1-1. How to Use cdk refactor

After refactoring your code, run cdk refactor before deploying. With the --dry-run flag, it only displays the list of resources detected as moved, without applying anything. Once you confirm that the moves are as intended, run the command again without --dry-run. After you answer the confirmation prompt, the moves are applied, and finally a deployment runs automatically to align template details, such as path metadata, with your code. As described later, cdk refactor fails when the changes include resource additions, removals, or property changes, so this deployment modifies neither the moved resources nor any other resources.

NOTE: The command requires the --unstable=refactor flag, and its options and behavior may change in the future.

Here is an example run with --dry-run:

$ npx cdk refactor --unstable=refactor --dry-run

The following resources were moved or renamed:

┌─────────────────┬────────────────────────┬──────────────────────────────────────┐
│ Resource Type   │ Old Construct Path     │ New Construct Path                   │
├─────────────────┼────────────────────────┼──────────────────────────────────────┤
│ AWS::SQS::Queue │ MyStack/Queue/Resource │ MyStack/Notifications/Queue/Resource │
├─────────────────┼────────────────────────┼──────────────────────────────────────┤
│ AWS::SNS::Topic │ MyStack/Topic/Resource │ MyStack/Notifications/Topic/Resource │
└─────────────────┴────────────────────────┴──────────────────────────────────────┘
Enter fullscreen mode Exit fullscreen mode

1-2. Cases Where cdk refactor Cannot Be Used

While convenient, the command cannot be used in cases like the following:

  • Moves that involve resource additions, removals, or property changes
  • Moves of resources whose properties contain strings auto-generated from the construct tree path (such as the family name of an Amazon ECS task definition or the origin ID of an Amazon CloudFront distribution)
  • Moves of unsupported resources (such as custom resources)

Regarding the first case, cdk refactor handles only resource moves. If other changes are involved, split your code changes into two steps: first perform only the move in your code and apply it with cdk refactor, then add resources or change properties and apply those with cdk deploy.

The second case actually comes down to the same reason as the first. A path-derived string changes its value along with the path. As a result, even if your code does nothing but move the resource, the change is treated as a move involving a property change.

As for the third case, the CloudFormation stack refactoring feature that cdk refactor uses internally has unsupported resources. A typical example is custom resources, and L2 constructs that use custom resources internally, such as BucketDeployment, cannot be moved either.

The error displayed when the changes involve more than moves is as follows:

A refactor operation cannot add, remove or update resources.
Only resource moves and renames are allowed.
Enter fullscreen mode Exit fullscreen mode

2. The cdk orphan + cdk import Commands

The second method is the combination of the cdk orphan and cdk import commands, which works even in configurations where cdk refactor cannot be used. The flow is to temporarily detach the resource you want to protect from replacement from the stack, and then bring it back into its new position after refactoring. The steps are as follows:

  1. Use cdk orphan to detach the resource you want to protect from replacement from the stack
  2. Refactor your code
  3. Use cdk import to bring the detached resource into its new position
  4. Apply the remaining changes with cdk deploy

2-1. Running cdk orphan

cdk orphan is a command that removes a resource from the stack's management without deleting it. Internally, it sets the CloudFormation DeletionPolicy to Retain so that the target resource is not deleted, replaces references from other resources with actual values such as the queue URL or ARN, and then removes the resource from the template. As a result, the resource leaves the stack's management but continues to exist in your AWS account.

NOTE: The command requires the --unstable=orphan flag, and its options and behavior may change in the future.

Specify the target resource by its path and run the command as shown below. It displays the list of resources to be detached, and the process starts after you answer the confirmation prompt. In our example, the queue is the only resource we want to protect from replacement, so we specify only the queue. We leave the topic attached and let the final cdk deploy in the steps replace it.

$ npx cdk orphan --unstable=orphan MyStack/Queue

Stack: MyStack
Resources to orphan (1):
  Queue4A7E3555 (AWS::SQS::Queue) - /MyStack/Queue/Resource
Do you wish to orphan these resources? This will perform 3 CloudFormation deployments. (y/n) y

...

✅ Resources orphaned from MyStack
Enter fullscreen mode Exit fullscreen mode

2-2. Running cdk import

Once cdk orphan is complete, refactor your code and then run cdk import. A prompt asks for the identifier of each resource to import: enter the URL for the queue, and for the topic, press Enter without typing anything to skip it. This brings the detached queue into the stack under its new logical ID while keeping its messages. Here is an example run:

$ npx cdk import

...

Perform import? (y/n) y

MyStack/Notifications/Queue/Resource (AWS::SQS::Queue): enter QueueUrl (empty to skip) https://sqs.us-east-1.amazonaws.com/123456789012/MyStack-Queue4A7E3555-pNg4kbUdhvRE
MyStack/Notifications/Topic/Resource (AWS::SNS::Topic): enter TopicArn (empty to skip)
Skipping import of MyStack/Notifications/Topic/Resource

 ✅  MyStack
Some resources were skipped. Run another cdk import or a cdk deploy to bring the stack up-to-date with your CDK app definition.
Enter fullscreen mode Exit fullscreen mode

2-3. Final cdk deploy and Points to Note

Finally, run cdk deploy to apply the remaining changes, including the topic's replacement, and the whole procedure is complete.

Note that this method has the following three caveats:

  • cdk import can only bring in resource types that support CloudFormation resource import
  • Always finish refactoring your code before running cdk import
  • Do not run a regular deployment between cdk orphan and cdk import

Regarding the first point, the list of supported resource types is available in the CloudFormation resource import documentation.

As for the second point, cdk import synthesizes your local code to determine where to import resources. If you run it with the pre-refactoring code, the resource is imported back into its original position.

Regarding the third point, if you deploy in between, the queue definition in your code is created as a new, empty resource separate from the existing queue.


3. Setting the Construct ID to Default

The third method is to set the resource's construct ID to Default inside the custom construct you extract to. Since logical IDs are generated from paths, the idea is to assign construct IDs so that the path stays the same after refactoring. Unlike the two methods so far, no additional commands are needed.

3-1. L1 Constructs and Paths

As background for this method, let's take a closer look at the path of the queue in our example. Inside an L2 construct such as Queue, an L1 construct that corresponds one-to-one with the CloudFormation resource is actually defined, with the construct ID Resource. Therefore, the actual path used to calculate the logical ID is MyStack/Queue/Resource, and Queue4A7E3555 was generated from it. This is also why the paths in the example run in 1-1 ended with /Resource.

3-2. Keeping the Logical ID with Default

When AWS CDK determines a logical ID, it excludes the string Default from the path components used in the calculation. Taking advantage of this behavior, set the construct ID of the queue whose logical ID you want to keep to Default inside the extracted custom construct. Then, give the Notifications construct the same construct ID as the pre-extraction queue, namely Queue.

With this, the path of the queue's L1 construct becomes MyStack/Queue/Default/Resource. Since Default is excluded from the calculation, the logical ID remains Queue4A7E3555, the same as before the extraction. All that's left is to run cdk deploy as usual. In this deployment, nothing happens to the queue, whose logical ID is unchanged, and only the other changes, such as the topic's replacement, are applied. The code is as follows:

export class Notifications extends Construct {
  constructor(scope: Construct, id: string) {
    super(scope, id);

    // The path becomes MyStack/Queue/Default/Resource, so the logical ID does not change
    new Queue(this, 'Default');

    // The topic's logical ID changes, but we accept its replacement
    new Topic(this, 'Topic');
  }
}

// Give the construct the same ID as the pre-extraction queue, 'Queue'
new Notifications(this, 'Queue');
Enter fullscreen mode Exit fullscreen mode

3-3. Conditions for Applying This Method

To apply this method, the following two conditions must be met:

  • Only one resource inside the extracted custom construct needs to keep its logical ID
  • The construct ID of the extracted custom construct is the same as the original ID of the resource or construct being moved

Regarding the first condition, only one Default can exist under the same parent. In the example above as well, the queue is the only resource whose logical ID is kept, and the topic's replacement is accepted as assumed. On the other hand, you can have any number of resources that do not need to keep their logical IDs inside the construct. Along with newly added resources, you can also move existing resources whose replacement is acceptable, like the topic in our example.

As for the second condition, in the example above we gave Notifications the construct ID Queue, so this method cannot be used if you want to give the construct a different name as part of the extraction.

This method is effective for extractions centered on a single resource you want to protect from replacement. The relationship between paths and logical IDs so far is summarized below.

Paths and logical IDs when the construct ID is set to Default


4. Overriding the Logical ID

The last method is to pin the logical ID, which AWS CDK normally determines automatically, to its pre-refactoring value. As long as the logical ID does not change, CloudFormation sees the resource as the same one. So all you need to do is apply the refactored code with a regular cdk deploy. During the deployment, nothing happens to the resources whose logical IDs are pinned, and only the other changes are applied. Since this method also works for resource types not supported by resource import, it serves as a last resort when the other methods cannot be used. There are two ways to pin the logical ID.

4-1. Pinning with overrideLogicalId

AWS CDK has a mechanism called an escape hatch, which lets you take out the L1 construct inside an L2 construct and manipulate it directly. As shown below, retrieve the L1 construct through defaultChild of the node property that every construct has, and pass the pre-refactoring logical ID to overrideLogicalId.

export class Notifications extends Construct {
  constructor(scope: Construct, id: string) {
    super(scope, id);

    const queue = new Queue(this, 'Queue');

    // Pin the logical ID to its pre-refactoring value (the topic is not pinned, accepting its replacement)
    const cfnQueue = queue.node.defaultChild as CfnQueue;
    cfnQueue.overrideLogicalId('Queue4A7E3555');

    new Topic(this, 'Topic');
  }
}
Enter fullscreen mode Exit fullscreen mode

4-2. Pinning with renameLogicalId

The other way is the stack's renameLogicalId method. This one operates on the logical ID itself, without going through constructs. Therefore, it also works for resources that cannot be retrieved with defaultChild, such as related resources that an L2 construct generates internally.

Use it in the following order. First, run cdk synth once on the code with only the refactoring applied, and check the newly generated logical ID in the output template. Then, as shown below, add a renameLogicalId call with the logical ID you checked as the first argument and the pre-refactoring logical ID you want to keep as the second.

export class MyStack extends Stack {
  constructor(scope: Construct, id: string, props?: StackProps) {
    super(scope, id, props);

    // Resource definitions omitted

    this.renameLogicalId('NotificationsQueue91395D8F', 'Queue4A7E3555');
  }
}
Enter fullscreen mode Exit fullscreen mode

4-3. Downside: The Pinning Code Remains

The downside of this method is that the code pinning the logical ID remains after the refactoring. Since it leaves a value that AWS CDK normally determines automatically explicitly specified, it is a good idea to leave a comment explaining why the ID is pinned.


How to Choose Among the Methods

If no resource you want to protect from replacement is moved, no countermeasure is needed and you can simply deploy. If one is needed, first consider cdk refactor, which completes with a single command. If your configuration does not allow it, consider cdk orphan + cdk import. If that is not an option either, use the construct ID Default approach when the extraction you want to perform meets the conditions in 3-3, and otherwise use logical ID overriding as a last resort.

Decision flow for choosing a method


Appendix 1. Accepting Replacement of Stateless Resources

As mentioned earlier, you do not need to prevent replacement for every resource. So which resources can tolerate replacement? Stateless resources, such as Lambda functions and the SNS topic in our example, hold no data, so replacing them is generally not a problem. What mainly needs protection are resources that lose their data when replaced, such as Amazon DynamoDB tables, Amazon S3 buckets, and the queue in our example. Accepting replacement for everything else greatly reduces the effort of refactoring.


Appendix 2. Leaving Physical Names Unspecified

When you accept replacement, be careful with how resource physical names are handled. A physical name is the actual resource name on AWS, such as a bucket name or a function name. In a CloudFormation replacement, the new resource is created first, and the old one is deleted afterward. Consequently, if your code specifies a physical name, the original resource with the same name still exists at creation time, and the deployment fails with an "already exists" error.

The basic way to avoid this is not to specify physical names. If omitted, a unique name is generated automatically, so names never collide during replacement. Leaving physical names unspecified also gives you more flexibility for refactoring.


Conclusion

In this article, I introduced the following four methods for refactoring safely in AWS CDK:

  1. The cdk refactor command
  2. The cdk orphan + cdk import commands
  3. Setting the construct ID to Default
  4. Overriding the logical ID

Knowing these methods lets you clean up your code with confidence, even for stacks that are already released. Alongside them, accepting the replacement of stateless resources and leaving physical names unspecified are also effective policies. I hope you will incorporate these techniques into your daily development.

Top comments (0)