DEV Community

rubendob
rubendob

Posted on

Terraform: Checksum Mismatch Between S3 and DynamoDB

Recently, I ran into an interesting issue with a Terraform remote backend.

Out of nowhere, every Terraform operation (plan, refresh, or apply) started failing with an error related to the remote state.

Terraform was reporting a checksum mismatch between the state stored in S3 and the value stored in DynamoDB.

The error looked like this:

Error: state data in S3 does not have the expected content

The checksum calculated for the state stored in S3 does not match the checksum
stored in DynamoDB.

Calculated checksum: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
Stored checksum:     YYYYYYYYYYYYYYYYYYYYYYYYYYYYYYYY
Enter fullscreen mode Exit fullscreen mode

The Problem

In this setup, Terraform was using:

  • An S3 bucket to store the terraform.tfstate file.
  • A DynamoDB table to manage locking and metadata associated with the remote state.

When Terraform accesses the backend, it:

  1. Reads the state stored in S3.
  2. Calculates its checksum.
  3. Compares that value with the Digest stored in DynamoDB.

In our case, those two values did not match.

Because of this inconsistency, Terraform refused to continue in order to avoid working with a state that could potentially be corrupted or outdated.

Initial Checks

Before changing anything, we went through a few basic checks:

  • Confirm that no terraform apply operation was currently running.
  • Confirm that no CI/CD pipeline was using the same remote state.
  • Reproduce the issue both locally and from the CI/CD pipelines.

Since the same error appeared regardless of where Terraform was executed, we could quickly rule out a problem with the local machine or the pipeline configuration.

Finding the Mismatch

The next step was to inspect the DynamoDB table associated with the remote backend.

Using the AWS CLI, we searched for the entry related to the affected state:

aws dynamodb scan \
  --table-name <lock-table>
Enter fullscreen mode Exit fullscreen mode

Among the results, we found an entry similar to this:

{
  "LockID": {
    "S": "<state-path>-md5"
  },
  "Digest": {
    "S": "old-value"
  }
}
Enter fullscreen mode Exit fullscreen mode

The value stored in Digest matched exactly the checksum Terraform was reporting as incorrect.

This confirmed that the current state stored in S3 and the digest registered in DynamoDB were out of sync.

The Fix

Terraform was already telling us which checksum it expected:

Calculated checksum: correct-value
Stored checksum:     old-value
Enter fullscreen mode Exit fullscreen mode

After verifying that the state file stored in S3 was valid, and confirming that no Terraform operation was currently running, we manually updated the Digest field in DynamoDB so that it matched the checksum calculated by Terraform.

For example:

aws dynamodb update-item \
  --table-name <lock-table> \
  --key '{"LockID":{"S":"<state-path>-md5"}}' \
  --update-expression "SET Digest = :d" \
  --expression-attribute-values '{":d":{"S":"correct-value"}}'
Enter fullscreen mode Exit fullscreen mode

This type of manual change should be done carefully.

You should not update the Digest just to make the error disappear. First, make sure that the state stored in S3 is the correct one and that there are no concurrent Terraform operations running against the same backend.

Result

After updating the Digest, we ran:

terraform plan
Enter fullscreen mode Exit fullscreen mode

Terraform was able to access the remote backend again, and the plan completed successfully.

We did not need to:

  • Recreate the state.
  • Restore a previous state version.
  • Modify the Terraform code.
  • Recreate any infrastructure resources.

The issue was simply a mismatch between the checksum stored in DynamoDB and the actual contents of the state file stored in S3.

Lessons Learned

When you see an error like:

state data in S3 does not have the expected content
Enter fullscreen mode Exit fullscreen mode

it is easy to immediately assume that the terraform.tfstate file is corrupted.

However, that is not always the case.

The state stored in S3 may be perfectly valid, while the inconsistency exists only in the Digest value stored in DynamoDB.

Before making any changes, I recommend checking the following:

  1. Confirm that no Terraform executions are currently running.
  2. Confirm that no CI/CD pipeline is using the same backend.
  3. Verify the state stored in S3.
  4. Inspect the corresponding entry in DynamoDB.
  5. Compare Terraform's Calculated checksum with the Stored checksum.

Only after completing those checks should you consider manually correcting the Digest.

In this case, a few minutes spent understanding how the remote backend works saved us from wasting much more time debugging Terraform code that was not actually the problem.

Links

Top comments (0)