Recently, I ran into an interesting issue with a Terraform remote backend.
Out of nowhere, every Terraform operation (plan, refresh, or apply) started failing with an error related to the remote state.
Terraform was reporting a checksum mismatch between the state stored in S3 and the value stored in DynamoDB.
The error looked like this:
Error: state data in S3 does not have the expected content
The checksum calculated for the state stored in S3 does not match the checksum
stored in DynamoDB.
Calculated checksum: XXXXXXXXXXXXXXXXXXXXXXXXXXXXXXXX
Stored checksum: YYYYYYYYYYYYYYYYYYYYYYYYYYYYYYYY
The Problem
In this setup, Terraform was using:
- An S3 bucket to store the
terraform.tfstatefile. - A DynamoDB table to manage locking and metadata associated with the remote state.
When Terraform accesses the backend, it:
- Reads the state stored in S3.
- Calculates its checksum.
- Compares that value with the
Digeststored in DynamoDB.
In our case, those two values did not match.
Because of this inconsistency, Terraform refused to continue in order to avoid working with a state that could potentially be corrupted or outdated.
Initial Checks
Before changing anything, we went through a few basic checks:
- Confirm that no
terraform applyoperation was currently running. - Confirm that no CI/CD pipeline was using the same remote state.
- Reproduce the issue both locally and from the CI/CD pipelines.
Since the same error appeared regardless of where Terraform was executed, we could quickly rule out a problem with the local machine or the pipeline configuration.
Finding the Mismatch
The next step was to inspect the DynamoDB table associated with the remote backend.
Using the AWS CLI, we searched for the entry related to the affected state:
aws dynamodb scan \
--table-name <lock-table>
Among the results, we found an entry similar to this:
{
"LockID": {
"S": "<state-path>-md5"
},
"Digest": {
"S": "old-value"
}
}
The value stored in Digest matched exactly the checksum Terraform was reporting as incorrect.
This confirmed that the current state stored in S3 and the digest registered in DynamoDB were out of sync.
The Fix
Terraform was already telling us which checksum it expected:
Calculated checksum: correct-value
Stored checksum: old-value
After verifying that the state file stored in S3 was valid, and confirming that no Terraform operation was currently running, we manually updated the Digest field in DynamoDB so that it matched the checksum calculated by Terraform.
For example:
aws dynamodb update-item \
--table-name <lock-table> \
--key '{"LockID":{"S":"<state-path>-md5"}}' \
--update-expression "SET Digest = :d" \
--expression-attribute-values '{":d":{"S":"correct-value"}}'
This type of manual change should be done carefully.
You should not update the Digest just to make the error disappear. First, make sure that the state stored in S3 is the correct one and that there are no concurrent Terraform operations running against the same backend.
Result
After updating the Digest, we ran:
terraform plan
Terraform was able to access the remote backend again, and the plan completed successfully.
We did not need to:
- Recreate the state.
- Restore a previous state version.
- Modify the Terraform code.
- Recreate any infrastructure resources.
The issue was simply a mismatch between the checksum stored in DynamoDB and the actual contents of the state file stored in S3.
Lessons Learned
When you see an error like:
state data in S3 does not have the expected content
it is easy to immediately assume that the terraform.tfstate file is corrupted.
However, that is not always the case.
The state stored in S3 may be perfectly valid, while the inconsistency exists only in the Digest value stored in DynamoDB.
Before making any changes, I recommend checking the following:
- Confirm that no Terraform executions are currently running.
- Confirm that no CI/CD pipeline is using the same backend.
- Verify the state stored in S3.
- Inspect the corresponding entry in DynamoDB.
- Compare Terraform's
Calculated checksumwith theStored checksum.
Only after completing those checks should you consider manually correcting the Digest.
In this case, a few minutes spent understanding how the remote backend works saved us from wasting much more time debugging Terraform code that was not actually the problem.
Top comments (0)