Two hundred destroys that needed 40 seconds of real work hung for 90 minutes. The platform team kicked off a terraform apply to remove stale config entries from an internal service, watched a trickle of deletes complete for twelve minutes, then the same Still destroying... lines repeat with the elapsed counters climbing and nothing completing at all, and waited until someone finally ran kill -9. By that point state and the config service disagreed, the DynamoDB lock was still held, and nobody was sure which of the 200 entries had actually been deleted. The custom Terraform provider doing the destroys had a synchronous HTTP call with no context timeout, and the backend behind it was rate-limiting at 5 RPS. Neither side was wrong on its own. The contract between them was broken.
Problem signals:
- Still destroying... lines for the same addresses repeat past any plausible duration, the elapsed counters climbing and no Destruction complete line arriving
- The backend service is healthy on its dashboard but throttling requests at a low RPS limit
- kill -9 on the terraform process leaves the DynamoDB state lock held forever
- After force-unlock, terraform state list shows resources that no longer exist in the cloud
- The custom provider in use was written internally and has no timeouts {} block support documented
What the team thought was happening, and what was actually happening
Forty seconds of work, ninety minutes of Still destroying...
The first assumption was that the internal config service was hung. It was not. Its dashboard showed it healthy and serving requests, just slowly. The second assumption was that the heartbeat lines meant terraform was making steady progress. That one was half true: for twelve minutes Terraform was making progress, just nowhere near the rate the backend could have served, and after that it was making none. The backend was rate-limiting at 5 requests per second, so 200 entries is 40 seconds of real work at that ceiling. The team waited 90 minutes.
The reason for the gap was a custom Terraform provider written by a previous platform team. Its DeleteResource function looked roughly like the snippet below. No context. No timeout. No retry-with-backoff. When the backend returned a 429, the provider's HTTP client did its own internal retry, swallowed the error, and tried again. Forever. Because the provider never returned from Delete, Terraform's supervisor saw a working call and waited. Note what that does and does not look like from the terminal: Terraform's own CLI UI hook prints module.config.config_entry.this["abc-123"]: Still destroying... [id=abc-123, 03m20s elapsed] every 10 seconds for every in-flight resource, driven by a timer in the CLI's UI hook rather than by any callback the provider makes, so the lines keep scrolling the whole time. There is no provider-side progress channel in terraform-plugin-sdk v2 whose absence could explain the hang; the defect is solely that Delete never returns. Those heartbeat lines are also the single best diagnostic available before the kill, because they name which addresses are stuck, how many are in flight, and how far the elapsed counter has climbed past any sane per-resource duration.
func resourceConfigEntryDelete(d *schema.ResourceData, meta interface{}) error {
client := meta.(*ConfigClient)
id := d.Id()
// No context. No timeout. No bound on retries.
for {
err := client.DeleteEntry(id)
if err == nil {
return nil
}
if isRateLimited(err) {
time.Sleep(1 * time.Second)
continue
}
return err
}
}
The shape of the broken Delete function (reconstructed from the provider source)
What this should have been is below. The schema.ResourceTimeout block lets users set a timeouts {} block on the resource. The context carries that deadline. When the deadline expires, the provider returns an error, Terraform reports which resource failed to destroy and leaves it in state unchanged, so the next plan re-proposes the destroy.
func resourceConfigEntryDelete(ctx context.Context, d *schema.ResourceData, meta interface{}) diag.Diagnostics {
client := meta.(*ConfigClient)
id := d.Id()
return diag.FromErr(retry.RetryContext(ctx, d.Timeout(schema.TimeoutDelete), func() *retry.RetryError {
err := client.DeleteEntryWithContext(ctx, id)
if err == nil || isNotFound(err) {
return nil // already gone counts as deleted
}
if isRateLimited(err) {
return retry.RetryableError(err)
}
return retry.NonRetryableError(err)
}))
}
What the Delete function should look like. A not-found response means the entry is already gone, which is the outcome a delete wants.
The half-finished destroy and the stuck DynamoDB lock
Why kill -9 left us worse off
Nobody pressed Ctrl-C first. Terraform's first interrupt asks the provider to stop and writes the latest state snapshot, so a kill that follows loses less; this provider's legacy Delete ignored the stop, so it would have taken a second Ctrl-C to end the run. That second interrupt cancels the operation. Terraform warns that data loss may have occurred, but its cleanup still runs and normally releases the lock; kill -9 runs nothing at all. When the engineer finally ran kill -9 on the terraform process, two things happened that compounded the problem. First, the DynamoDB lock entry stayed exactly where it was. Terraform normally releases its lock when it exits on Ctrl-C, and never on SIGKILL. So the next person who ran terraform plan got the familiar error and assumed someone else was still working on it. They were not. The lock was a ghost.
Second, the deletes had been trickling through at a small fraction of the backend's ceiling. Every worker retried on a flat one-second sleep, ignoring Retry-After, and the backend's limiter counted rejected requests against the same bucket. The retry traffic kept that bucket saturated and starved the deletes that would otherwise have gone through. Observed throughput was roughly 60 of the 200 entries in the first 12 minutes, about one delete every 12 seconds against a 5 RPS ceiling. At minute 12 the limiter's penalty for clients that keep hitting it escalated, and after that nothing got through at all. That gap between the nominal limit and the achieved rate, and then the stall, is why 40 seconds of work turned into ninety minutes. Terraform writes state to the remote backend as an apply progresses, whenever a resource finishes and at least 20 seconds (by default) have passed since the last write, so nearly every completed delete was recorded. The exceptions sat at the edges. The last two deletes finished within 20 seconds of the write before them, and the stall meant nothing finished afterwards to trigger another. And one of the deletes still in flight when the kill landed had already been processed by the backend: its success response was lost when the connection dropped, the client's internal retry re-sent the DELETE, every retry since had drawn a 429, and the provider never saw that the entry was gone. State and the backend now disagreed at the edges, and the provider's Read function carried a separate defect of its own - on a not-found response it re-wrote the prior state instead of calling d.SetId("") - so a refresh would reconcile drift on the entries that still existed but would silently leave every already-deleted entry sitting in state.
Before doing anything else we confirmed the terraform process was actually dead on the operator's machine. ps aux | grep terraform, on the actual machine, not a tmux pane from yesterday. We have force-unlocked locks that turned out to belong to a process still doing useful work, and the damage is worse than a stuck lock. Once confirmed dead, terraform force-unlock with the lock ID from the error message released DynamoDB. (The S3 backend now deprecates the DynamoDB lock table in favour of use_lockfile; force-unlock releases either.)
# 1. Confirm no terraform process is running on the operator's machine
ssh operator-host 'ps aux | grep -v grep | grep terraform'
# 2. Release the lock (lock ID comes from the error message)
terraform force-unlock 7c4a3e22-1b9d-4e8a-b6d7-9f2a8c5e4d11
# 3. See what state thinks vs what the cloud actually has
terraform plan -refresh-only
# 4. Apply the refresh so surviving entries pick up attribute drift. With this
# provider's Read defect it will NOT drop already-deleted entries from state,
# so it is not a substitute for the tfstate-vs-backend diff below.
terraform apply -refresh-only
The recovery sequence after confirming the process is dead
Scripting state rm and import for 200 entries
Reconciling state against a half-finished destroy
After the refresh-only apply, state and cloud still did not agree, and it is worth being precise about why. terraform apply -refresh-only reconciles state only through the provider's Read, and this provider's Read never signalled deletion: on a not-found response it returned the prior state unchanged instead of calling d.SetId(""). So the refresh picked up attribute drift on entries that still existed and was a no-op for every entry that had been deleted. Two populations came out of it cleanly. Entries that still existed in both tfstate and cloud, because the destroy never reached them, could be destroyed normally once the patched provider was in place. Entries the hung apply had deleted and Terraform had recorded on a successful return were already gone from both sides and needed nothing further, and that group is the work Terraform itself persisted, not anything the refresh did. Everything else - every entry deleted during the hung apply that Terraform never got to record - was still sitting in state, and no amount of refreshing was going to move it. Those had to be found by diffing tfstate against the backend and removed with terraform state rm. If you can ship a provider patch first, the cleaner recovery order is to fix Read and re-run the refresh before diffing; a Delete that treats not-found as success would also let a plain destroy clear them.
That third population is where it got annoying, because nothing in the Terraform output marks those entries as unrecorded. The entries had been deleted from cloud by the hung apply, but the provider's Read function never signalled the deletion. On a not-found response it returned the prior state unchanged instead of calling d.SetId(""), and the client's own response cache could still serve an entry that had already been removed, so Terraform kept the resource in state either way. Those entries were ghosts in tfstate. There were three of them: the two deletes that finished after the last state write, and the in-flight delete the backend had already processed. Three is a small number, but only two of them would have printed a Destruction complete line, to a terminal nobody had logged, so we found all three by diffing state against the backend and scripted the terraform state rm.
# Pull current tfstate resource list. A 200-entry set is managed with for_each,
# so the addresses look like module.config.config_entry.this["abc-123"] - the
# backend ID is the instance key, never the resource name.
terraform state list | grep config_entry > tfstate_entries.txt
# Derive the backend IDs from the state addresses instead of manufacturing
# addresses from the IDs.
sed -n 's|.*\["\(.*\)"\]$|\1|p' tfstate_entries.txt | sort > tfstate_ids.txt
# Pull live entries from the backend, one page at a time, with a pause between calls
: > live_entries.txt
page=1
while :; do
body=$(curl -sf "$CONFIG_API/entries?page=$page&per_page=100") || { echo "fetch failed on page $page"; exit 1; }
count=$(printf '%s' "$body" | jq '.entries | length')
[ "$count" -eq 0 ] && break
printf '%s' "$body" | jq -r '.entries[].id' >> live_entries.txt
page=$((page + 1))
sleep 1
done
# Guard: an empty live list means the fetch broke, not that the backend is empty.
# Without this, every resource in tfstate looks like a ghost.
if [ ! -s live_entries.txt ]; then
echo "live_entries.txt is empty, refusing to generate state rm commands"
exit 1
fi
# IDs in tfstate but not in cloud: these are ghosts. Raw ID against raw ID.
comm -23 tfstate_ids.txt <(sort live_entries.txt) > ghost_ids.txt
# Sanity guard: if every state entry is classified as a ghost, the comparison is
# broken (format mismatch), not the backend. Abort rather than empty the state.
total=$(wc -l < tfstate_ids.txt)
ghosts=$(wc -l < ghost_ids.txt)
if [ "$ghosts" -eq "$total" ]; then
echo "all $total state entries classified as ghosts, refusing to run state rm"
exit 1
fi
# Map the surviving IDs back to their Terraform addresses and remove them
while read id; do
terraform state rm "module.config.config_entry.this[\"$id\"]"
done < ghost_ids.txt
Generating the state rm commands from a diff between tfstate and the live backend
For the inverse case (entry exists in cloud but not in tfstate), the recovery is terraform import. We did not hit this on this incident but we have hit it on similar ones, and the same diff approach works in the other direction. The general pattern for any half-finished Terraform operation against a custom provider is laid out in our Terraform state recovery playbook.
The contract every custom Terraform provider has to honor
What the provider should have done
A custom Terraform provider is a contract. Terraform's whole supervision model assumes the provider plays by it. The contract is short: Create, Read, Update, and Delete each accept a context, each respect the user's timeouts {} block, each emit clear errors when something goes wrong, and each return in bounded time. When a provider violates the contract, Terraform's user-facing behavior degrades in ways that look like Terraform bugs but are not.
Internal providers skip the contract more often than vendor ones, because the team that writes the provider also runs the backend it talks to, and they convince themselves they have full visibility. They do not. terraform-cli is a separate process. It cannot see your retry loop. It cannot see your in-flight HTTP call. All it sees is a function that has not yet returned. The fix for this provider was five changes:
| Step | What it does |
|---|---|
| 1. Accept context on every CRUD function | The legacy Create/Read/Update/Delete fields on schema.Resource, with the func(*schema.ResourceData, interface{}) error signature, are deprecated in terraform-plugin-sdk v2, not removed: they still compile and run, which is exactly how the broken Delete above survives in a live provider. Migrate to CreateContext/ReadContext/UpdateContext/DeleteContext. The legacy variants get no context and therefore can never honor a timeout, which is the actual reason to migrate. |
| 2. Declare and honor timeouts on every resource | Add a Timeouts: &schema.ResourceTimeout{Create: schema.DefaultTimeout(5 * time.Minute), Delete: schema.DefaultTimeout(5 * time.Minute)} block on every resource schema. Use d.Timeout(schema.TimeoutDelete) inside the function. |
| 3. Replace internal retry loops with retry.RetryContext | The retry helper respects the context deadline, backs off exponentially from half a second up to ten seconds, and separates retryable from non-retryable errors. Hand-rolled for-loops over time.Sleep do none of that. The helper does not read Retry-After, though: a backend that sends one is better served by a rate limiter inside the provider's client. |
| 4. Pin the fixed version via .terraform.lock.hcl | Release a new patch version of the provider, update the lockfile, and remove the old version from your internal registry so nobody can fall back to it. |
| 5. Make Read signal deletion on a not-found response | This is a separate defect from the missing timeout, and cards 1 and 2 do not fix it: a context deadline governs whether a call returns, never what it returns. A Read that treats a 404 or a 429 as "no change" and re-writes the prior state leaves ghosts that no refresh will clear. On a not-found response, call d.SetId("") and return nil so Terraform drops the resource from state, surface throttling errors instead of swallowing them, and do not serve reads from a client-side response cache. |
The apply pattern itself also needed a change, and the obvious one would not have helped. Terraform runs at most 10 operations at once by default, and each resource's timeout starts only when Terraform calls its Delete, so the queue behind the first ten never eats into anyone's timeout. What matters against a 5 RPS backend is how many calls are in flight. Batching with -target would not change that: a batch of 10 targets runs the same 10 concurrent deletes the default does, and HashiCorp reserves -target for exceptional circumstances. Bulk destroys now run with a lower -parallelism, the provider's client carries its own rate limiter so a burst queues inside the provider instead of tripping the backend's, and the backend team is adding a bulk delete endpoint, which the provider will wrap as a single resource operation instead of looping.
The relationship that broke and what fixes each side
When a custom provider has left your state in an unknown shape
If you are looking at a hung apply right now
Hung Terraform applies against internal providers are the kind of incident that sounds boring in a postmortem and feels terrifying in the moment. You cannot tell if the apply is still doing useful work or stuck forever. You cannot kill it without risking a half-finished state, though one Ctrl-C first makes Terraform write what it knows, and a second one ends the run and normally releases the lock. You cannot force-unlock until you are certain the process is dead. And once you do recover, you do not actually know which resources got modified and which did not, because the Still destroying... heartbeat names which resources are in flight but says nothing about what the provider did to them, and the kill threw away updates Terraform was still holding in memory.
We run these recovery engagements often enough that the script above is templated. The no-timeout custom provider pattern shows up in maybe one in five of the Terraform recoveries we have done this year, almost always with internal providers written years ago by an engineer who has since left. The fix is mechanical once you know the shape of the failure: confirm process death, force-unlock, refresh-only plan, diff state against cloud, reconcile with state rm and import, then patch the provider so it cannot happen again.
If you are staring at a hung apply right now and you are not sure whether to kill it, book an infrastructure review with our team and we will be on a bridge with you the same day. If the apply is already dead and you are sorting through the wreckage, the same engagement covers the state reconciliation and the provider fix together.
Originally published at https://infraforge.agency/insights/terraform-apply-hung-custom-provider-no-timeout/.
If your team is dealing with similar infrastructure debt, we offer infrastructure reviews and recovery engagements — see /review.
Top comments (0)