Disclosure: I work at Eon, on the content side. Our team ran these assessments and this is our data. I didn't write the queries, but I sat with the engineer who did, and the methodology below includes what we couldn't see. Argue with it.
Most backup tools produce a coverage report, and most of those reports work from tags. They check the resources a policy already enrolled, which is usually everything tagged environment=production, and report how many of those have a recent recovery point. The number comes back green.
That number is real. It's just answering a smaller question than the one you asked.
We wanted to know how much smaller. Across 78 companies' cloud environments on AWS, Google Cloud, and Azure, our team connected through a read-only role, discovered every resource, classified which ones were production, and checked each one for a recovery copy. 101,340 live resources, 53,709 of them production. Pulled September 15, 2026.
Only 14% of AWS production resources had a production tag
This is the finding that explains all the others.
Of about 47,000 AWS production resources:
- 14% carried a tag saying production
- 66% had no environment tag at all
- 27% had no tags of any kind
- 19% had an environment tag containing something that isn't an environment
That last group is the one you'll recognize. Real values found in the environment field: base, stable, research, corporate.
Somebody typed those. The field was required, it was 6pm, and they put something in it. Then a backup policy read that field to decide what to protect.
A coverage report keyed to environment=production would have found about 1 in 7 of these resources. The rest never enter the count, so they can't show up as unprotected. They don't show up at all.
61.2% had no recovery copy
Of the 53,709 production resources, 61.2% had no backup policy, no recovery point, and no native cloud snapshot the assessment could detect.
Three definitions, in case you think we picked the flattering one:
| Definition | Share |
|---|---|
| No policy, no recovery point, no native snapshot detected | 61.2% |
| No backup policy at all | 62.8% |
| No policy, or a policy that's attached but blocked or over a limit | 65.4% |
There's no reasonable definition that gets this below 61%.
The unprotected resources are the old ones
This is the number that changed how I think about the problem. I'd assumed the unprotected resources would be new ones. Spun up last sprint, policy hasn't caught up, on someone's list.
- Median age of a production resource with no recovery copy: 502 days
- Median age of one with a recovery copy: 363 days
- Share of unprotected resources older than 12 months: 55%
So they're about four months older on average. Not a backlog. A resident population that nobody is watching.
It's also the best argument against the obvious objection, that these resources are protected by something the assessment couldn't see. A resource that's been running 16 months with no policy, no tag, and no documentation isn't usually one somebody carefully set up a separate backup plan for.
0.9% of the gaps were on purpose
Skipping backups is sometimes correct. A box that Terraform rebuilds on every deploy, a cache that refills from upstream, a low-value system where the cost genuinely isn't worth it. Fine.
So we counted how often that decision was actually recorded. 0.9% of production resources had a documented exclusion. The other 61.2% had no policy and no note.
The distinction only matters on one day, but it matters a lot that day. "We assessed it and accepted the risk" is an answer. "We didn't know it was there" is a different conversation.
Nobody was clean
Among the 51-odd environments with 25 or more production resources:
- Typical environment: 64.5% of production unprotected
- More than half were above 50%
- Zero reached 100%
- Best in the set: 99.8%
The median tracks the aggregate almost exactly, so this isn't a few disasters dragging an average.
Why this is worse than it was two years ago
For a long time, leaving a low-tier resource unprotected was a defensible cost decision. What broke production was a human mistake or a hardware fault, both rare, both usually fixable by redeploying from code and reloading from upstream.
An AI agent with valid credentials has a different blast radius. It writes at machine speed to anything those credentials reach, and it has no idea how you tiered your estate. The cache, the queue, the read replica, and the database behind them go together.
Terraform brings back the table. Not the rows.
What to actually do about it
Count from the cloud side, not from the backup tool. Pull live inventory from every account and region via read-only APIs, then match each production resource against a recovery point. Starting from your backup tool's inventory guarantees you only count what it already knows.
Don't classify production from one tag. Look at the name, the account, what depends on it, and who gets paged when it breaks. Expect the untagged resources to be the oldest ones.
Turn every gap into a written decision. Attach a policy or write down why not, with an owner and a review date. A documented exclusion is defensible. An undecided gap isn't.
Keep recovery copies outside the blast radius. Credentials broad enough to modify production data are often broad enough to delete snapshots in the same account. Bucket versioning protects against overwrites, not against a credential that can delete versions.
Re-run the count on a schedule. New resources arrive weekly. Unless a policy attaches at creation, the coverage you proved Tuesday is stale by Friday, and in our data those gaps persist for about 16 months.
Methodology, including what we couldn't see
The numbers come from cloud data protection assessments run across 78 companies' cloud environments. Each connects through a single read-only IAM role and discovers and classifies every resource. No agents, no appliances, no manual tagging. Results aggregated and anonymized; nothing identifies a company, account, or resource.
Cohort: 78 environments, 101,340 live resources across AWS, GCP, and Azure, 53,709 production. Excluded: sandboxes, inactive environments, deleted resources, disconnected accounts, and child resources that inherit a parent's protection (an EBS volume attached to an EC2 instance counts once, through the instance). Distribution stats use environments with 25+ production resources.
How "production" was determined: from resource metadata — name, IDs, tags, the account it runs in, and how the rest of that account was already classified. Tags are one input among several, so a resource with no tags still gets classified.
The honest limitation: native snapshot detection is uneven. The assessment sees native snapshots on 98% of RDS instances, 42% of FSx file systems, 24% of EC2 instances, and almost none of the other resource types. So for S3, DynamoDB, EFS, EKS, and anything on GCP or Azure, "no recovery copy" means no policy and no recovery point in our data. A native backup, or a backup made by another tool, could exist outside that view. The age and tagging findings are why I don't think that accounts for much of the 61%, but weigh it yourself.
Also worth saying: a classifier is a model of an environment, not the environment, and it has an error rate. And counting unprotected resources says nothing about whether the protected ones would actually restore. That's a separate study, and honestly a harder one.
Full write-up with the rest of the findings is here. I'll answer what I can in the comments and pull in the engineer who ran the queries for anything I can't. And if your environment looks nothing like this, I'd like to hear what you're doing differently.
Top comments (0)