<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Rick Wise</title>
    <description>The latest articles on DEV Community by Rick Wise (@cloudwiseteam).</description>
    <link>https://dev.to/cloudwiseteam</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3582447%2Fe7a88946-c7a3-4aad-9242-6d52380c09f1.png</url>
      <title>DEV Community: Rick Wise</title>
      <link>https://dev.to/cloudwiseteam</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/cloudwiseteam"/>
    <language>en</language>
    <item>
      <title>SR&amp;ED documentation as a solo founder: log technological uncertainty while you build</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Thu, 24 Sep 2026 13:27:37 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/sred-documentation-as-a-solo-founder-log-technological-uncertainty-while-you-build-b4</link>
      <guid>https://dev.to/cloudwiseteam/sred-documentation-as-a-solo-founder-log-technological-uncertainty-while-you-build-b4</guid>
      <description>&lt;p&gt;The SR&amp;amp;ED claim took me longer than the feature I was claiming for.&lt;/p&gt;

&lt;p&gt;Not the form. The reconstruction. CRA wants the technological uncertainty documented: what you didn't know, what you tried, why the obvious approach failed. I had none of that written down. I had commits.&lt;/p&gt;

&lt;p&gt;Rebuilding that narrative from about 14 months of git history and AI-assisted coding sessions took roughly three weeks I had not budgeted. This post is about the habit that would have made it a few days: a running log, two lines per dead end, written on the day the dead end happens.&lt;/p&gt;

&lt;p&gt;Context, so you can weigh this properly. I'm the solo founder of CloudWise, an AWS cost optimization tool: 189 waste checks across 40+ AWS services, built by one person, pre-revenue, no paying customers. I paused feature development in August 2026 to work on distribution. I filed the SR&amp;amp;ED claim myself, without a consultant, and later received a CRA T2 notice of reassessment that resulted in a refund. I have filed exactly once. Everything below is one founder's experience, not advice on whether your work is eligible.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem: reconstructing SR&amp;amp;ED documentation after the fact
&lt;/h2&gt;

&lt;p&gt;If you are mid-build, your record of the work is probably the same as mine was: a git log, some pull request descriptions, a ticket tracker, and a pile of chat sessions with an AI coding assistant.&lt;/p&gt;

&lt;p&gt;That record answers one question well: what changed. It answers a different question badly: what did you not know at the time, and how did you find out.&lt;/p&gt;

&lt;p&gt;A commit message like &lt;code&gt;fix: handle untagged resources in attribution&lt;/code&gt; tells you a fix landed. It does not tell you whether that fix was the first approach or the fourth, why the earlier ones failed, or what each failure taught you. The failed approaches are often not in the history at all. They lived on a branch you deleted, or in a session you never saved, or only in your head.&lt;/p&gt;

&lt;p&gt;So the three weeks were not writing time. They were archaeology: reading my own diffs in order, working out which ones marked a change of approach, then trying to recover why I had abandoned the previous one. Fourteen months later, the why is the part that has decayed the most.&lt;/p&gt;

&lt;h2&gt;
  
  
  What CRA actually asks you to document on the T661
&lt;/h2&gt;

&lt;p&gt;I am going to stay close to the source here, because this is the part where founder blog posts tend to freelance.&lt;/p&gt;

&lt;p&gt;The claim form is the T661. Its project section asks, in order, for three things: the scientific or technological uncertainties you faced, the work you did to try to overcome them, and the advancement you achieved or attempted. CRA's &lt;a href="https://www.canada.ca/en/revenue-agency/services/forms-publications/publications/t4088/guide-form-t661-scientific-research-experimental-development-expenditures-claim-guide-form-t661.html" rel="noopener noreferrer"&gt;guide to Form T661 (T4088)&lt;/a&gt; walks through those lines and their word limits. The same guide has you indicate what supporting evidence exists for the work, and notes that you keep that evidence rather than submit it with the claim.&lt;/p&gt;

&lt;p&gt;CRA's &lt;a href="https://www.canada.ca/en/revenue-agency/services/scientific-research-experimental-development-tax-incentive-program/sred-policies-guidelines/guidelines-eligibility-work-sred-tax-incentives.html" rel="noopener noreferrer"&gt;guidelines on the eligibility of work&lt;/a&gt; describe technological uncertainty as a situation where it is unknown whether, or how, a result can be achieved because the available technological knowledge is insufficient. The same page tells claimants to keep the evidence that is generated as the work progresses.&lt;/p&gt;

&lt;p&gt;Read those two pages yourself before you read anyone's summary of them, including mine. They are shorter than you expect.&lt;/p&gt;

&lt;p&gt;The practical consequence for technological uncertainty in an SR&amp;amp;ED claim is this: the form asks for a narrative with a specific shape. Uncertainty, then systematic work, then what was learned. A git log has a different shape. It is a list of things that worked well enough to commit.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I had vs. what I needed
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;What the form asks about&lt;/th&gt;
&lt;th&gt;What I had&lt;/th&gt;
&lt;th&gt;What I needed&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;What was uncertain at the start&lt;/td&gt;
&lt;td&gt;Nothing written. Memory.&lt;/td&gt;
&lt;td&gt;A dated note stating what I did not know and why existing approaches did not settle it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What I tried, in order&lt;/td&gt;
&lt;td&gt;Commits for the approaches that survived&lt;/td&gt;
&lt;td&gt;A record of the approaches that did not survive, with the reason each was dropped&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Why the obvious approach failed&lt;/td&gt;
&lt;td&gt;Occasionally a PR description&lt;/td&gt;
&lt;td&gt;One sentence, written the day it failed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What was learned&lt;/td&gt;
&lt;td&gt;The final code&lt;/td&gt;
&lt;td&gt;A plain statement of the conclusion, including "this does not work"&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;When the work happened&lt;/td&gt;
&lt;td&gt;Commit timestamps&lt;/td&gt;
&lt;td&gt;The same, which was the one thing git gave me for free&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Timestamps were the only column where my existing record was enough. Everything else I rebuilt.&lt;/p&gt;

&lt;p&gt;The AI-assisted coding sessions were the other source. Where a session survived, it sometimes held the reasoning I was missing, because I had explained the problem to the assistant in plain language before asking for help. But I had never kept them with this purpose in mind, so coverage was patchy and finding the relevant ones was slow.&lt;/p&gt;

&lt;p&gt;One more thing I learned during the reconstruction: most of what I had built did not belong in the claim at all. Standard integration work, UI, and infrastructure plumbing did not qualify. What qualified was a narrow slice, cross-account cost attribution where tagging is incomplete and the answer has to be inferred, because that was where I did not know whether the approach would work. Sorting fourteen months of work into those two piles was part of the reconstruction. A log would have done that sorting for me as I went.&lt;/p&gt;

&lt;h2&gt;
  
  
  The lightweight SR&amp;amp;ED documentation habit: two lines per dead end
&lt;/h2&gt;

&lt;p&gt;Here is the whole system. One file in the repo, &lt;code&gt;docs/uncertainty-log.md&lt;/code&gt;, append-only. An entry only when an approach fails or you hit something you do not know how to do. Not a journal. Not a standup.&lt;/p&gt;

&lt;p&gt;Each entry is two lines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## 2026-03-14&lt;/span&gt;
Tried: inferring the owning account for untagged resources from the creating IAM principal.
Failed because: a shared deploy role creates resources for several accounts, so the principal does not identify the owner. Next: try usage-pattern correlation. (branch: attribution-v2, commit abc1234)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That entry is an illustration of the format in my problem area, not a record from my actual claim.&lt;/p&gt;

&lt;p&gt;The rules I would follow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Write it the day it fails.&lt;/strong&gt; The reason is obvious today and gone in a month.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Line one is what you tried. Line two is why it failed and what that told you.&lt;/strong&gt; If you cannot write line two, you have not understood the failure yet, which is worth knowing anyway.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Link the commit or branch.&lt;/strong&gt; The log carries the reasoning, git carries the dates and the diffs. Together they cover each other's gaps.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Log the question before the answer.&lt;/strong&gt; When you start on something you do not know how to do, write one line saying so. That line is the start of the narrative the form asks for.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Do not log routine work.&lt;/strong&gt; If the path was known and it was just effort, it does not go in. This keeps the file short and gives you, or whoever prepares your claim, a first rough sort of uncertain work from routine work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Save the AI sessions that contain reasoning.&lt;/strong&gt; Export them into a folder next to the log, named by date. They are timestamped explanations of what you did not know, written at the time. I had these by accident. Have them on purpose.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cost: a couple of minutes per dead end. I would have traded that for three weeks without thinking about it.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Start the log on day one of the build&lt;/strong&gt;, not when I first hear the word SR&amp;amp;ED.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Stop deleting failed branches.&lt;/strong&gt; Tag them and leave them. A failed approach with a commit history is evidence. A deleted one is an anecdote.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write PR descriptions that say what was ruled out&lt;/strong&gt;, not only what was done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Read the T661 guide before the build, not after.&lt;/strong&gt; Knowing the shape of the questions changes what you bother to write down.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Budget the time.&lt;/strong&gt; Even with a good log, the claim is real work. Without one, it was three weeks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Separate uncertain work from routine work as I go&lt;/strong&gt;, because that line decided what went into the claim, and drawing it afterwards was slow.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  What this post is not
&lt;/h2&gt;

&lt;p&gt;I am a founder who filed once. I am not an SR&amp;amp;ED advisor, I cannot assess your work, and I cannot tell you whether you qualify. If you are a Canadian founder working on SR&amp;amp;ED for startups and you think some of your work involved real technological uncertainty, talk to a qualified claim preparer and read CRA's own pages linked above.&lt;/p&gt;

&lt;p&gt;What I can tell you is narrower and I am confident in it: whatever you end up claiming, and whoever prepares it, the record you keep while building is the raw material. Two lines per dead end. Start today.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Rick Wise is the solo founder of &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;CloudWise&lt;/a&gt;, an AWS cost optimization tool with 189 waste checks across 40+ AWS services. He filed one SR&amp;amp;ED claim, himself, and is not a tax advisor. He helps technical founders reconstruct the engineering narrative from their repo and AI coding sessions at &lt;a href="https://sredlab.ca/?ref=cw-blog" rel="noopener noreferrer"&gt;SR&amp;amp;ED Lab&lt;/a&gt;. Your accountant or claim preparer still files.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>productivity</category>
      <category>git</category>
      <category>saas</category>
      <category>career</category>
    </item>
    <item>
      <title>How We Prove an AWS Waste Check Works: Fire on a Real Cluster, Stay Silent on Its Twin</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 23 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/how-we-prove-an-aws-waste-check-works-fire-on-a-real-cluster-stay-silent-on-its-twin-5275</link>
      <guid>https://dev.to/cloudwiseteam/how-we-prove-an-aws-waste-check-works-fire-on-a-real-cluster-stay-silent-on-its-twin-5275</guid>
      <description>&lt;p&gt;A unit test tells you your code does what you think AWS does. It can't tell you what AWS actually does.&lt;/p&gt;

&lt;p&gt;For a tool that reads CloudWatch metrics and billing data to decide whether a resource is wasted, that gap is the whole product. So we test each check the same way: build the wasted thing for real, and make the check prove itself on it.&lt;/p&gt;

&lt;p&gt;The rule is two halves:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Fixture:&lt;/strong&gt; a real resource that &lt;em&gt;should&lt;/em&gt; trigger the check. The check has to fire on it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Control:&lt;/strong&gt; the healthy version of the same thing. The check has to stay silent on it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fire without silence and you've built an alarm that always rings. Silence without fire and you've built nothing. You need both, measured against live AWS, before you trust the check.&lt;/p&gt;

&lt;p&gt;Here is what that looked like for one check: idle DocumentDB.&lt;/p&gt;

&lt;h2&gt;
  
  
  The check, and the real cluster we pointed it at
&lt;/h2&gt;

&lt;p&gt;The rule: flag a DocumentDB cluster with &lt;strong&gt;zero connections&lt;/strong&gt; and combined read plus write IOPS &lt;strong&gt;under 10 per second&lt;/strong&gt;, averaged over 7 days.&lt;/p&gt;

&lt;p&gt;We stood up a real &lt;code&gt;db.t3.medium&lt;/code&gt; cluster in a separate test AWS account and left it alone. It costs $1.87 a day and it's a short-lived test resource, not something we keep. Measured on 2026-09-19:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric (7-day window)&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;th&gt;Gate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DatabaseConnections&lt;/td&gt;
&lt;td&gt;0.0000&lt;/td&gt;
&lt;td&gt;must be 0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combined read + write IOPS&lt;/td&gt;
&lt;td&gt;5.46 / s&lt;/td&gt;
&lt;td&gt;under 10 / s&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is an idle cluster by any reasonable reading. The check should fire.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the real cluster told us that the unit tests couldn't
&lt;/h2&gt;

&lt;p&gt;It did not fire. Not on the first try, and not for one reason. Three separate bugs were stacked in front of it, and every one of them lived in the gap between our code and AWS's behavior.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Layer&lt;/th&gt;
&lt;th&gt;What was wrong&lt;/th&gt;
&lt;th&gt;How it hid&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Wiring&lt;/td&gt;
&lt;td&gt;The detector called a metrics method that didn't exist on the live data provider, so it always got nothing back. Fixed 2026-07-09.&lt;/td&gt;
&lt;td&gt;A missing answer reads as "no data", which reads as "nothing wrong here".&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Units&lt;/td&gt;
&lt;td&gt;CloudWatch reports DocumentDB IOPS as a &lt;strong&gt;rate&lt;/strong&gt; (Count per Second). We read it with &lt;code&gt;Statistics=Sum&lt;/code&gt; over 7 days, which adds a rate up as if it were a running total. A steady ~5.15 per second over 10,080 one-minute datapoints sums to roughly &lt;strong&gt;52,000&lt;/strong&gt;, against a gate expecting a small number.&lt;/td&gt;
&lt;td&gt;Our regression tests pinned the same &lt;code&gt;Sum&lt;/code&gt;-shaped mock, so they enshrined the bug.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Naming&lt;/td&gt;
&lt;td&gt;Cost data calls the product "Amazon DocumentDB (with MongoDB compatibility)". Our service-to-detector map only knew "Amazon DocumentDB" and "AmazonDocDB", and the lookup is exact, so on every scheduled scan the DocumentDB detector was skipped, even on accounts whose costs listed DocumentDB.&lt;/td&gt;
&lt;td&gt;A skipped detector reports nothing, which looks exactly like an account with no waste.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The second row is the one worth sitting with. The tests were not sloppy. They were written against a mock that matched the code's assumption, so they could only ever confirm the assumption. Only a real cluster reporting a real rate could disagree, and it did.&lt;/p&gt;

&lt;p&gt;The fix for the second bug reads the &lt;strong&gt;average&lt;/strong&gt; of the rate instead of the sum, and gates on combined IOPS averaging under 10 per second with zero connections.&lt;/p&gt;

&lt;h2&gt;
  
  
  The first real firing, and why we almost missed it
&lt;/h2&gt;

&lt;p&gt;With all three fixed, the 2026-09-19 06:01 UTC scan ran the DocumentDB detector on that cluster for the first time and it fired: idle cluster, list price $56.94 a month for one &lt;code&gt;db.t3.medium&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;You would not have seen it in the findings list, though. Our scan reconciles every finding against the account's actual billing before showing it, and the test account was on DocumentDB's free trial, so the finding reconciled to $0.00 and was dropped from the customer-facing list on purpose. It only existed in the pre-floor debug output we keep for exactly this kind of question.&lt;/p&gt;

&lt;p&gt;That's the system working as designed, and it's also why "look at the findings page" is not a test. We had to read the layer beneath it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The control: the same cluster, connected
&lt;/h2&gt;

&lt;p&gt;A healthy twin usually means a second resource. For this check the healthy state is simply "someone is using it", and the two states are mutually exclusive by construction: the check requires zero connections, so any connection at all rules idle out.&lt;/p&gt;

&lt;p&gt;So the control was the same cluster with one variable changed. On 2026-09-19 we opened a handful of connections from a short-lived Lambda inside the VPC and deleted it the same day.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Fixture (idle)&lt;/th&gt;
&lt;th&gt;Control (connected)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DatabaseConnections&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;0 to 6 per minute while connected, 7-day average 0.0169&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Combined IOPS&lt;/td&gt;
&lt;td&gt;5.46 / s&lt;/td&gt;
&lt;td&gt;5.44 / s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scan on 2026-09-20&lt;/td&gt;
&lt;td&gt;fired the day before&lt;/td&gt;
&lt;td&gt;detector ran, &lt;strong&gt;0 findings&lt;/strong&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Nothing moved except the connections. IOPS stayed under the gate, and one afternoon of traffic was enough to pull the 7-day average off zero and silence the check. Fired on the fixture, silent on the control.&lt;/p&gt;

&lt;p&gt;One honest caveat, because it would be easy to leave out: those two observations are separated in time, not simultaneous. It's one cluster measured in two states, and the code guarantees the states can't both be true. That's a weaker claim than "two clusters side by side", and we say so in the write-up.&lt;/p&gt;

&lt;h2&gt;
  
  
  Controls need checking too
&lt;/h2&gt;

&lt;p&gt;The same sweep caught the opposite failure. A Glue job we built as a healthy control set off the "missing timeout" check. We had left the timeout out when creating it, and Glue quietly defaults to 2,880 minutes, which is exactly the number the check looks for. The control was never healthy.&lt;/p&gt;

&lt;p&gt;A control that fires isn't a control. We now set the timeout explicitly, and we verify every control is silent across &lt;em&gt;all&lt;/em&gt; checks, not only its own.&lt;/p&gt;

&lt;h2&gt;
  
  
  What we took from it
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;A passing unit test is a statement about your assumptions.&lt;/strong&gt; Real resources are the only thing that can contradict them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test the whole path, not the function.&lt;/strong&gt; All three bugs lived outside the detector's decision logic: in how metrics were fetched and read, and in how the check got selected to run.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prove silence as hard as you prove firing.&lt;/strong&gt; A check nobody has watched stay quiet is a guess.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;It's cheap.&lt;/strong&gt; One small cluster for roughly a week costs about what a lunch does, and it found bugs our mock-based tests couldn't.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's the method we use when we validate a check against real AWS. CloudWise runs 189 waste checks across 40+ AWS services, and if you want to see what a read-only pass finds in your own account, it takes about five minutes.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Free, read-only AWS waste scan: &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;cloudcostwise.io&lt;/a&gt; — five minutes, no card.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>testing</category>
      <category>devops</category>
      <category>finops</category>
    </item>
    <item>
      <title>Your Container Tag Is Lying to You: A Mutable ECR Tag Put the Wrong Build in Production</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 09 Sep 2026 13:00:00 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/your-container-tag-is-lying-to-you-a-mutable-ecr-tag-put-the-wrong-build-in-production-1o57</link>
      <guid>https://dev.to/cloudwiseteam/your-container-tag-is-lying-to-you-a-mutable-ecr-tag-put-the-wrong-build-in-production-1o57</guid>
      <description>&lt;p&gt;On 2026-08-21, a routine docs-only PR merged to &lt;code&gt;main&lt;/code&gt;. CI built the frontend image, pushed it, and called &lt;code&gt;start-deployment&lt;/code&gt; on the App Runner service — and got rejected. The pipeline went red. &lt;strong&gt;"Deploy to Production" failed.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;That part is normal; pipelines go red. What isn't normal is what CloudTrail and ECR showed once we went looking: production wasn't broken by the red step. It was broken by the deploy that &lt;em&gt;succeeded&lt;/em&gt; 3 seconds earlier — from a laptop, not CI — and the pipeline's red error message never once mentioned it.&lt;/p&gt;

&lt;h2&gt;
  
  
  The timeline, from CloudTrail and ECR, not from memory
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Time (EDT)&lt;/th&gt;
&lt;th&gt;Event&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;10:52:25&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;StartDeployment&lt;/code&gt; on &lt;code&gt;cloudwise-production-frontend&lt;/code&gt;, called by an IAM user from a local machine (&lt;code&gt;aws-cli/2.34.2&lt;/code&gt;, macOS, arm64) — not CI. At this instant the &lt;code&gt;production&lt;/code&gt; tag points at the &lt;strong&gt;previous&lt;/strong&gt; build.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10:52:43&lt;/td&gt;
&lt;td&gt;CI pushes the new frontend image and moves the &lt;code&gt;production&lt;/code&gt; tag to point at it.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10:52:44&lt;/td&gt;
&lt;td&gt;CI calls &lt;code&gt;StartDeployment&lt;/code&gt; as its own deploy role → &lt;code&gt;InvalidRequestException: Can't start a deployment … because it isn't in RUNNING state&lt;/code&gt; → exit 254 → &lt;strong&gt;pipeline red&lt;/strong&gt;.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10:55:49&lt;/td&gt;
&lt;td&gt;The manual deployment from 10:52:25 &lt;strong&gt;succeeds&lt;/strong&gt; — having resolved the &lt;code&gt;production&lt;/code&gt; tag back at 10:52:25, before CI moved it. It deploys the &lt;strong&gt;old&lt;/strong&gt; image.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Net state: the &lt;code&gt;production&lt;/code&gt; tag pointed at the image CI had just built. That image was never deployed. App Runner was &lt;code&gt;RUNNING&lt;/code&gt; on the previous build, and the pipeline was red for a reason that named none of this.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;18 seconds&lt;/strong&gt; is the whole story. That's the window between the manual deploy resolving the tag and CI moving it. Move a mutable tag inside that window and the registry and the running container permanently disagree about which image "production" means.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why a mutable tag is the actual defect
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;ECR_TAG=production&lt;/code&gt; for the frontend repository always points at the newest push. App Runner doesn't pin a digest — it resolves whatever &lt;code&gt;production&lt;/code&gt; points to &lt;strong&gt;at the moment a deployment starts&lt;/strong&gt;, not at the moment the container swap completes. Those are two different instants, and normally nothing happens in between.&lt;/p&gt;

&lt;p&gt;Two deployments racing to start closes that gap. One of them read the tag old; the other wrote it new. Whichever deployment actually executes the swap runs whatever it resolved, and there's no signal anywhere in App Runner's own state that says "the tag has moved since I started."&lt;/p&gt;

&lt;p&gt;This had already been flagged as a hazard in the abstract — two CI runs landing close together can serve an older image, all green, no laptop involved. This incident reached the same failure by a different route: &lt;strong&gt;the second actor wasn't CI at all&lt;/strong&gt;, so none of the existing concurrency guards between pipeline runs applied.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix: verify the digest, don't trust the tag (PR #1181)
&lt;/h2&gt;

&lt;p&gt;Two independent changes, both in &lt;code&gt;.github/workflows/deploy-environment.yml&lt;/code&gt;:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. Bounded retry on &lt;code&gt;start-deployment&lt;/code&gt;.&lt;/strong&gt; A &lt;code&gt;RUNNING&lt;/code&gt;-state rejection now retries for up to 20 minutes instead of failing on the first attempt — a wait can only delay a deploy, never corrupt one, so it's a safe response to "something else is deploying right now." But a retry alone doesn't fix the race; it just makes CI patient. The actual fix is next.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Re-read the tag's digest on every retry attempt, and refuse to deploy if it moved.&lt;/strong&gt; Before each &lt;code&gt;start-deployment&lt;/code&gt; call, CI re-reads what &lt;code&gt;production&lt;/code&gt; currently points to and compares it against the digest of the image &lt;em&gt;this run&lt;/em&gt; pushed. If they no longer match — because something else moved the tag in the interim — CI stops and errors loudly instead of deploying an image it can no longer identify:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;"The &lt;code&gt;production&lt;/code&gt; tag no longer points at the image this run built. Built &lt;code&gt;&amp;lt;digest&amp;gt;&lt;/code&gt;, tag now &lt;code&gt;&amp;lt;digest&amp;gt;&lt;/code&gt;. Something pushed to this tag outside this pipeline; deploying now would ship an unknown image."&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That single check closes the 18-second window: CI will now only ever call &lt;code&gt;start-deployment&lt;/code&gt; when it can prove the tag still means what it thinks it means.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. A post-swap check — warn-only, on purpose.&lt;/strong&gt; After the deployment completes, CI reads the digest App Runner actually deployed and compares it against what CI pushed. If they don't match, it logs a warning and keeps going rather than failing the pipeline.&lt;/p&gt;

&lt;p&gt;That's a deliberate asymmetry, and it's the design decision worth explaining: the &lt;em&gt;pre-swap&lt;/em&gt; digest check is a hard gate, because it's checkable before anything ships and a mismatch there means CI is about to deploy the wrong thing. The &lt;em&gt;post-swap&lt;/em&gt; check can't be a hard gate, because by the time it runs, the running container already &lt;strong&gt;is&lt;/strong&gt; what it is — the thing that ran ahead was the registry tag, not production. Failing the pipeline red at that point would block a legitimate deploy for a discrepancy that already happened and that the running service had no part in causing. Warn, don't block, on a fact you can no longer change.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that makes it worth writing about
&lt;/h2&gt;

&lt;p&gt;The merge that triggered this was &lt;strong&gt;docs-only&lt;/strong&gt;. The image that never got deployed differed from the one that did by markdown and release notes — nothing user-visible. Production was healthy and serving the whole time.&lt;/p&gt;

&lt;p&gt;That's luck, not design, and it's the actual point: the same race on a real frontend change ships a red pipeline &lt;strong&gt;and&lt;/strong&gt; production quietly serving the old build, and the red error points at App Runner deployment state — not at "prod is stale." You'd fix the wrong problem. &lt;code&gt;InvalidRequestException: isn't in RUNNING state&lt;/code&gt; is a true statement about App Runner. It was never a true statement about what mattered, which was that the tag we were about to deploy had already been consumed by someone else. A pipeline that goes red for the right reason but tells you the wrong story is worse than one that goes red honestly — you close the ticket on the error you can see instead of the one that actually happened.&lt;/p&gt;

&lt;p&gt;Boring fixes are the ones worth shipping: retry with a bound, verify a digest instead of trusting a tag, and only fail hard on the check that happens before anything ships.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;We hit this on our own production App Runner service, running the app you'd be scanning your AWS bill with. Curious what a read-only pass over your own account finds? Free, read-only AWS waste scan: &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;cloudcostwise.io&lt;/a&gt; — five minutes, no card.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cicd</category>
      <category>docker</category>
      <category>devops</category>
    </item>
    <item>
      <title>Catching a Cost Spike the Same Day It Happens: Inside Our Z-Score Anomaly Detector</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Thu, 27 Aug 2026 13:03:55 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/catching-a-cost-spike-the-same-day-it-happens-inside-our-z-score-anomaly-detector-4mmo</link>
      <guid>https://dev.to/cloudwiseteam/catching-a-cost-spike-the-same-day-it-happens-inside-our-z-score-anomaly-detector-4mmo</guid>
      <description>&lt;p&gt;AWS Cost Explorer's own data can lag up to 24 hours behind real usage, and nobody's on call for "the bill." A forgotten load test, a runaway Auto Scaling event, a script that didn't tear down its own infrastructure — by the time a human opens the console, whatever ran already ran.&lt;/p&gt;

&lt;p&gt;CloudWise's anomaly detector is built to close that gap: a daily scheduled scan, not a dashboard you have to remember to check. Here's the actual statistics behind it — the real thresholds from &lt;code&gt;lambdas/anomaly_detector/handler.py&lt;/code&gt; — and the exact test case, already in our shipped suite, that proves it catches a spike.&lt;/p&gt;

&lt;h2&gt;
  
  
  The mechanic: z-score, not a percentage
&lt;/h2&gt;

&lt;p&gt;The naive version of this feature checks "is today more than X% above yesterday?" That breaks immediately: a service that costs $2/day naturally swings 200% day to day out of pure noise, while a service that costs $4,000/day moving 15% is a real four-figure problem. Percentage-of-yesterday has no sense of a service's &lt;em&gt;normal&lt;/em&gt; variance.&lt;/p&gt;

&lt;p&gt;The detector instead computes a &lt;a href="https://en.wikipedia.org/wiki/Standard_score" rel="noopener noreferrer"&gt;z-score&lt;/a&gt; per AWS service, per account, every day:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# lambdas/anomaly_detector/handler.py
&lt;/span&gt;&lt;span class="n"&gt;SPIKE_THRESHOLD_STD_DEV&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;   &lt;span class="c1"&gt;# standard deviations from the mean
&lt;/span&gt;&lt;span class="n"&gt;MIN_ABSOLUTE_CHANGE_USD&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;5.0&lt;/span&gt;   &lt;span class="c1"&gt;# minimum dollar change to trigger, at all
&lt;/span&gt;&lt;span class="n"&gt;MIN_DATA_POINTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;7&lt;/span&gt;             &lt;span class="c1"&gt;# minimum days of history required
&lt;/span&gt;
&lt;span class="n"&gt;mean&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;statistics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;historical&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;std_dev&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;statistics&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stdev&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;historical&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="nf"&gt;len&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;historical&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;
&lt;span class="n"&gt;change&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;today_cost&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;mean&lt;/span&gt;
&lt;span class="n"&gt;z_score&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;today_cost&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;mean&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="n"&gt;std_dev&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;std_dev&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;

&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;z_score&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="n"&gt;SPIKE_THRESHOLD_STD_DEV&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="c1"&gt;# ...flag it
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;historical&lt;/code&gt; is the trailing daily costs for that service, excluding today. If a service doesn't have at least 7 days of history yet, it's skipped — no history means no baseline, and a false "anomaly" on day one of a brand-new resource is worse than useless. And &lt;strong&gt;before z-score even runs&lt;/strong&gt;, there's a $5 absolute-change floor: if today's cost moved by less than five dollars, the detector doesn't care how many standard deviations that represents. A service that normally costs $0.02/day moving to $0.11/day is a 450% swing, a huge z-score, and completely irrelevant to anyone's bill. The dollar floor is what keeps a statistically-driven detector from paging you about noise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Severity isn't binary
&lt;/h2&gt;

&lt;p&gt;Once something clears the 2.0 z-score bar, it gets bucketed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;SEVERITY_THRESHOLDS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;CRITICAL&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;4.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;HIGH&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;3.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;MEDIUM&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;2.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;LOW&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Yes, &lt;code&gt;LOW&lt;/code&gt; is below the 2.0 trigger threshold in the table — that band exists for a second call site that widens the net for manual/on-demand scans. The scheduled daily job only ever emits &lt;code&gt;MEDIUM&lt;/code&gt; and up, because the scheduled job's trigger &lt;em&gt;is&lt;/em&gt; 2.0.) The severities aren't decoration — they're what a Shield-tier customer's Slack alert leads with, and what determines whether the message reads as "worth a look this week" or "look now."&lt;/p&gt;

&lt;h2&gt;
  
  
  The test that proves it
&lt;/h2&gt;

&lt;p&gt;Rather than a hypothetical dollar figure, here's the exact fixture from our own shipped test suite (&lt;code&gt;lambdas/anomaly_detector/test_handler.py::test_detects_anomaly_with_high_z_score&lt;/code&gt;) — real, committed, running in CI on every change to this file:&lt;/p&gt;

&lt;p&gt;An EC2 line item costs $9–$11/day for seven straight days: $10, $11, $9, $10.5, $9.5, $10, $10. Mean &lt;strong&gt;$10.00&lt;/strong&gt;, standard deviation &lt;strong&gt;$0.65&lt;/strong&gt;. Then it jumps to &lt;strong&gt;$100&lt;/strong&gt; in a day.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;change     = 100 - 10        = $90
z_score    = (100 - 10) / 0.65 = 139.4
percentage = 90 / 10 * 100     = 900%
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A z-score of 139.4 isn't a borderline call by any measure — the test asserts the detector flags it as &lt;code&gt;CRITICAL&lt;/code&gt; or &lt;code&gt;HIGH&lt;/code&gt; severity, and it does, deterministically, every run. The absolute dollars here are small on purpose: it's a unit test, not a customer account. The statistical shape — a service ten times its normal cost, standard deviations off its own baseline — is exactly what the 2.0 threshold exists to catch on day one, whether the line item is $10 or $10,000.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is Shield-tier, and why it's read-only
&lt;/h2&gt;

&lt;p&gt;The detector runs once a day via EventBridge, over the last 30 days of each Shield-tier account's cost data, and writes findings to an alerts table before dispatching Slack and email notifications. It doesn't touch anything in your AWS account — no remediation, no API calls beyond reading Cost Explorer data CloudWise already ingests for the dashboard. That's deliberate: detecting a spike and &lt;em&gt;deciding what to do about it&lt;/em&gt; are different problems, and this one only does the first.&lt;/p&gt;

&lt;p&gt;The mechanic is boring on purpose — mean, standard deviation, a threshold, a dollar floor. Boring is what you want from something that's deciding whether to interrupt your weekend.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>observability</category>
      <category>python</category>
    </item>
    <item>
      <title>I Shipped a Redesign That Orphaned My Own Product</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:53:46 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/i-shipped-a-redesign-that-orphaned-my-own-product-5g74</link>
      <guid>https://dev.to/cloudwiseteam/i-shipped-a-redesign-that-orphaned-my-own-product-5g74</guid>
      <description>&lt;p&gt;This month I rebuilt CloudWise around an AI agent.&lt;/p&gt;

&lt;p&gt;Not a chatbot bolted into a corner. The agent &lt;strong&gt;is&lt;/strong&gt; the product now. You open the app and you're talking to it. "Where's my money going?" — and it pulls your AWS spend, ranks the waste, shows you the dollars per month per finding, and tells you what's safe to fix.&lt;/p&gt;

&lt;p&gt;It looked done. The screens were beautiful. The agent answered. I shipped it as the front door.&lt;/p&gt;

&lt;p&gt;Then I used it like a customer would. And I discovered I had quietly orphaned half my own product.&lt;/p&gt;

&lt;p&gt;This is the story of that month, because it taught me the single most useful lesson I've learned as a solo founder: &lt;strong&gt;a redesign isn't done when it looks done.&lt;/strong&gt;&lt;/p&gt;




&lt;h2&gt;
  
  
  What I built
&lt;/h2&gt;

&lt;p&gt;For most of this year, CloudWise looked like every other AWS cost tool: a dashboard. Charts, tables, a sidebar of pages. Useful, but passive. You had to know what to look for.&lt;/p&gt;

&lt;p&gt;The bet I made for the overhaul was that the &lt;em&gt;interface&lt;/em&gt; should be a conversation, not a dashboard. Most people don't want to read a cost dashboard. They want to ask a question and get an answer.&lt;/p&gt;

&lt;p&gt;So the centerpiece became a conversational workspace backed by a real tool-calling agent loop on Claude (running on AWS Bedrock). "Real" is the important word. This isn't a model that summarizes a page of text. It's a model with tools — it actually queries your cost data, your findings, your Reserved Instance and Savings Plan coverage, and composes an answer from live numbers. It remembers your account between conversations.&lt;/p&gt;

&lt;p&gt;You can ask it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Where's my money going?" → it breaks down spend by service, in gold, with the month-over-month delta.&lt;/li&gt;
&lt;li&gt;"What's safe to fix?" → it ranks your waste findings: idle NAT Gateways, forgotten SageMaker notebooks, unattached EBS snapshots, each with a dollar figure and a safe-to-fix flag.&lt;/li&gt;
&lt;li&gt;"Compare my two accounts." → and it does.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That last one matters more than it sounds, and it's where the trouble started.&lt;/p&gt;




&lt;h2&gt;
  
  
  The part where it looked finished
&lt;/h2&gt;

&lt;p&gt;By the start of June, the redesign was, on paper, complete. New design system. Dark, native. A conversational workspace. A guided tour. A redesigned dashboard. The agent loop. Cross-session memory. I'd migrated every screen.&lt;/p&gt;

&lt;p&gt;I flipped the new workspace to be the post-login front door and moved on.&lt;/p&gt;

&lt;p&gt;Here's the uncomfortable truth about building alone: &lt;strong&gt;you stop seeing your own product.&lt;/strong&gt; You navigate it the way the author navigates it — from the inside, knowing every shortcut, never actually starting cold the way a real user does.&lt;/p&gt;

&lt;p&gt;So I made myself start cold. I logged in like a brand-new customer and tried to do the boring things. Switch to my other AWS account. Open settings. Change a notification. Log out.&lt;/p&gt;

&lt;p&gt;I couldn't.&lt;/p&gt;




&lt;h2&gt;
  
  
  The front door had hidden the house
&lt;/h2&gt;

&lt;p&gt;None of those features were gone. Every settings page still existed. The account switcher still existed. Logout still existed. The old navigation still existed, sitting in the codebase, fully functional.&lt;/p&gt;

&lt;p&gt;They just weren't &lt;em&gt;reachable&lt;/em&gt; from the place users now landed.&lt;/p&gt;

&lt;p&gt;The new workspace shell had a clean little user footer — a static label with the account name. No menu. No logout. No link to subscription, notifications, alerts, API keys, AWS accounts, or password. From the new front door, there was literally no path to any of it.&lt;/p&gt;

&lt;p&gt;Worse, the half of the product that &lt;em&gt;did&lt;/em&gt; still have navigation rendered in the &lt;strong&gt;old&lt;/strong&gt; chrome. So a new user would land in a slick dark conversational workspace, click one thing, and get bounced into the previous design — a completely different layout. Two products wearing different clothes, stitched together at a seam the user falls straight through.&lt;/p&gt;

&lt;p&gt;I wrote it down plainly in my own audit doc at the time, because I needed to see it without flinching:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The redesign did not delete your features — the cutover orphaned them.&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;That sentence reframed the entire month. This wasn't a teardown. It was a &lt;em&gt;finishing&lt;/em&gt; problem. The work wasn't to rebuild — it was to take ownership of every job the old interface used to do, and make the new one do it better.&lt;/p&gt;




&lt;h2&gt;
  
  
  The agent was also lying about multi-account
&lt;/h2&gt;

&lt;p&gt;While I was in there, I found a real bug — the kind that only surfaces when you use the product for real, with more than one account.&lt;/p&gt;

&lt;p&gt;CloudWise is a multi-account tool. You connect your AWS Organization and it discovers all your accounts. But the agent was collapsing them. When you asked about cost, it blended every account together with no way to scope to one. And when you asked about Reserved Instance and Savings Plan coverage, it reported coverage for &lt;strong&gt;exactly one account&lt;/strong&gt; — the first one it happened to grab — and silently ignored the rest.&lt;/p&gt;

&lt;p&gt;That's not a cosmetic miss. That's the tool confidently giving you a wrong answer about money. If you had three accounts and asked "how's my RI coverage," it answered for one-third of your footprint and didn't tell you.&lt;/p&gt;

&lt;p&gt;The backend could already filter by account. The old reports page could already do multi-account selection. The &lt;em&gt;agent&lt;/em&gt; — the new centerpiece — was strictly less capable than both. The redesign had, in this one spot, made the product worse while looking like it made it better.&lt;/p&gt;

&lt;p&gt;So the fix had three parts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;A header account switcher&lt;/strong&gt; that scopes the entire workspace — single or multi-select, capped by your plan tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Per-account arguments on the agent's tools&lt;/strong&gt;, so the model can honor that switcher &lt;em&gt;and&lt;/em&gt; answer in-conversation requests like "compare my two accounts."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Fixing the coverage bug&lt;/strong&gt; so it reports across all accounts, not the first one.
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Before: commitments reported for whatever account came first.
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_commitments&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;accounts&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;first&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;accounts&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;fetch_coverage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;first&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;   &lt;span class="c1"&gt;# silently ignores the rest
&lt;/span&gt;
&lt;span class="c1"&gt;# After: the agent scopes to what you asked for — one, some, or all.
&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;get_commitments&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;accounts&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;scope_ids&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;targets&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;accounts&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_id&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;scope_ids&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;scope_ids&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="n"&gt;accounts&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nf"&gt;fetch_coverage&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;account_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;a&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;targets&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A small diff. A big difference in whether you can trust the answer.&lt;/p&gt;




&lt;h2&gt;
  
  
  Thirty days of taking ownership
&lt;/h2&gt;

&lt;p&gt;That's what the month actually was. Not "add features." Re-own the basics, and make the agent the real front door instead of a beautiful demo sitting on top of a product it had stopped being responsible for.&lt;/p&gt;

&lt;p&gt;In order, here's what changed:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;One app shell, not two.&lt;/strong&gt; Unified chrome around every authenticated page — one persistent sidebar, the account switcher, every settings link, and logout — so you never fall through the seam between the new workspace and the old layout again. The workspace became one surface &lt;em&gt;inside&lt;/em&gt; a consistent app, not a separate world.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The account switcher scopes everything.&lt;/strong&gt; Pick an account (or several) in the header and the whole workspace re-scopes to it — and the agent answers for exactly that selection.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The agent learned to count accounts.&lt;/strong&gt; Per-account tool arguments, the coverage bug fixed, and the ability to genuinely compare accounts on request.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A deep, cross-page guided tour.&lt;/strong&gt; Instead of a checklist nobody reads, a docked "CloudWise agent" companion walks a new user through the real product — navigating the actual screens across seventeen steps, not narrating a slideshow.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The whole product went dark-mode-native.&lt;/strong&gt; Every deep page — cost reports, remediation, savings plans, settings, the dashboard — rebuilt on one design system. Money is always gold. No more half-migrated screens where one page is dark and the next is white.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A shareable cost-health score.&lt;/strong&gt; One number, 0–100, for how efficiently you're running AWS — with a public, sanitized share card carrying no account details or PII. Because the first question every engineer actually asks is "are we good, or not?"&lt;/p&gt;

&lt;p&gt;Across the month that came out to a few dozen shipped changes and around fifty releases — versions 1.54 through 1.104. Most of them were not glamorous. Most of them were re-owning a job the old interface used to do, quietly, that the new one had dropped on the floor.&lt;/p&gt;




&lt;h2&gt;
  
  
  The lesson, stated plainly
&lt;/h2&gt;

&lt;p&gt;I've shipped a lot of software in twenty-five years. I still fell for this one, which is why I think it's worth writing down.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;"Looks done" and "is the product" are not the same state.&lt;/strong&gt; They can be a hundred releases apart.&lt;/p&gt;

&lt;p&gt;A redesign is not finished when the new screens are beautiful and the happy path works. It's finished when the new thing has quietly taken over &lt;em&gt;every&lt;/em&gt; job the old thing did — logout, account switching, the boring settings page nobody screenshots — and does each of them at least as well. Until then you don't have a new product. You have a beautiful front door on a house whose rooms you've locked.&lt;/p&gt;

&lt;p&gt;And the cruelest part: &lt;strong&gt;users don't grade you on the demo.&lt;/strong&gt; They grade you on the one ordinary day they need the single feature you forgot to bring across. The day they need to switch accounts, or check their RI coverage across all three, or just sign out. That's the moment your redesign is actually judged — and it's never the moment you rehearsed.&lt;/p&gt;

&lt;p&gt;So now I have a rule. Before I call any cutover done, I log in cold and do the boring things. All of them. If I can't sign out, I'm not done. No matter how good the front door looks.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;CloudWise is an AWS cost optimization tool for startups — 191 automated waste checks, a real agent that runs against your actual usage, air-gapped mode for security teams, starting at $19/month. If you want to ask an agent where your AWS money is going, it's at &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;cloudcostwise.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>buildinpublic</category>
      <category>saas</category>
    </item>
    <item>
      <title>EBS Snapshot Sprawl: The Waste Cost Explorer Can't Show You</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:51:19 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/ebs-snapshot-sprawl-the-waste-cost-explorer-cant-show-you-5had</link>
      <guid>https://dev.to/cloudwiseteam/ebs-snapshot-sprawl-the-waste-cost-explorer-cant-show-you-5had</guid>
      <description>&lt;p&gt;Last week's short covered the 101 on old EBS snapshots: the 90-day threshold, the $0.05/GB-month rate, &lt;code&gt;aws ec2 delete-snapshot&lt;/code&gt;. If you saw it, you already know snapshots are cheap-per-unit and expensive-in-aggregate. What that 42 seconds couldn't fit in is the actual reason snapshot sprawl is so hard to find in the first place — and it isn't the price. It's that Cost Explorer structurally cannot show you who's responsible for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost Explorer shows you a bill, not a culprit
&lt;/h2&gt;

&lt;p&gt;Group your AWS costs by usage type and EBS snapshots show up as one line: &lt;code&gt;EBS:SnapshotUsage&lt;/code&gt;, rolled up per account and region. That's it. Not per-snapshot, not per-volume, not per-AMI. If that line is $340/month, Cost Explorer will tell you the total and nothing about which of your 200 snapshots — or which of your 40 AMIs pinning them — put it there.&lt;/p&gt;

&lt;p&gt;Part of why this is so opaque is how the billing actually works. EBS snapshots are &lt;strong&gt;incremental&lt;/strong&gt;: the first snapshot of a volume captures every block, but every snapshot after that only stores blocks that changed since the previous one. Delete an "old" snapshot in the middle of a chain and AWS doesn't just drop its unique blocks — it can merge the still-referenced blocks from that snapshot into the next one to keep the chain valid. The result is a genuinely well-designed storage model that happens to make per-snapshot cost attribution close to meaningless from the outside. You can't look at snapshot #14 in a chain of 20 and know what deleting it actually frees, and Cost Explorer doesn't even try — it just gives you the account-wide sum and moves on.&lt;/p&gt;

&lt;p&gt;So the number you see going up every month is real. The reason is invisible from the billing console. You have to go look at the resources directly.&lt;/p&gt;

&lt;h2&gt;
  
  
  &lt;code&gt;deregistered ≠ deleted&lt;/code&gt; — the AMI-pinning trap
&lt;/h2&gt;

&lt;p&gt;Here's the gap that costs teams the most and gets found the least: deregistering an AMI does not delete the snapshot backing it.&lt;/p&gt;

&lt;p&gt;When you run &lt;code&gt;CreateImage&lt;/code&gt; against an EC2 instance, AWS creates an AMI and, silently, one or more EBS snapshots to back it — you'll find the linkage in the snapshot's own description field, something like &lt;code&gt;Created by CreateImage(i-0abc123def456789) for ami-0fedcba987654321&lt;/code&gt;. Deregister that AMI later (cleaning up an old release, retiring a pipeline, whatever the reason) and AWS removes the AMI. The snapshot stays. Forever. Nothing in the console flags it, nothing in Cost Explorer changes shape, and nothing tells you the thing that snapshot was created for no longer exists.&lt;/p&gt;

&lt;p&gt;This is exactly the pattern our &lt;code&gt;AMI_ORPHANED_SNAPSHOT&lt;/code&gt; detector checks for. It's not a guess — it's a direct read of the relationship AWS itself records: the detector regex-matches each snapshot's description against the &lt;code&gt;CreateImage(...)  for (ami-...)&lt;/code&gt; pattern, pulls the referenced AMI ID, and checks it against the account's currently-registered AMIs. If the AMI isn't there anymore, the snapshot is flagged — at higher priority than a generic "old snapshot" check, because an orphaned-AMI snapshot has a &lt;strong&gt;certain&lt;/strong&gt; reason to be dead, not just an age-based guess.&lt;/p&gt;

&lt;p&gt;That priority ordering matters. Our storage detector dedups three overlapping checks against the same snapshot inventory: AMI-orphan first, then volume-orphan (the source volume was deleted), then plain age (over the 90-day default threshold, at the $0.05/GB-month rate from last week's short). A snapshot only gets counted once, under whichever explanation is strongest — an AMI-orphaned snapshot isn't also reported as merely "old," because "old" undersells why it's actually safe to delete.&lt;/p&gt;

&lt;h2&gt;
  
  
  Lifecycle policies are the actual fix — for snapshots you haven't made yet
&lt;/h2&gt;

&lt;p&gt;None of the above is a criticism of EBS snapshots as a backup mechanism. They're cheap, they're incremental, and they're the right tool. The problem is entirely operational: nothing deletes them automatically unless you tell it to.&lt;/p&gt;

&lt;p&gt;AWS Data Lifecycle Manager (DLM) exists for exactly this — attach a policy to a tag or resource type and it will create snapshots on a schedule and expire them on a schedule, so "backup taken 400 days ago for a server that's been gone for 399 of them" stops being possible going forward. If you're not running DLM policies today, that's the highest-leverage 20-minute fix here, full stop.&lt;/p&gt;

&lt;p&gt;But DLM only prevents new sprawl. It does nothing for the snapshots already sitting in your account from AMIs someone deregistered two years ago, or backups nobody automated before DLM was set up. That backlog needs to be found once, by hand or by a scan, before a policy can keep it clean going forward. That's the gap our detector is built for — not a replacement for lifecycle policies, the thing that finds what predates them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding yours
&lt;/h2&gt;

&lt;p&gt;If you want to see this in your own account instead of grepping AMI descriptions by hand, &lt;a href="https://cloudcostwise.io?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=ebs-snapshot-sprawl-waste-cost-explorer-cant-show-you" rel="noopener noreferrer"&gt;a free scan&lt;/a&gt; checks this along with 190+ other waste patterns across 40+ AWS services — read-only, five minutes, nothing gets deleted without you clicking it. If you'd rather watch the 90-day/$0.05-per-GB basics first, we posted a 30-second walkthrough of that detector last week.&lt;/p&gt;

&lt;p&gt;Or check the AMI-pinning trap yourself right now:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Snapshots whose description names a CreateImage/AMI pairing&lt;/span&gt;
aws ec2 describe-snapshots &lt;span class="nt"&gt;--owner-ids&lt;/span&gt; self &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Snapshots[?contains(Description, 'CreateImage')].[SnapshotId,Description]"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; table

&lt;span class="c"&gt;# Currently-registered AMI IDs&lt;/span&gt;
aws ec2 describe-images &lt;span class="nt"&gt;--owners&lt;/span&gt; self &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--query&lt;/span&gt; &lt;span class="s2"&gt;"Images[].ImageId"&lt;/span&gt; &lt;span class="nt"&gt;--output&lt;/span&gt; text
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cross-reference the AMI ID in each snapshot's description against the second list. Anything missing has been quietly billing you since the day someone cleaned up an AMI and assumed the cleanup was done.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
    <item>
      <title>The Bug That Turned Every Bad Password Into a Server Outage</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:50:30 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/the-bug-that-turned-every-bad-password-into-a-server-outage-4dda</link>
      <guid>https://dev.to/cloudwiseteam/the-bug-that-turned-every-bad-password-into-a-server-outage-4dda</guid>
      <description>&lt;p&gt;Type your password wrong on CloudWise's login page, and for a while, the server told you it had a nervous breakdown.&lt;/p&gt;

&lt;p&gt;Not "invalid credentials." Not even a plain 401. An HTTP &lt;strong&gt;500&lt;/strong&gt; — the code reserved for "something on our end is broken" — for the most routine failure mode there is: a human mistyping a password.&lt;/p&gt;

&lt;p&gt;This happened on staging. It happened on production. It happened to &lt;code&gt;POST /api/v1/auth/login&lt;/code&gt; with wrong creds, an unknown email, an unconfirmed account, or an account mid-password-reset. Four different "you did something ordinary" situations, all reported back as "we did something wrong." It went undetected until an automated gate caught it and refused to let a release ship. This is that bug, the fix, and why the distinction matters more than it sounds like it should.&lt;/p&gt;

&lt;h2&gt;
  
  
  The bug
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;_handle_cognito_auth_error&lt;/code&gt;, the function that translates AWS Cognito's auth exceptions into an HTTP response, raised a &lt;code&gt;CloudWiseException&lt;/code&gt; without ever setting a &lt;code&gt;status_code&lt;/code&gt;. No status code means the default. The default is 500.&lt;/p&gt;

&lt;p&gt;So every one of these:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Wrong password&lt;/li&gt;
&lt;li&gt;Email that isn't registered&lt;/li&gt;
&lt;li&gt;Account that hasn't confirmed its email yet&lt;/li&gt;
&lt;li&gt;Account stuck in a forced password reset&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;came back as a server error. Reproduced live on both environments — for example, &lt;code&gt;POST /api/v1/auth/login&lt;/code&gt; with bad credentials on staging returned:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;HTTP 500
{"detail":"Invalid email or password"}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;and on production:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;HTTP 500
{"detail":"Authentication failed: User does not exist."}
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read that second one again. Production wasn't just returning the wrong status code — it was telling an anonymous caller whether a given email address had an account. That's a second, smaller bug riding along inside the first one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the status code is the actual bug
&lt;/h2&gt;

&lt;p&gt;It's tempting to shrug this off — the message was right there in the body, &lt;code&gt;"Invalid email or password"&lt;/code&gt;, so what's the harm in the wrong three-digit prefix?&lt;/p&gt;

&lt;p&gt;The harm is that the status code isn't decoration. It's the part of the response that infrastructure reads without understanding a word of the payload. A monitoring dashboard doesn't parse &lt;code&gt;detail&lt;/code&gt;. It counts 5xx rates. To CloudWise's own alerting, every single mistyped password looked identical to an actual outage — same bucket, same page-worthy signal, same "something is on fire" pattern, forever, at whatever rate normal users normally fat-finger their passwords. That's not a rare event. It's baseline noise, and the bug was quietly dressing it up as baseline crisis.&lt;/p&gt;

&lt;p&gt;It also meant every one of those routine rejections logged at &lt;code&gt;error&lt;/code&gt; level — the log level meant for things that need a human to look at them, now firing constantly for people who just needed to try again.&lt;/p&gt;

&lt;p&gt;And then it did something worse than annoy a dashboard: it blocked a production release.&lt;/p&gt;

&lt;h2&gt;
  
  
  How it actually got caught
&lt;/h2&gt;

&lt;p&gt;CloudWise's release pipeline runs an E2E gate against staging before anything is allowed to promote to production. One of the negative-path assertions in that gate — &lt;code&gt;ci-gate-read.spec.ts:93&lt;/code&gt; — logs in with a wrong password and checks that the app shows an error and stays on &lt;code&gt;/auth/login&lt;/code&gt;, the way a real login form should behave.&lt;/p&gt;

&lt;p&gt;That test didn't expect a 500. Nothing in a normal auth flow should return one for a wrong password. The gate went red, and release &lt;code&gt;1.105.0&lt;/code&gt; sat there, unable to auto-promote to production, because the pipeline correctly refused to trust a build where the login form's error handling was behaving like a crash.&lt;/p&gt;

&lt;p&gt;That's the part worth sitting with: this bug was already live in production, unnoticed, for who knows how long. It took a &lt;em&gt;different&lt;/em&gt; release's gate run to surface it — not because the gate was looking for this specific bug, but because it was asserting the right general behavior (bad password → clean error, stay put) and the actual behavior didn't match. A negative-path test doesn't need to know about your bug in advance. It just needs to check that the ordinary failure case looks ordinary.&lt;/p&gt;

&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;The fix, in &lt;code&gt;backend/app/services/cognito_auth_service.py&lt;/code&gt;, maps Cognito's exceptions to the status codes they should have had all along:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;NotAuthorizedException&lt;/code&gt; (wrong password) and &lt;code&gt;UserNotFoundException&lt;/code&gt; (unknown email) → &lt;strong&gt;401&lt;/strong&gt;, both returning the exact same message: &lt;code&gt;"Invalid email or password"&lt;/code&gt;. Same message for both cases on purpose — the response can no longer be used to tell whether an email is registered, closing that account-enumeration leak.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;UserNotConfirmedException&lt;/code&gt; and &lt;code&gt;PasswordResetRequiredException&lt;/code&gt; → &lt;strong&gt;403&lt;/strong&gt; — the account exists, but the request is correctly rejected for a reason that isn't "try a different password."&lt;/li&gt;
&lt;li&gt;Anything else — a genuinely unexpected Cognito error — still defaults to 500. That default is correct in that case. An unrecognized failure mode &lt;em&gt;is&lt;/em&gt; a server-side concern worth alerting on. The bug was never that 500 existed; it was that the four most common, most expected auth rejections were routed into it by omission.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Expected-rejection logging dropped from &lt;code&gt;error&lt;/code&gt; to &lt;code&gt;warning&lt;/code&gt;, so the logs now reflect what actually happened: a routine, anticipated rejection, not an incident.&lt;/p&gt;

&lt;p&gt;The test suite got the fix that should have caught this the first time. &lt;code&gt;test_handle_cognito_auth_error_*&lt;/code&gt; previously asserted on the error &lt;em&gt;message&lt;/em&gt; only — never the status code. That's exactly how a 500-instead-of-401 slips through code review and CI both: the message text looked fine, so nobody noticed the number in front of it was wrong. The tests now assert status codes explicitly, plus two cases that weren't covered before: the no-leak &lt;code&gt;UserNotFoundException&lt;/code&gt; path, and &lt;code&gt;PasswordResetRequiredException&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is the sibling of an earlier fix, CLO-48, which cleaned up the same bug class on the refresh-token path — an expired refresh token was logging as &lt;code&gt;ERROR&lt;/code&gt; twice when a clean 401 was all that was warranted. That one didn't touch login. This one closes the login path CLO-48 didn't cover.&lt;/p&gt;

&lt;h2&gt;
  
  
  The takeaway
&lt;/h2&gt;

&lt;p&gt;Status codes are a contract between your API and everything that reads it without understanding it — monitors, alerting rules, retry logic, browsers, other services. A wrong password is a client error. Treating it as a server error doesn't just mislabel one response; it teaches every downstream system that watches your 5xx rate to distrust the signal, right when you need that signal to mean something.&lt;/p&gt;

&lt;p&gt;The thing that actually caught this wasn't a code review, a manual QA pass, or a customer complaint. It was an automated negative-path assertion doing exactly what negative-path assertions are for: checking that the boring, expected failure looks boring and expected. It's staying in the gate. Nobody has to remember to test this again — the pipeline already refuses to ship a build that gets it wrong.&lt;/p&gt;

&lt;p&gt;If you want to see what else that kind of scrutiny turns up, &lt;a href="https://cloudcostwise.io?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=login-returned-500-bug-story" rel="noopener noreferrer"&gt;a free, read-only scan&lt;/a&gt; of your AWS account takes about five minutes and changes nothing without you approving it first.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>testing</category>
      <category>webdev</category>
      <category>buildinpublic</category>
    </item>
    <item>
      <title>Inside a Real $11,871/mo AWS Waste Scan</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:50:30 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/inside-a-real-11871mo-aws-waste-scan-5536</link>
      <guid>https://dev.to/cloudwiseteam/inside-a-real-11871mo-aws-waste-scan-5536</guid>
      <description>&lt;p&gt;Most cost-optimization content shows you a screenshot of a dashboard and asks you to trust the number. So instead, here's an actual scan CloudWise runs internally as a demo profile — a mid-size SaaS company's AWS account, $28,500/mo in total spend, run through all 191 detectors. It came back with &lt;strong&gt;$11,871/mo in identified savings across 18 findings.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;I want to walk through what's actually in that list, because the shape of it surprised me even after building the detectors.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the 18 findings live
&lt;/h2&gt;

&lt;p&gt;Split by category, not dollar amount:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Database — 9 findings.&lt;/strong&gt; By far the largest bucket. A mix of idle RDS instances, an oversized analytics database running at 12% CPU, stale manual snapshots, and a cluster of ElastiCache-specific findings (idle replicas, an engine migration opportunity, a serverless-fit opportunity, and one very large data-tiering opportunity — more on that below).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Compute — 2 findings.&lt;/strong&gt; An idle bastion host nobody SSHs into anymore, and one oversized EC2 fleet running at 15% CPU.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Storage — 2 findings.&lt;/strong&gt; An unattached EBS volume left over from a migration, and a batch of aging EBS snapshots.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commitment Management — 2 findings.&lt;/strong&gt; A Savings Plan expiring in 52 days, and a Convertible Reserved Instance sitting on previous-generation hardware with a free exchange available.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network — 2 findings.&lt;/strong&gt; Two idle Global Accelerators — one genuinely idle, one just disabled but still billing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Purchase Optimization — 1 finding.&lt;/strong&gt; Production RDS instances that have been running on-demand, 24/7, for 90+ days with no Reserved Instance covering them.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nine of eighteen findings sitting in "Database" isn't a coincidence — it's the category where the widest range of waste patterns overlap: idle instances, oversized instances, stale snapshots, and now a whole sub-family of ElastiCache-specific checks (engine choice, replication topology, traffic shape, data tiering) that didn't exist as separate detectors a year ago.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one number that matters most
&lt;/h2&gt;

&lt;p&gt;Of the $11,871/mo total, &lt;strong&gt;one single finding accounts for 68% of it: $8,057/mo&lt;/strong&gt;, from an ElastiCache data-tiering opportunity.&lt;/p&gt;

&lt;p&gt;The cluster in question runs 4× &lt;code&gt;cache.r6g.16xlarge&lt;/code&gt; nodes — memory-only, no local SSD — for 1,676 GiB of total cache capacity. ElastiCache's R6gd family adds local NVMe storage and &lt;em&gt;tiers&lt;/em&gt; data between RAM and SSD automatically, based on access frequency. For a dataset where most of the data isn't accessed on every request (true of almost every real cache), that means the same effective capacity fits on far less hardware: 1× &lt;code&gt;cache.r6gd.16xlarge&lt;/code&gt; instead of 4× &lt;code&gt;cache.r6g.16xlarge&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That's not a "you forgot to delete something" finding. It's an architecture-level rightsizing call that requires reading how your own cache is actually used before you touch it — which is exactly why it's flagged as &lt;code&gt;confidence: medium&lt;/code&gt;, not &lt;code&gt;high&lt;/code&gt;, and comes with an explicit risk note: SSD-resident data has slightly higher latency, so it's a good fit only when a meaningful chunk of your data is cold.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the next four biggest opportunities have in common
&lt;/h2&gt;

&lt;p&gt;Set aside the $8,057/mo ElastiCache finding and look at what else shows up near the top of this scan: an expiring Savings Plan, an unpurchased Reserved Instance opportunity on production RDS, an oversized analytics database at 12% CPU, and an oversized EC2 fleet at 15% CPU.&lt;/p&gt;

&lt;p&gt;None of those are "click delete." Every one of them is a standing decision that has to be made again, deliberately, on some recurring cadence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A Savings Plan doesn't renew itself — someone has to decide, before it lapses, whether the workload underneath it still looks the same.&lt;/li&gt;
&lt;li&gt;An RDS instance rightsized today can be wrong again in six months if the workload grows, shrinks, or changes shape.&lt;/li&gt;
&lt;li&gt;"Buy the Reserved Instance" is itself a bet on the next 12 months looking like the last 3 — which is exactly the kind of call &lt;a href="https://dev.to/blog/aws-reserved-instance-commitment-risk"&gt;CloudWise's Commitment Risk Score&lt;/a&gt; exists to check before you make it.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compare that to the &lt;em&gt;idle&lt;/em&gt;-resource findings in this same scan — the bastion host, the unattached EBS volume, the stale snapshots, the disabled accelerators. Those are real money too, and they're the easiest to fix: find it, delete it, done. But they're structurally small, because once you delete something, it stays deleted. It doesn't come back next quarter.&lt;/p&gt;

&lt;p&gt;Rightsizing and commitment decisions aren't like that. The right instance size and the right commitment level are correct &lt;em&gt;for right now&lt;/em&gt; — and they drift the moment your traffic pattern, team, or architecture changes, which for a growing company is constantly. That's the actual thesis worth taking from a scan like this: &lt;strong&gt;waste renews itself.&lt;/strong&gt; The idle-resource sweep is a one-time cleanup. The rightsizing and commitment layer is a standing job, and it's where most of the money in this particular scan actually lives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for how you should read &lt;em&gt;your&lt;/em&gt; scan
&lt;/h2&gt;

&lt;p&gt;If you run a scan and the top finding is "delete this idle volume," fix it and move on — there's nothing recurring about it. But if your biggest findings look more like this account's — an oversized instance, an expiring commitment, a data-tiering opportunity that requires understanding your own traffic — treat that as a signal that the fix isn't a single action, it's a review cadence you don't currently have. That's a different kind of problem, and it's worth naming as one.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;CloudWise is an AWS cost optimization tool for startups — 191 automated waste checks across 42 AWS services, read-only by design, starting at $19/month. Run a free scan at &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;cloudcostwise.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
    <item>
      <title>We Cut Lambda Cold Starts 56% — Three Wrong Turns Before the Real SnapStart Fix</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 19 Aug 2026 13:43:07 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/we-cut-lambda-cold-starts-56-three-wrong-turns-before-the-real-snapstart-fix-1bd1</link>
      <guid>https://dev.to/cloudwiseteam/we-cut-lambda-cold-starts-56-three-wrong-turns-before-the-real-snapstart-fix-1bd1</guid>
      <description>&lt;p&gt;CloudWise's &lt;code&gt;/dashboard&lt;/code&gt; took up to 2.3 seconds to load cold. That's the API Lambda's own cliff — not network, not the frontend. AWS Lambda SnapStart was already turned on. It was already restoring a frozen snapshot in about 600ms, which is the number SnapStart is supposed to deliver. And for two deploys, turning it on made no measurable difference to the thing that actually made the dashboard feel slow. This is the debugging path that got from there to a measured 56% cut — including the two attempts that didn't work, because the wrong-turn part is the part worth reading.&lt;/p&gt;

&lt;h2&gt;
  
  
  What was actually slow
&lt;/h2&gt;

&lt;p&gt;We measured every layer before touching anything (full numbers in &lt;code&gt;docs/redesign/clo114-dashboard-speed-analysis.md&lt;/code&gt;). Four dashboard API calls, already firing in parallel — parallelizing them further was a non-issue, wall-clock already tracked the slowest call. Server compute, warm, was fine: ~200–380ms. SnapStart's own restore was fine: ~550–620ms.&lt;/p&gt;

&lt;p&gt;The problem was one specific number from CloudWatch &lt;code&gt;REPORT&lt;/code&gt; lines: on a cold (post-restore) invocation, the first request's own &lt;code&gt;Duration&lt;/code&gt; — not the restore, the handler work after the restore — was &lt;strong&gt;1,124ms at p50, up to 2,805ms at the tail&lt;/strong&gt;. SnapStart had already paid for the expensive part (a frozen Python interpreter + imports) and handed back a restored environment in 600ms. Something in the first real request was still doing ~1,100ms of work the snapshot didn't cover.&lt;/p&gt;

&lt;h2&gt;
  
  
  The diagnosis: init phase vs. handler phase
&lt;/h2&gt;

&lt;p&gt;SnapStart snapshots whatever ran at &lt;em&gt;module import / init&lt;/em&gt;. Anything deferred to the request handler is still paid, lazily, on the first real invocation — that's the whole cliff. &lt;code&gt;get_settings()&lt;/code&gt; already runs at import in &lt;code&gt;app/main.py&lt;/code&gt;, and constructing &lt;code&gt;Settings&lt;/code&gt; performs the full AWS Parameter Store load — so config reads were already inside the snapshot, free. What wasn't: the boto3 session and DynamoDB client (built lazily in a factory, inside the request path) and FastAPI's &lt;code&gt;startup_event&lt;/code&gt; — Mangum runs &lt;code&gt;lifespan="auto"&lt;/code&gt; on the first request, not at import.&lt;/p&gt;

&lt;p&gt;The fix looked obvious: move both into the init phase, so they're captured in the snapshot. We rejected the alternative — Lambda provisioned concurrency — outright: it bills a kept-warm 2GB instance 24/7 regardless of traffic, and it's redundant with SnapStart, which already restores in ~600ms. Paying a standing bill to paper over a problem SnapStart half-solves is exactly the kind of waste we built this product to catch in other people's accounts. So: init-phase priming, &lt;code&gt;~$0&lt;/code&gt; incremental cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong turn #1: it shipped clean and did nothing
&lt;/h2&gt;

&lt;p&gt;PR #641 landed the priming code. Deploy went green, gates passed, prod promoted. Re-measured: cold &lt;code&gt;Duration&lt;/code&gt; p50 &lt;strong&gt;1,142ms&lt;/strong&gt; — statistically the same as the 1,124ms baseline. Priming had shipped and, as far as the numbers were concerned, changed nothing.&lt;/p&gt;

&lt;p&gt;Worse: we couldn't even tell &lt;em&gt;why&lt;/em&gt;. Every diagnostic line we'd added — "hooks registered," "SnapStart prime," even FastAPI's own unconditional startup log — was completely absent from CloudWatch across 90 minutes of cold starts. The working theory for a while was that SnapStart's init-phase logs simply don't reliably surface in CloudWatch, which would have made this nearly undebuggable from logs alone.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrong turn #2 (well, half a turn): it was a log level
&lt;/h2&gt;

&lt;p&gt;It wasn't a SnapStart logging quirk. It was &lt;code&gt;logging.info()&lt;/code&gt;. The Lambda's effective log level was dropping &lt;code&gt;INFO&lt;/code&gt; — CloudWatch showed our &lt;code&gt;[WARNING]&lt;/code&gt; lines and nothing below them, meaning every one of our priming diagnostics had been silently filtered the whole time. We moved the key diagnostic and priming log lines to &lt;code&gt;WARNING&lt;/code&gt; (PR #648) and immediately got a real signal for the first time in this investigation. Lesson, underlined: when a fix "does nothing and also produces no logs," check the log level before you start reasoning about distributed systems semantics. We spent longer on the second hypothesis than the first.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual bug: &lt;code&gt;after_restore&lt;/code&gt; cleared, but never rebuilt
&lt;/h2&gt;

&lt;p&gt;With real diagnostics finally visible, we split the cold &lt;code&gt;REPORT&lt;/code&gt; lines into two groups: requests that hit a &lt;em&gt;fresh init&lt;/em&gt; (rare — a brand-new execution environment) versus requests that hit a &lt;em&gt;restored&lt;/em&gt; snapshot (the common case SnapStart exists for). The restored-and-then-requested group was still averaging &lt;strong&gt;~1,064ms&lt;/strong&gt; — basically the original cliff, just hiding in a bucket we hadn't isolated before.&lt;/p&gt;

&lt;p&gt;The cause was in our own restore hook. A boto3 session captured in a snapshot carries frozen credentials and, once a client exists, frozen TLS connection pools. After a real restore, the execution environment is new: credentials need refreshing, any frozen socket is dead. So &lt;code&gt;reset_for_restore()&lt;/code&gt; correctly &lt;em&gt;cleared&lt;/em&gt; the cached session and clients on &lt;code&gt;after_restore&lt;/code&gt; — that part was right, and necessary for correctness. What it didn't do was rebuild them. So the next request found an empty session, built one lazily, and paid almost exactly the cost priming was supposed to eliminate — just relocated from "first request ever" to "first request after every restore," which for a Lambda that scales to zero between sparse dashboard loads is most of them.&lt;/p&gt;

&lt;p&gt;The fix (PR #649): &lt;code&gt;reset_for_restore()&lt;/code&gt; clears &lt;em&gt;and eagerly rebuilds&lt;/em&gt; — with the freshly-restored credentials — so the next request finds the session already built, not merely uncorrupted.&lt;/p&gt;

&lt;h2&gt;
  
  
  The result
&lt;/h2&gt;

&lt;p&gt;Post-restore first-request &lt;code&gt;Duration&lt;/code&gt;, n=7 samples per version, staging:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Version&lt;/th&gt;
&lt;th&gt;mean&lt;/th&gt;
&lt;th&gt;p50&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Baseline (no priming)&lt;/td&gt;
&lt;td&gt;1,183ms&lt;/td&gt;
&lt;td&gt;1,124ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clear-only (the bug)&lt;/td&gt;
&lt;td&gt;1,064ms&lt;/td&gt;
&lt;td&gt;1,060ms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Eager re-prime (the fix)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;652ms&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;490ms&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;p50 1,124 → 490ms, a 56% reduction.&lt;/strong&gt; Diagnostics on every restored environment now confirm the full chain worked as intended: &lt;code&gt;hooks_registered=True&lt;/code&gt;, &lt;code&gt;session_already_built=True&lt;/code&gt;, &lt;code&gt;restore_reprimed=True&lt;/code&gt;, zero errors.&lt;/p&gt;

&lt;p&gt;It's a real win, and it's not a complete one. About 3 of 7 restores in that sample still land at ≥500ms, even with the session confirmed pre-built — most likely the first actual DynamoDB &lt;em&gt;operation&lt;/em&gt; opening its own TLS connection, since boto3 connects lazily on first call rather than at client construction. Chasing that down would mean priming a real DynamoDB round-trip inside the restore hook itself, for what's probably diminishing returns against a ~490ms floor that's already good enough. We banked the 56% rather than shipping a seventh iteration.&lt;/p&gt;

&lt;p&gt;Alongside this, a frontend session-storage stale-while-revalidate cache (shipped separately, PR #639) hides both the network round-trip and any residual cold cliff on repeat visits — the dashboard's daily-batched data makes brief staleness safe — and a keep-warm EventBridge ping keeps instances hot enough that cold restores stay rare in the first place.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual lesson
&lt;/h2&gt;

&lt;p&gt;None of the three wrong turns here were exotic. "The fix looks right but does nothing" was a log level. "The clear-only version regressed" was a hook that did half its job and returned success anyway, because clearing genuinely is correct and necessary — it just isn't sufficient. The instrumentation that finally cracked it wasn't clever; it was making the diagnostic state (&lt;code&gt;hooks_registered&lt;/code&gt;, &lt;code&gt;session_already_built&lt;/code&gt;, &lt;code&gt;restore_reprimed&lt;/code&gt;) observable on the request path instead of trusting init-phase logs that, it turned out, we weren't even looking at correctly.&lt;/p&gt;

&lt;p&gt;If you're chasing a Lambda cold start and the fix you shipped isn't moving the number: check what log level is actually filtering your diagnostics before you start doubting your architecture.&lt;/p&gt;

&lt;p&gt;This is the same instinct behind the product: measure what's actually happening in your AWS account before acting on it. If you want to see what that looks like pointed at your own bill, &lt;a href="https://cloudcostwise.io?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=we-cut-lambda-cold-starts-56-percent-snapstart-lessons" rel="noopener noreferrer"&gt;a free, read-only scan&lt;/a&gt; takes about five minutes and changes nothing without you approving it first.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>lambda</category>
      <category>snapstart</category>
      <category>performance</category>
    </item>
    <item>
      <title>Why Read-Only Is the Only Safe Way to Let AI Near Your AWS Account</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 12 Aug 2026 13:54:04 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/why-read-only-is-the-only-safe-way-to-let-ai-near-your-aws-account-jj5</link>
      <guid>https://dev.to/cloudwiseteam/why-read-only-is-the-only-safe-way-to-let-ai-near-your-aws-account-jj5</guid>
      <description>&lt;p&gt;Every AWS cost tool eventually asks you for the same terrifying thing: an IAM role. And every vendor says the same reassuring word about it: "read-only." I want to tell you exactly what that word means when CloudWise says it, because "read-only" gets used loosely enough in this industry that the word alone shouldn't be enough to trust anyone — including us. So this post is the IAM, not the marketing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The role you actually grant
&lt;/h2&gt;

&lt;p&gt;When you connect an AWS account through CloudWise's one-click setup, you launch exactly one CloudFormation stack: the monitoring stack. There is no button in onboarding that launches anything else — &lt;code&gt;startConnect('monitoring')&lt;/code&gt; is the only path the connect flow calls.&lt;/p&gt;

&lt;p&gt;That stack's policy (&lt;code&gt;cloudwise-cur-setup-template.yaml&lt;/code&gt;) is a bespoke allow-list we wrote and maintain, not the AWS-managed &lt;code&gt;ReadOnlyAccess&lt;/code&gt; policy. That distinction matters more than it sounds like it should: &lt;code&gt;ReadOnlyAccess&lt;/code&gt; is enormous and vague — it grants read access to almost every AWS service, including ones CloudWise has no reason to ever look at. Our policy is scoped to what a cost scan actually needs, action by action: &lt;code&gt;ec2:DescribeInstances&lt;/code&gt;, &lt;code&gt;rds:DescribeDBInstances&lt;/code&gt;, &lt;code&gt;s3:ListAllMyBuckets&lt;/code&gt;, &lt;code&gt;ce:GetCostAndUsage&lt;/code&gt;, &lt;code&gt;compute-optimizer:GetEC2InstanceRecommendations&lt;/code&gt;, and so on — read the whole thing at &lt;a href="https://cloudcostwise.io/security/permissions?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=read-only-is-the-only-safe-way-to-let-ai-near-your-aws-account" rel="noopener noreferrer"&gt;cloudcostwise.io/security/permissions&lt;/a&gt;, which renders the live action count straight off the template, not a number we typed into a page and forgot to update. Every single statement in that policy is a &lt;code&gt;Get*&lt;/code&gt;, &lt;code&gt;Describe*&lt;/code&gt;, &lt;code&gt;List*&lt;/code&gt;, or &lt;code&gt;BatchGet*&lt;/code&gt; call — plus &lt;code&gt;sts:GetCallerIdentity&lt;/code&gt; and &lt;code&gt;iam:SimulatePrincipalPolicy&lt;/code&gt;, which CloudWise uses to check what a role &lt;em&gt;can&lt;/em&gt; do without ever calling it. There is no &lt;code&gt;Put&lt;/code&gt;, &lt;code&gt;Create&lt;/code&gt;, &lt;code&gt;Delete&lt;/code&gt;, &lt;code&gt;Update&lt;/code&gt;, &lt;code&gt;Attach&lt;/code&gt;, or &lt;code&gt;Modify&lt;/code&gt; verb anywhere in that role's policy. It cannot make a single write call against your account. Not "shouldn't" — cannot; IAM will reject the attempt at the API layer before it reaches any resource.&lt;/p&gt;

&lt;p&gt;(One nuance, for the pedants, because I'd rather you catch it than an auditor: the CUR bucket in that same template does have a bucket policy permitting &lt;code&gt;s3:PutObject&lt;/code&gt; — granted to AWS's own &lt;code&gt;billingreports.amazonaws.com&lt;/code&gt; service principal, not to CloudWise. That's how AWS itself delivers your Cost and Usage Report into a bucket you own. It's AWS writing to you, not us writing to anything.)&lt;/p&gt;

&lt;h2&gt;
  
  
  Remediation is a different stack, a different role, and a different click
&lt;/h2&gt;

&lt;p&gt;If you want CloudWise to actually &lt;em&gt;fix&lt;/em&gt; waste — stop an idle NAT gateway, delete an orphaned snapshot, right-size an instance — that is never part of onboarding. It requires deploying a second, separate CloudFormation stack (&lt;code&gt;CloudWise-Remediation&lt;/code&gt;, using &lt;code&gt;cloudwise-remediation-role.yaml&lt;/code&gt;) that you reach only from &lt;code&gt;settings/remediation&lt;/code&gt; or &lt;code&gt;setup/permissions&lt;/code&gt;, deliberately after the point where you've already seen what a read-only scan finds. Nothing in the sign-up or connect flow grants this role. You have to go looking for it.&lt;/p&gt;

&lt;p&gt;That role's policy carries an explicit &lt;strong&gt;Deny&lt;/strong&gt; block that overrides everything else, no matter what any other statement in the policy says:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;iam:*&lt;/code&gt;, &lt;code&gt;organizations:*&lt;/code&gt;, &lt;code&gt;sts:*&lt;/code&gt; — CloudWise can never touch identity, org structure, or assume other roles&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;secretsmanager:GetSecretValue&lt;/code&gt; / &lt;code&gt;PutSecretValue&lt;/code&gt; / &lt;code&gt;CreateSecret&lt;/code&gt; / &lt;code&gt;UpdateSecret&lt;/code&gt; — no access to your secrets, ever&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;cloudtrail:DeleteTrail&lt;/code&gt;, &lt;code&gt;cloudtrail:StopLogging&lt;/code&gt;, &lt;code&gt;config:DeleteConfigRule&lt;/code&gt;, &lt;code&gt;config:StopConfigurationRecorder&lt;/code&gt;, &lt;code&gt;guardduty:DeleteDetector&lt;/code&gt; — your audit and detection trail can't be turned off&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;s3:DeleteBucket&lt;/code&gt;, &lt;code&gt;ec2:DeleteVpc&lt;/code&gt;, &lt;code&gt;ec2:DeleteSubnet&lt;/code&gt;, &lt;code&gt;ec2:DeleteSecurityGroup&lt;/code&gt;, &lt;code&gt;rds:DeleteDBInstance&lt;/code&gt; — no deleting the structural stuff that would actually hurt&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ec2:AuthorizeSecurityGroupIngress/Egress&lt;/code&gt;, &lt;code&gt;ec2:RevokeSecurityGroupIngress/Egress&lt;/code&gt;, &lt;code&gt;ec2:CreateSecurityGroup&lt;/code&gt; — CloudWise cannot open, close, or create a network path&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;kms:CreateKey/CreateGrant/Encrypt/Decrypt/GenerateDataKey*&lt;/code&gt;, &lt;code&gt;ssm:*&lt;/code&gt; — no touching encryption or Systems Manager&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A Deny statement in IAM always wins, regardless of what any Allow statement elsewhere in the same policy — or any other policy — grants. That's not a UI promise. It's how the policy evaluates at the API layer.&lt;/p&gt;

&lt;p&gt;You'll also see a small number of &lt;code&gt;Create*&lt;/code&gt; actions in that same policy file — things like &lt;code&gt;ec2:RunInstances&lt;/code&gt; or &lt;code&gt;rds:CreateDBInstance&lt;/code&gt;. Those exist for one purpose: &lt;strong&gt;rollback&lt;/strong&gt;. Before executing an approved action, CloudWise records enough state to reverse it — restart an instance it stopped, recreate a resource it deleted, restore a secret's scheduled deletion. If a fix goes wrong, or you change your mind, there's a way back. That's a narrower and more honest thing than "can create resources," and it's also narrower than "can never create resources" — so I'm not going to round it either direction. Read the actual file if you want the specifics; that's the point of publishing it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Nothing executes without your click
&lt;/h2&gt;

&lt;p&gt;Here's the part I actually care about you believing: CloudWise proposes fixes. It does not execute them on its own initiative, ever, under any circumstance we ship today.&lt;/p&gt;

&lt;p&gt;The execution path (&lt;code&gt;lambdas/remediation_executor/handler.py&lt;/code&gt;) checks a status field before it will run a single mutating API call, and that status has to read &lt;code&gt;approved&lt;/code&gt;. The only thing that writes &lt;code&gt;approved&lt;/code&gt; is a separate approval-gateway Lambda, triggered by a verified, authenticated action from you. There is no code path today that sets an action to &lt;code&gt;approved&lt;/code&gt; automatically. If you never click approve, the action sits there, proposed, forever, and nothing happens to your account.&lt;/p&gt;

&lt;p&gt;There's also a second, quieter layer under that: every mutating call CloudWise's execution role makes is tagged with a session tag — &lt;code&gt;aws:PrincipalTag/cloudwise-action&lt;/code&gt; — that CloudWise itself sets at execution time, and the role's policy can further condition on that tag. I want to be precise about what this is and isn't: it's real defense-in-depth inside our own execution path, not a lever you hold. You don't set that tag; we do. The thing you actually control is upstream of it — the approve click that has to happen before any of this fires at all.&lt;/p&gt;

&lt;h2&gt;
  
  
  What "read-only" means when we say it
&lt;/h2&gt;

&lt;p&gt;So, precisely:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Onboarding grants exactly one role&lt;/strong&gt;, scoped to &lt;code&gt;Get&lt;/code&gt;/&lt;code&gt;Describe&lt;/code&gt;/&lt;code&gt;List&lt;/code&gt;/&lt;code&gt;BatchGet&lt;/code&gt; actions, verifiable line-by-line at &lt;a href="https://cloudcostwise.io/security/permissions?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=read-only-is-the-only-safe-way-to-let-ai-near-your-aws-account" rel="noopener noreferrer"&gt;cloudcostwise.io/security/permissions&lt;/a&gt;. It cannot write to your account. This is the role every CloudWise customer has, whether or not they ever touch remediation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Remediation is opt-in, separate, and denies the dangerous stuff outright&lt;/strong&gt; — identity, org structure, secrets, audit trails, network rules, deletion of anything structural — regardless of what else the policy grants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nothing mutates without your explicit approval.&lt;/strong&gt; Propose, then execute only what you approve. Not "propose, then execute unless you object."&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of that requires you to trust our intentions. It requires you to trust IAM evaluation semantics, which is a much smaller ask, and one you can verify yourself against files we publish rather than a page of prose we wrote about ourselves.&lt;/p&gt;

&lt;p&gt;If you want to see what the read-only role actually finds in your account, &lt;a href="https://cloudcostwise.io?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=read-only-is-the-only-safe-way-to-let-ai-near-your-aws-account" rel="noopener noreferrer"&gt;a free, read-only scan&lt;/a&gt; takes about five minutes, grants nothing beyond what's described above, and changes nothing without you approving it first.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>security</category>
      <category>iam</category>
      <category>ai</category>
    </item>
    <item>
      <title>NAT Gateway: The Bill Nobody Reads</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Mon, 20 Jul 2026 16:03:01 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/nat-gateway-the-bill-nobody-reads-4pc2</link>
      <guid>https://dev.to/cloudwiseteam/nat-gateway-the-bill-nobody-reads-4pc2</guid>
      <description>&lt;p&gt;Group your AWS bill by service in Cost Explorer and NAT Gateway charges don't get their own row. They're folded into &lt;strong&gt;Amazon Virtual Private Cloud&lt;/strong&gt;, sitting next to VPN connections, Transit Gateway attachments, and PrivateLink endpoints. Unless you break the view down by usage type — &lt;code&gt;NatGateway-Hours&lt;/code&gt;, &lt;code&gt;NatGateway-Bytes&lt;/code&gt; — the number you're actually paying for a gateway that might be doing nothing is invisible.&lt;/p&gt;

&lt;p&gt;We &lt;a href="https://dev.to/blog/aws-nat-gateway-costs"&gt;wrote up the mechanics of that bill back in March&lt;/a&gt; — the $0.045/hour base charge, the $0.045/GB data processing fee, the multi-AZ multiplication that turns one gateway into three. That post is still accurate and worth reading if you want the full pricing breakdown. This one is about something different: what CloudWise's detector actually checks before it tells you a NAT Gateway is dead weight, and where that check still falls short.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the detector actually looks at
&lt;/h2&gt;

&lt;p&gt;The idle-NAT-gateway check lives in &lt;code&gt;cloudwise_scan_core&lt;/code&gt;'s network detector, alongside the unattached-EIP and idle-load-balancer checks — it's one pass over a VPC's network resources, not a dedicated NAT scanner. For every NAT Gateway in &lt;code&gt;available&lt;/code&gt; state, it pulls two CloudWatch metrics over a 7-day window: &lt;code&gt;ActiveConnectionCount&lt;/code&gt; and &lt;code&gt;BytesOutToDestination&lt;/code&gt;. If both are zero for the full week, it flags the gateway as &lt;code&gt;IDLE_NAT_GATEWAY&lt;/code&gt; at &lt;code&gt;HIGH&lt;/code&gt; confidence.&lt;/p&gt;

&lt;p&gt;Two design choices worth calling out:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why 7 days, not 1.&lt;/strong&gt; A single quiet day doesn't mean a gateway is unused — it might serve a batch job that runs Sundays, or a staging environment nobody touches on weekends. A full week with zero connections and zero bytes is a much harder signal to explain away.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why the estimate is conservative.&lt;/strong&gt; The finding's &lt;code&gt;monthly_savings&lt;/code&gt; uses a flat $32.40 — the hourly base charge times a 720-hour month, not the $32.85 you'd get from AWS's actual 730-hour average. It also doesn't add anything for data processing, because by definition a gateway that's flagged idle processed zero bytes in the lookback window. The number CloudWise shows you is the floor, not an estimate padded to look impressive.&lt;/p&gt;

&lt;p&gt;The finding's &lt;code&gt;risk&lt;/code&gt; field is blunt about the tradeoff, too: deleting a NAT Gateway immediately cuts internet egress for every private-subnet resource routed through it. The action is offered — &lt;code&gt;aws ec2 delete-nat-gateway&lt;/code&gt; — but nothing executes it without a human approving first. That's the same read-only-first posture behind every detector we ship, not something special-cased for NAT Gateways.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it doesn't do yet
&lt;/h2&gt;

&lt;p&gt;Here's the honest gap: the detector tells you a NAT Gateway earned its keep for zero dollars this week. It doesn't tell you what to replace it with.&lt;/p&gt;

&lt;p&gt;That decision genuinely depends on what's routing through it:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;If the only traffic is S3 or DynamoDB, a &lt;strong&gt;Gateway VPC Endpoint&lt;/strong&gt; replaces it for free — no hourly charge, no per-GB fee, ever.&lt;/li&gt;
&lt;li&gt;If it's other AWS services (Secrets Manager, ECR, CloudWatch Logs), an &lt;strong&gt;Interface VPC Endpoint&lt;/strong&gt; runs $0.01/hour per AZ plus $0.01/GB — about 78% cheaper per gigabyte than NAT, though you're paying a small hourly fee per endpoint per AZ instead of one gateway.&lt;/li&gt;
&lt;li&gt;If it's a non-production environment with real internet egress needs but low traffic and no requirement for managed HA, a self-managed &lt;strong&gt;NAT instance&lt;/strong&gt; (something like a &lt;code&gt;t4g.nano&lt;/code&gt;) can run for a few dollars a month — you trade the 24/7 base charge for patching it yourself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;CloudWise doesn't make that call for you today. It's a fair next detector to build — matching an idle gateway's actual destination traffic against what a Gateway Endpoint could cover for free — but until it exists, the decision after "this is idle" is still on you. We'd rather say that plainly than imply the tool does more than it does.&lt;/p&gt;

&lt;h2&gt;
  
  
  Finding yours
&lt;/h2&gt;

&lt;p&gt;If you want to see this in your own account without reading CloudWatch dashboards by hand, &lt;a href="https://cloudcostwise.io?utm_source=blog&amp;amp;utm_medium=hub&amp;amp;utm_campaign=nat-gateway-bill-nobody-reads" rel="noopener noreferrer"&gt;a free scan&lt;/a&gt; checks this along with 190+ other waste patterns across 40+ AWS services — read-only, five minutes, nothing gets deleted without you clicking it. We &lt;a href="https://youtube.com/shorts/z7bP_wMccek?feature=share" rel="noopener noreferrer"&gt;posted a 30-second walkthrough of this exact detector&lt;/a&gt; if you'd rather watch than read.&lt;/p&gt;

&lt;p&gt;Or just run the two-metric check yourself:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aws cloudwatch get-metric-statistics &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--namespace&lt;/span&gt; AWS/NATGateway &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--metric-name&lt;/span&gt; BytesOutToDestination &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--dimensions&lt;/span&gt; &lt;span class="nv"&gt;Name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;NatGatewayId,Value&lt;span class="o"&gt;=&lt;/span&gt;nat-0abc123def456 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--start-time&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; &lt;span class="nt"&gt;-v-7d&lt;/span&gt; +%Y-%m-%dT%H:%M:%S&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--end-time&lt;/span&gt; &lt;span class="si"&gt;$(&lt;/span&gt;&lt;span class="nb"&gt;date&lt;/span&gt; &lt;span class="nt"&gt;-u&lt;/span&gt; +%Y-%m-%dT%H:%M:%S&lt;span class="si"&gt;)&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--period&lt;/span&gt; 86400 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--statistics&lt;/span&gt; Sum
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Seven zeros in a row is $32.40 a month for a gateway that isn't gating anything.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
    <item>
      <title>Why I Built a Risk Score Instead of a Buy Button</title>
      <dc:creator>Rick Wise</dc:creator>
      <pubDate>Wed, 08 Jul 2026 17:26:13 +0000</pubDate>
      <link>https://dev.to/cloudwiseteam/why-i-built-a-risk-score-instead-of-a-buy-button-54eg</link>
      <guid>https://dev.to/cloudwiseteam/why-i-built-a-risk-score-instead-of-a-buy-button-54eg</guid>
      <description>&lt;p&gt;The obvious feature to build here is a button. "You're spending $4,200/month on &lt;code&gt;m5.xlarge&lt;/code&gt; — buy a 1-year Savings Plan and save 30%." One click, instant discount, everybody's happy.&lt;/p&gt;

&lt;p&gt;I almost built that button. I'm glad I didn't.&lt;/p&gt;

&lt;p&gt;Here's the problem with the button: it's only looking at the last three months. It has no idea whether the workload it's telling you to commit to will still exist in month nine. And a Reserved Instance or Savings Plan isn't a coupon — it's a bet, paid up front or amortized monthly, that a specific slice of your infrastructure will look roughly the same for a year or three. Get that bet wrong and the "savings" tool just talked you into a liability.&lt;/p&gt;

&lt;p&gt;So instead of a buy button, I built a &lt;strong&gt;Commitment Risk Score&lt;/strong&gt; — a feature whose entire job is to sometimes tell you &lt;em&gt;not&lt;/em&gt; to buy anything yet.&lt;/p&gt;

&lt;h2&gt;
  
  
  The feature that argues with itself
&lt;/h2&gt;

&lt;p&gt;The Commitment Risk Score pulls four signals straight from Cost Explorer for an account: how much your top instance families churn month to month, how volatile your spend is, how long your individual resources actually live, and how well you're already using the commitments you have. It weights them (35/25/25/15) into a single 0–100 score, and that score maps to a recommendation — anywhere from "3-year Convertible RI, take the discount" down to "on-demand and Spot only, buy nothing."&lt;/p&gt;

&lt;p&gt;That last outcome is the part that made this feature interesting to build. Most cost-optimization tools are graded on how much they tell you to save. This one is graded on how honest it is about when a "savings" purchase would actually be a mistake. It costs about $0.04/account/month to compute — four &lt;code&gt;GetCostAndUsage&lt;/code&gt; calls, refreshed weekly — the recommendation quality has nothing to do with the compute cost, so there was no excuse to cut a signal to save pennies.&lt;/p&gt;

&lt;p&gt;I wrote up the full math — the Jaccard-distance churn calculation, the coefficient-of-variation volatility threshold, the worked examples — in a separate deep dive, because the mechanics deserved their own space: &lt;a href="https://dev.to/blog/aws-reserved-instance-commitment-risk"&gt;The Commitment-Risk Score: should you buy that RI?&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What almost went wrong
&lt;/h2&gt;

&lt;p&gt;The instinct that almost got me was averaging. My first pass weighted all four signals close to evenly, because "why not, they're all relevant." It took building the worked examples to see the problem: a team mid-Graviton-migration with high churn but a stable dollar total would score &lt;em&gt;fine&lt;/em&gt; on an even-weighted average, because volatility and churn partially cancel out in the wrong direction. The churn signal needed to dominate — 35%, not 25% — because a changing instance mix is structurally the most common way a commitment gets stranded, regardless of what the topline spend number is doing.&lt;/p&gt;

&lt;p&gt;The existing-waste signal ended up smallest at 15%, which felt backwards at first — isn't "you're already wasting money" the most damning fact? It is, but it's also the least &lt;em&gt;predictive&lt;/em&gt; one. It tells you about the past, not whether the next commitment will strand. It stayed in the score because it's a legitimate red flag, but it doesn't get to drown out the forward-looking signals.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is the right shape for a cost tool
&lt;/h2&gt;

&lt;p&gt;CloudWise's whole premise is read-only: we look at your AWS usage and billing data, we never touch your infrastructure, and every dollar figure we show you needs to survive you checking it against Cost Explorer yourself. A recommendation engine that only ever says "buy more" doesn't survive that scrutiny for long — eventually it recommends a commitment that strands, and you stop trusting the number.&lt;/p&gt;

&lt;p&gt;A feature that's willing to say "not yet, here's why" is the one worth trusting the next time it says "yes, buy it." That was true before I shipped this, and building it just made me believe it more.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;CloudWise is an AWS cost optimization tool for startups — 191 automated waste checks including commitment-risk scoring, read-only by design, starting at $19/month. Run a free scan at &lt;a href="https://cloudcostwise.io" rel="noopener noreferrer"&gt;cloudcostwise.io&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>aws</category>
      <category>finops</category>
      <category>devops</category>
      <category>cloud</category>
    </item>
  </channel>
</rss>
