<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yuuki Yamashita</title>
    <description>The latest articles on DEV Community by Yuuki Yamashita (@_76130e67067eab4c8510).</description>
    <link>https://dev.to/_76130e67067eab4c8510</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3963934%2Ff567e490-409e-4254-8600-f596ed5e7e99.png</url>
      <title>DEV Community: Yuuki Yamashita</title>
      <link>https://dev.to/_76130e67067eab4c8510</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_76130e67067eab4c8510"/>
    <language>en</language>
    <item>
      <title>I Asked Claude Code Why AI Labs Suddenly Want to Slow AI Down</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Mon, 14 Sep 2026 01:38:18 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-asked-claude-code-why-ai-labs-suddenly-want-to-slow-ai-down-gmg</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-asked-claude-code-why-ai-labs-suddenly-want-to-slow-ai-down-gmg</guid>
      <description>&lt;h2&gt;
  
  
  The trigger
&lt;/h2&gt;

&lt;p&gt;On September 12, 2026, three competitors, Anthropic, OpenAI, and xAI, all said roughly the same thing within hours of each other: AI development needs to slow down. Companies that fight over the same enterprise contracts and the same researchers do not usually agree on anything in public. So I opened Claude Code and asked it to walk through what was actually going on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round one: the nuclear weapons comparison
&lt;/h2&gt;

&lt;p&gt;My first question was whether this looked like the nuclear non-proliferation playbook: the people who built the weapon are the ones who later argue loudest for controlling it.&lt;/p&gt;

&lt;p&gt;Claude Code's answer split the comparison into what holds and what does not. What holds: the leading developer warning about the danger of what it built, and a first mover shaping the resulting regulation in a way that locks in its own position, is a real structural echo of the NPT era. What breaks down: nuclear proliferation has a physical bottleneck, uranium enrichment, that can be detected from orbit. AI capability spreads at the speed of compute and data, with no equivalent tripwire. Nuclear weapons are held specifically to never be used; frontier AI models are built specifically to be used, constantly, for profit. The incentive structures point in opposite directions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round two: fact-checking the essay itself
&lt;/h2&gt;

&lt;p&gt;Rather than take the "slow down" framing at face value, I had Claude Code pull the actual source: Dario Amodei's essay "We Must Pace the Frontier." Worth doing, because most coverage ("AI CEOs want to slow AI down") undersells what he is actually proposing. He is explicit that pacing "does not mean halting model training." The real ask is third-party testing across four risk categories, permanent embedded evaluators inside Anthropic, and an antitrust carve-out so labs can coordinate on safety without violating competition law.&lt;/p&gt;

&lt;p&gt;That last item is the one critics keyed in on. Chamath Palihapitiya's response on X: "Dario makes the case to stop open source and concentrate enormous technological and economic power with Anthropic." Whether or not you buy that read, it is a materially different claim than "Anthropic wants AI to be safer," and it is worth knowing both versions exist before writing a hot take.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round three: the detail I almost missed
&lt;/h2&gt;

&lt;p&gt;I asked Claude Code to check whether anything else happened around that week that might be relevant. It surfaced a threat intelligence report Anthropic had published two days earlier, on September 10: seven Chinese AI labs, including Moonshot AI (maker of the Kimi models), accused of large-scale distillation of Claude, GPT, Gemini, and Grok. The specific claim against Moonshot: roughly 300,000 requests routed to Claude over ten days through a network of over 5,000 fraudulent accounts, with Claude's answers served back to users labeled as Kimi's own. US intelligence agencies backed the concern.&lt;/p&gt;

&lt;p&gt;Two days between that report and the "slow the frontier" essay is a short gap. Claude Code was careful not to claim a confirmed causal link, correctly, since none of the primary sources draw one. But it flagged the sequence as worth including, because both stories run on the same anxiety: capability spreading somewhere it should not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round four: the rumor that did not survive fact-checking
&lt;/h2&gt;

&lt;p&gt;The same week, a rumor was circulating that Huawei's founder and his family had fled China. I asked Claude Code to fold that into the piece too. It refused to treat it as established, and pushed back before writing anything: the only source was a single screenshot posted to Chinese social media, the original poster had called it unverified, Huawei and Chinese authorities had not commented, and Chinese social media reaction skewed skeptical. It is in this piece as an unconfirmed rumor that circulated the same week, not as evidence of anything.&lt;/p&gt;

&lt;p&gt;It would have been easy to let an interesting-sounding rumor slide into a paragraph about geopolitical pressure without that check. Separating what is verified from what is circulating on social media, before it goes into a draft, is the step that gets skipped most often when writing under deadline.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgi29dk8bnf1x85bommb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgi29dk8bnf1x85bommb.jpg" alt=" " width="799" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;No single motive explains three competitors agreeing in public. Some of the fear about near-term AGI timelines is probably genuine. The specific regulatory ask looks a lot like regulatory capture. And the timing next to the distillation report suggests a geopolitical containment story running underneath the safety framing. All three can be true at once, which is a less satisfying headline than any one of them alone, but it is the one the primary sources actually support.&lt;/p&gt;

&lt;p&gt;Sources: &lt;a href="https://darioamodei.com/post/we-must-pace-the-frontier" rel="noopener noreferrer"&gt;Dario Amodei, "We Must Pace the Frontier"&lt;/a&gt;, &lt;a href="https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing" rel="noopener noreferrer"&gt;Axios&lt;/a&gt;, &lt;a href="https://x.com/chamath/status/2098780471966802037" rel="noopener noreferrer"&gt;Chamath Palihapitiya on X&lt;/a&gt;, &lt;a href="https://www.coindesk.com/tech/2026/09/09/moonshot-s-kimi-rattled-markets-u-s-agencies-now-say-it-was-trained-on-american-models" rel="noopener noreferrer"&gt;CoinDesk on the Moonshot/Kimi distillation report&lt;/a&gt;. The Huawei rumor traces to a single unverified screenshot with no mainstream corroboration; see &lt;a href="https://hidamaricolumn.com/huawei-ren-escape-rumor/" rel="noopener noreferrer"&gt;this Japanese-language writeup&lt;/a&gt; for the closest thing to a source trail.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>anthropic</category>
      <category>claudecode</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Tested Whether cdkd Really Deploys Faster Than cdk deploy</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 03 Sep 2026 15:58:25 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-tested-whether-cdkd-really-deploys-faster-than-cdk-deploy-25i4</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-tested-whether-cdkd-really-deploys-faster-than-cdk-deploy-25i4</guid>
      <description>&lt;p&gt;A tool claiming "up to 15x faster than cdk deploy" showed up in my feed a while back. Drop-in replacement, it said: keep your CDK app exactly as it is, just swap &lt;code&gt;cdk deploy&lt;/code&gt; for &lt;code&gt;cdkd deploy&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I've learned to be skeptical of "Nx faster" claims. So I actually deployed something real to AWS with both tools and timed it. Short version: it really is that fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  What cdkd actually is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/go-to-k/cdkd" rel="noopener noreferrer"&gt;cdkd&lt;/a&gt; deploys an existing AWS CDK app without going through CloudFormation. It calls the AWS SDK directly instead. It's built by &lt;a href="https://github.com/go-to-k" rel="noopener noreferrer"&gt;go-to-k&lt;/a&gt; (Kenta Goto), an AWS DevTools Hero and CDK top contributor who also maintains &lt;code&gt;cls3&lt;/code&gt; (a fast S3 bucket emptier) and &lt;code&gt;delstack&lt;/code&gt; (for cleaning up stuck CloudFormation/CDK stacks) — tools that quietly fix the annoying parts of working with AWS. cdkd feels like the biggest one yet, and I mean that as a compliment grounded in actually using it, not a throwaway one.&lt;/p&gt;

&lt;p&gt;The mechanism is straightforward. cdkd runs the exact same CDK synth step as the CDK CLI, producing the same CloudFormation template. What changes is everything after that: instead of handing the template to CloudFormation, cdkd's own engine reads the resource dependency graph (&lt;code&gt;Ref&lt;/code&gt;, &lt;code&gt;Fn::GetAtt&lt;/code&gt;), builds a DAG, and fires AWS SDK / Cloud Control API calls directly, in parallel, as soon as each resource's dependencies are satisfied.&lt;/p&gt;

&lt;p&gt;Worth saying up front: cdkd calls itself not production-ready, dev/test only. This isn't a "replace CloudFormation in prod" pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  I actually ran both, on real AWS
&lt;/h2&gt;

&lt;p&gt;cdkd's own README backs up the 15x number with a VPC + Lambda + SQS + CloudFront benchmark. So I wrote that same stack as a CDK app and deployed it twice — &lt;code&gt;DeployRaceCfn&lt;/code&gt; via &lt;code&gt;cdk deploy&lt;/code&gt;, &lt;code&gt;DeployRaceCdkd&lt;/code&gt; via &lt;code&gt;cdkd deploy&lt;/code&gt; — to the same AWS account, same region (ap-northeast-1).&lt;/p&gt;

&lt;p&gt;The stack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VPC (2 AZ + NAT Gateway) with a Lambda inside it, fronted by a Function URL&lt;/li&gt;
&lt;li&gt;CloudFront, origin set to that Function URL&lt;/li&gt;
&lt;li&gt;SQS + EventSourceMapping + a consumer Lambda&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;First attempt failed. The account had hit its VPC limit (five, the default):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resource handler returned message: "The maximum number of VPCs has been reached.
(Service: Ec2, Status Code: 400, ...)"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other test stacks in the same account were sitting on VPCs I'd forgotten about. I tore down the cdkd stack (already measured, no longer needed) to free a slot and reran. cdkd writes a structured event log to S3 on every run (&lt;code&gt;cdkd events&lt;/code&gt;), so even the failed attempt was easy to diagnose after the fact.&lt;/p&gt;

&lt;p&gt;Timing compares the deploy phase only. Synth is identical work either way (same &lt;code&gt;aws-cdk-lib&lt;/code&gt;), so I excluded it, matching how cdkd's own benchmarks are measured. The &lt;code&gt;cdk deploy&lt;/code&gt; timeline comes from CloudFormation's &lt;code&gt;DescribeStackEvents&lt;/code&gt;; the &lt;code&gt;cdkd deploy&lt;/code&gt; timeline comes from &lt;code&gt;cdkd events &amp;lt;stack&amp;gt; --run &amp;lt;id&amp;gt; --format json&lt;/code&gt;, reading the &lt;code&gt;RESOURCE_STARTED&lt;/code&gt; / &lt;code&gt;RESOURCE_SUCCEEDED&lt;/code&gt; events it records itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;cdk deploy&lt;/th&gt;
&lt;th&gt;cdkd deploy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Time&lt;/td&gt;
&lt;td&gt;479.3s&lt;/td&gt;
&lt;td&gt;95.0s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resources created&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.0x.&lt;/strong&gt; The one extra resource on the CloudFormation side is &lt;code&gt;AWS::CDK::Metadata&lt;/code&gt;, a bookkeeping resource that only exists there — both sides build the same 33 real resources.&lt;/p&gt;

&lt;p&gt;Watching &lt;code&gt;cdkd deploy&lt;/code&gt; run, IAM roles and route tables land in a burst right at the start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[1/33] ✓ RaceQueueE818AC65 (AWS::SQS::Queue) created
[2/33] ✓ RaceVpcIGW94C1C01D (AWS::EC2::InternetGateway) created
[3/33] ✓ RaceVpcPublicSubnet1EIPC3B3497E (AWS::EC2::EIP) created
[4/33] ✓ ConsumerFunctionServiceRole68E8FEB1 (AWS::IAM::Role) created
[5/33] ✓ MainFunctionServiceRole8C918DF0 (AWS::IAM::Role) created
...
CloudFront Distribution Distribution830FAC52 accepted (not waiting for Deployed; pass --full-wait to wait)
Deployment Summary:
  Created: 33 / Duration: 94.99s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meanwhile &lt;code&gt;cdk deploy&lt;/code&gt; takes 16.6 seconds just to get its first VPC. By that point cdkd already has the Lambda wired up to SQS.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the time actually goes
&lt;/h2&gt;

&lt;p&gt;cdk deploy hands the template to CloudFormation, which creates resources one at a time. cdkd reads the same template's dependency graph itself and calls the AWS SDK / Cloud Control API directly, in parallel, as soon as dependencies clear. Cutting out the CloudFormation middleman is most of the story, but two specific waits explain most of the 384-second gap.&lt;/p&gt;

&lt;p&gt;The NAT Gateway is the first one. cdk deploy works through the SQS/Lambda side of the graph before it gets around to waiting on the NAT Gateway to come up. cdkd hits that wait much earlier. Same AWS-side wait either way — what differs is how early in the run you eat it.&lt;/p&gt;

&lt;p&gt;CloudFront is the bigger one. CloudFormation's default behavior waits until the distribution reaches &lt;code&gt;Deployed&lt;/code&gt; — full global propagation, three-plus minutes. cdkd's default returns as soon as &lt;code&gt;CreateDistribution&lt;/code&gt; is accepted (there's a &lt;code&gt;--full-wait&lt;/code&gt; flag if you want CloudFormation's behavior instead). Of the 479.3 seconds cdk deploy took, 183 of them — over three minutes — are spent solely on that one wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  This might actually change how people deploy
&lt;/h2&gt;

&lt;p&gt;I'll admit it: the slowness of iterating on a real AWS resource is part of why people reach for Vercel or Amplify instead. Build a VPC + Lambda + CloudFront stack in CDK and every check-your-work loop costs minutes, sometimes double digits of them.&lt;/p&gt;

&lt;p&gt;cdkd doesn't replace CloudFormation's state management, drift detection, or rollback handling, and it says so itself. But "CloudFormation in prod, cdkd while iterating" is now a real option, not a hypothetical. A CI pipeline that rebuilds a PR environment on every push, or just the apply-then-check loop on your own machine, running five times faster is enough to make "AWS-native is slower to iterate on than Vercel or Amplify" stop being true for a chunk of use cases.&lt;/p&gt;

&lt;p&gt;I also built an interactive replay of the actual measured timeline, plus a narrated demo video, if you want to see the two runs side by side.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00agsh7vxmj5wpz8sdvz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00agsh7vxmj5wpz8sdvz.png" alt=" " width="800" height="382"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cpxd8adrvnftf04goeq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cpxd8adrvnftf04goeq.png" alt=" " width="800" height="320"&gt;&lt;/a&gt;&lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/F9_ewhE9zY8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;cdkd is &lt;a href="https://github.com/go-to-k/cdkd" rel="noopener noreferrer"&gt;go-to-k/cdkd&lt;/a&gt; (Apache-2.0). Worth a star if this was useful.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cdk</category>
      <category>cloudformation</category>
      <category>devops</category>
    </item>
    <item>
      <title>Detecting Shadow AI Agents with AWS Agent Registry</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 01 Sep 2026 02:56:07 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/detecting-shadow-ai-agents-with-aws-agent-registry-o4c</link>
      <guid>https://dev.to/_76130e67067eab4c8510/detecting-shadow-ai-agents-with-aws-agent-registry-o4c</guid>
      <description>&lt;p&gt;AWS Agent Registry (GA August 2026) gives a team a private catalog for AI agents, MCP servers, and tools — semantic search, approval workflows, CloudTrail audit trails. What it doesn't give you out of the box is a way to find the agents that were &lt;em&gt;never registered in the first place&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;This is a pattern for closing that gap: scan the AWS account for AgentCore runtimes, diff them against what's actually in the registry, and route anything unregistered through a real approval flow before it becomes a permanent record. I built it as &lt;strong&gt;Shadow Agent Hunter&lt;/strong&gt;; this post is about the Agent Registry API details that made it work (and the ones that didn't, at first).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpiivmq47ddts83oh45h.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpiivmq47ddts83oh45h.jpg" alt=" " width="800" height="508"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The three API calls that matter
&lt;/h2&gt;

&lt;p&gt;Agent Registry splits into a control plane (&lt;code&gt;agent-registry-control&lt;/code&gt;) for managing registries and records, and a data plane (&lt;code&gt;agent-registry&lt;/code&gt;) for searching approved ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding what's already registered.&lt;/strong&gt; &lt;code&gt;list_registry_records&lt;/code&gt; returns every record regardless of status, so pending/draft records don't get re-flagged on a rescan either:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;

&lt;span class="n"&gt;control&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent-registry-control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;registered_names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;kwargs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;registryId&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nextToken&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_registry_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;registered_names&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;registryRecords&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nextToken&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I originally reached for the &lt;code&gt;provenance&lt;/code&gt; field here, expecting to link records back to their source AgentCore runtime ARN. Don't — &lt;code&gt;create_registry_record&lt;/code&gt; rejects a caller-supplied &lt;code&gt;provenance&lt;/code&gt; with &lt;code&gt;ValidationException: provenance cannot be set by the caller&lt;/code&gt;. It's populated only by the service's own auto-detection/sync integrations, not by a plain API call. Matching on record name is simpler anyway, since names are unique within a registry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Searching for duplicates.&lt;/strong&gt; &lt;code&gt;search_discoverable_registry_records&lt;/code&gt; does hybrid semantic + keyword search, ordered by relevance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;registry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent-registry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;registry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search_discoverable_registry_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;searchQuery&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runtime_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runtime_description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;registryIds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;registry_arn&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;maxResults&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There's no numeric score in the response — just an ordered list. With a small registry (say, two or three approved records), that means it always returns &lt;em&gt;something&lt;/em&gt;, even for genuinely unrelated runtimes, because there's nothing better to rank against. This isn't a bug so much as a reminder that semantic search needs a reasonably sized corpus to actually discriminate. It's also a decent argument for keeping a human in the approval loop rather than auto-rejecting on any hit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writing the approval.&lt;/strong&gt; Three sequential calls: create, submit, approve.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_registry_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;runtime_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;recordType&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CUSTOM&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;recordVersion&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;descriptors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runtimeArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;runtime_arn&lt;/span&gt;&lt;span class="p"&gt;})}},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;record_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;recordArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/record/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# not returned directly
&lt;/span&gt;
&lt;span class="c1"&gt;# record starts CREATING and must reach DRAFT before you can submit it
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_registry_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DRAFT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;submit_registry_record_for_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update_registry_record_status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;APPROVED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statusReason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reviewed and approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things worth flagging: &lt;code&gt;CreateRegistryRecordResponse&lt;/code&gt; only returns &lt;code&gt;recordArn&lt;/code&gt; and &lt;code&gt;status&lt;/code&gt; — no &lt;code&gt;recordId&lt;/code&gt; field, so you extract it from the ARN — and creation is asynchronous, so a record submitted for approval before it leaves &lt;code&gt;CREATING&lt;/code&gt; will fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-referencing with AgentCore Runtime and CloudTrail
&lt;/h2&gt;

&lt;p&gt;The other half of "shadow agent" detection is enumerating what's actually running, independent of the registry:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;agentcore&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-agentcore-control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;runtimes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;agentcore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_agent_runtimes&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agentRuntimes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The IAM action for this is &lt;code&gt;bedrock-agentcore:ListAgentRuntimes&lt;/code&gt; — note the namespace is &lt;code&gt;bedrock-agentcore&lt;/code&gt;, not &lt;code&gt;bedrock-agentcore-control&lt;/code&gt; like the SDK package name would suggest. Getting this wrong produces a plain &lt;code&gt;AccessDeniedException&lt;/code&gt; with no hint about the namespace mismatch.&lt;/p&gt;

&lt;p&gt;For attribution — who deployed an unregistered runtime — CloudTrail's &lt;code&gt;LookupEvents&lt;/code&gt; over &lt;code&gt;CreateAgentRuntime&lt;/code&gt; gets you there, with the usual 90-day retention caveat:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;cloudtrail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cloudtrail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cloudtrail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lookup_events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;LookupAttributes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AttributeKey&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EventName&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AttributeValue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CreateAgentRuntime&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Events&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;detail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CloudTrailEvent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;arn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;responseElements&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agentRuntimeArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;deployer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;userIdentity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;arn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Infra notes
&lt;/h2&gt;

&lt;p&gt;Agent Registry has no CDK construct as of September 2026 — no &lt;code&gt;AWS::AgentRegistry::Registry&lt;/code&gt; CloudFormation resource type exists yet. I provisioned the registry and its seed records with the boto3 script above rather than CDK. AgentCore Runtime, by contrast, has a stable L2 construct (&lt;code&gt;aws_bedrockagentcore.Runtime&lt;/code&gt;), and &lt;code&gt;AgentRuntimeArtifact.fromCodeAsset()&lt;/code&gt; deploys straight from a local Python directory with no Docker step.&lt;/p&gt;

&lt;p&gt;If you're running this on Vercel with OIDC federation to AWS: the OIDC provider is scoped per Vercel &lt;em&gt;team&lt;/em&gt;, not per project. A second project under the same team hitting &lt;code&gt;new iam.OpenIdConnectProvider(...)&lt;/code&gt; will fail deployment, since an AWS account only accepts one provider per issuer URL. Import the existing one with &lt;code&gt;iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn()&lt;/code&gt; instead of creating a second.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/oO9YQVkyDjY" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Result
&lt;/h2&gt;

&lt;p&gt;Scanning an account with a couple of intentionally-similar and intentionally-unrelated AgentCore runtimes: the similar one surfaces a real "possible duplicate" match against an existing approved record, and approving the unrelated one produces a genuine &lt;code&gt;APPROVED&lt;/code&gt; record you can see with &lt;code&gt;list-registry-records&lt;/code&gt; afterward — not a mock, the actual governance workflow AWS Agent Registry ships.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>cloud</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>The Legal Wall I Hit Building a YouTube Clone on AWS</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Wed, 26 Aug 2026 16:31:29 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/the-legal-wall-i-hit-building-a-youtube-clone-on-aws-1ocf</link>
      <guid>https://dev.to/_76130e67067eab4c8510/the-legal-wall-i-hit-building-a-youtube-clone-on-aws-1ocf</guid>
      <description>&lt;p&gt;Right now I'm testing how far you can push a YouTube-style video platform built entirely on AWS. Upload, transcode, stream — EC2, S3, and CloudFront handle all of that without much drama. Then I started sketching out a comment section and a direct-message feature between users, and I stopped typing.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Wait. Does this need a government license now?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That one question sent me down a rabbit hole through Japan's Telecommunications Business Act, Copyright Act, and a law with an even longer name that regulates online platforms. None of this is legal advice — I'm not a lawyer, just a developer who got nervous enough to read the actual statutes. If you're building something for real users at scale, talk to an actual lawyer. But if you're a solo developer wondering whether your side project is quietly illegal, this is the walkthrough I wish existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Japan's Telecommunications Business Act: the "are you a phone company" test
&lt;/h2&gt;

&lt;p&gt;Japan has a law called the &lt;a href="https://www.japaneselawtranslation.go.jp/en/laws/view/3648/en" rel="noopener noreferrer"&gt;Telecommunications Business Act&lt;/a&gt; (電気通信事業法). In the simplest possible terms: it's the law that decides whether your app is legally acting like a phone company, and if so, whether you need to tell the government about it.&lt;/p&gt;

&lt;p&gt;There are two tracks. Article 9 requires full registration, and it only kicks in if you own and operate your own transmission lines — think an actual telecom carrier laying fiber. Article 16 requires a lighter-weight notification, and it applies to almost everyone else who runs a communications service without owning that physical infrastructure. Since your app sits on AWS rather than your own fiber network, Article 16 (notification) is the one that could apply to you, not Article 9.&lt;/p&gt;

&lt;p&gt;Whether it actually applies comes down to one test: are you mediating communication between other people? A one-way video stream — you upload, others watch — is your own communication going out to viewers. It's not you relaying messages between two other people, so it generally falls outside the notification requirement. Add a DM feature, a live chat relay, or real-time comment broadcasting between users, though, and you start looking a lot more like a company that carries other people's messages, which is exactly what the law is watching for.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you're the only one who can upload, you're basically fine
&lt;/h2&gt;

&lt;p&gt;Here's the scorecard for a video app where you're the only content creator — think a personal portfolio site that happens to look like YouTube.&lt;/p&gt;

&lt;p&gt;Telecommunications Business Act notification isn't required, because you're not relaying anyone else's communication. The Information Distribution Platform Act (more on that below) doesn't apply either, since there's no user-generated content for anyone to complain about. Japan's Act on the Protection of Personal Information technically applies the moment you add user accounts or store watch history, but for a small non-commercial project, a basic privacy policy covers you in practice.&lt;/p&gt;

&lt;p&gt;If this describes your project, you can build it, ship it, and move on — the same way you'd treat any other side project.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moment strangers can upload, the rulebook gets thicker
&lt;/h2&gt;

&lt;p&gt;Turn on user uploads for everyone, and you've built a real user-generated-content platform. That changes things.&lt;/p&gt;

&lt;p&gt;The Telecommunications Business Act question gets murkier once you add DMs or chat, but realistically, the enforcement risk for a small side project is close to zero — call it a legal gray zone rather than a hard stop.&lt;/p&gt;

&lt;p&gt;The law that actually matters here is Japan's Information Distribution Platform Act (情報流通プラットフォーム対処法), which took effect on April 1, 2025. It's the successor to what used to be called the Provider Liability Limitation Act, and at its core it requires any platform hosting user content to have a way for people to report defamatory or copyright-infringing material and get it taken down. If you're running a UGC service, you need this regardless of size — but in practice, for a solo project, that requirement is satisfiable with a single email address people can send takedown requests to.&lt;/p&gt;

&lt;p&gt;The obligations scale up hard once you're huge. If your platform averages more than 10 million senders a month, or 20 million total, the government designates you a "large-scale specified telecommunications service provider" and you owe additional obligations like publishing your takedown response record. As of April 2025, the companies actually carrying that designation are Google, LY Corporation (LINE Yahoo), Meta, and TikTok. A solo developer isn't getting anywhere near that threshold, which is why the lightweight version — one inbox, checked occasionally — is a realistic bar to clear.&lt;/p&gt;

&lt;p&gt;There's also a quieter risk: if your platform lets anyone upload anything, someone eventually will upload something they don't own the rights to. If your code is open source on GitHub, that's a separate reputational problem worth thinking about even before the legal one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbtd0ukvhaerfrjcwtpj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbtd0ukvhaerfrjcwtpj.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The real trap was copyright law
&lt;/h2&gt;

&lt;p&gt;After all that, the thing that actually surprised me wasn't the telecom law or the platform law. It was copyright.&lt;/p&gt;

&lt;p&gt;Japan's &lt;a href="https://www.branche-ip.jp/2014/02/02/%E8%91%97%E4%BD%9C%E6%A8%A9%E6%B3%95%EF%BC%9A%E3%80%8C%E5%85%AC%E8%A1%86%E3%80%8D%E3%81%AB%E5%90%AB%E3%81%BE%E3%82%8C%E3%82%8B%E3%80%8C%E7%89%B9%E5%AE%9A%E3%81%8B%E3%81%A4%E5%A4%9A%E6%95%B0%E3%81%AE/" rel="noopener noreferrer"&gt;Copyright Act defines "the public"&lt;/a&gt; in a way that's broader than it sounds. Article 2, paragraph 5 states that "the public," for purposes of this law, includes "specific and numerous persons" — not just strangers off the street. In plain English: even if everyone who can see your content is someone you personally know and approved individually, if that group gets large enough, the law can treat it the same as posting it publicly. There's no hard headcount in the statute, but the widely cited rule of thumb among Japanese IP lawyers is that once you're past roughly 50 people, you're squarely in "many" territory, and past legal disputes have used numbers like 300+ as clearly qualifying.&lt;/p&gt;

&lt;p&gt;Then there's &lt;a href="https://note.com/copyrights/n/n2a269608a662" rel="noopener noreferrer"&gt;Article 23&lt;/a&gt;, covering the public transmission right. For anything automatically deliverable over a network — which covers basically all web and app content — this right also covers something Japanese law calls "making transmittable" (送信可能化). That means the right is triggered the moment content becomes available for someone to access, whether or not anyone has actually clicked play yet. "Nobody's actually watched it, so I'm fine" doesn't hold up as a legal argument under this framework.&lt;/p&gt;

&lt;p&gt;The counterweight is Article 30, the private-use exception. Using copyrighted material within a genuinely private circle — yourself, your family, people you live with — is exempt. Courts have described the boundary as "an extremely limited circle of personal relationships." A &lt;a href="https://www.businesslawyers.jp/articles/1247" rel="noopener noreferrer"&gt;2022 Supreme Court case&lt;/a&gt; involving JASRAC (Japan's music licensing body) and music schools is a good real-world illustration: the court found that a teacher performing music for students, one at a time, still counted as a "public" performance, because the audience rotated through an effectively unlimited stream of students over time. The lesson generalizes: an audience made of individually-approved people can still add up to "public" if it's large enough or churns enough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxbmrf4fqyfk5da0bldz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxbmrf4fqyfk5da0bldz.png" alt=" " width="800" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Passwords don't automatically make something private
&lt;/h2&gt;

&lt;p&gt;I went into this assuming that if I gated content behind a login and personally approved every viewer, I was safely inside the private-use exception. That turned out to be only half true.&lt;/p&gt;

&lt;p&gt;What decides the boundary isn't whether there's a password. It's who's on the other side of it, and how many of them there are. Family and people you live with land solidly inside the private-use exception — courts have consistently protected that "extremely limited circle." Once you're individually approving friends and acquaintances, the calculus shifts: the more people you add, and the less close the relationship, the more likely a court would call that group "specific and numerous" rather than private, regardless of whether you technically required a password to get in. A login screen controls access technically. It says nothing about whether the underlying use is legally private.&lt;/p&gt;

&lt;p&gt;The uncomfortable implication is that "approve anyone who asks" as a growth strategy — the design pattern, not any specific headcount — reads legally closer to "public" than "private," because the pool has no real ceiling.&lt;/p&gt;

&lt;p&gt;One important caveat, since a reader flagged this after the Japanese version of this post went semi-viral on X: all of this only matters if the content isn't fully yours to begin with. If you personally shot, wrote, and own every frame of what you're distributing, the public-transmission and private-use analysis above is close to irrelevant — copyright belongs to the creator, and you're free to distribute your own original work to as many people as you want. Where this actually bites is content that includes someone else's copyrighted material without permission: a recording that happens to pick up background music, a clip of broadcast TV, a screen recording that captures someone else's copyrighted app or footage. The moment third-party material is baked into what you're distributing, the "how many people, how close are you to them" analysis above is what determines your exposure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add payments or age gates, and yet more laws show up
&lt;/h2&gt;

&lt;p&gt;Everything above covers just the core of a video platform. Add features, and you pick up more regulatory surface.&lt;/p&gt;

&lt;p&gt;The moment you're storing watch history or account data, Japan's Act on the Protection of Personal Information kicks in — a privacy policy that states your purpose of use is the baseline expectation. If minors might realistically use your service, the Act on Development of an Environment that Provides Safe and Secure Internet Use for Young People brings in expectations around age verification and filtering. And if you add tipping, subscriptions, or ad revenue sharing — anything that moves money — the Payment Services Act and the Act on Specified Commercial Transactions apply separately.&lt;/p&gt;

&lt;p&gt;For a personal, non-commercial project, this tier is "know it exists, revisit it if you add the feature" rather than something to solve up front.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what does a solo developer actually need to do
&lt;/h2&gt;

&lt;p&gt;The pattern that emerged after all this reading: the real dividing line isn't "can other people upload," it's "is this actually public." A fully private deployment — access-controlled, URL not shared anywhere — falls inside the private-use exception with real room to spare. Even if a video you personally recorded happens to capture copyrighted material in the background, keeping it to private viewing is generally fine.&lt;/p&gt;

&lt;p&gt;The moment you deploy publicly, or post the URL on a blog where anyone can find it, you've crossed into public transmission, and the private-use exception no longer covers you — even if you're the only person who ever uploaded anything, placing unlicensed third-party material there can be infringing.&lt;/p&gt;

&lt;p&gt;Three things made this manageable for a solo project. First, keep access genuinely locked down — authentication plus a URL you don't publish — if you want to stay inside the private-use exception. Second, for anything you do deploy publicly or show in a GitHub README, use only material you made yourself or that's explicitly license-free; no copyrighted broadcast footage, no commercial music tracks. Third, add a line to your README or terms of service stating the app is intended for the creator's personal use and isn't designed for uploading third-party copyrighted material — it won't prevent a determined bad actor, but it's a reasonable statement of intent if the question ever comes up.&lt;/p&gt;

&lt;p&gt;Building something like automated audio fingerprinting to detect infringing uploads is overkill for a single-user, access-controlled personal project. Locking down access and being deliberate about demo content covers the realistic risk at this scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thought
&lt;/h2&gt;

&lt;p&gt;I genuinely expected video streaming to be a solved, boring technical problem by now. I did not expect to spend this much time in statute text. The good news is that a personal, single-user version of this project needs essentially no legal paperwork. The part that actually changed my mental model was realizing that "I put a password on it" doesn't automatically mean "this is private" under Japanese copyright law — that one took a re-read to sink in.&lt;/p&gt;

&lt;p&gt;Still testing how far this goes on AWS. More to come if there's more to find.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>japan</category>
      <category>legal</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Autonomy for $20, a Human Above It: A Pattern for AI Agents That Spend Money</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Wed, 26 Aug 2026 08:27:21 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/autonomy-for-20-a-human-above-it-a-pattern-for-ai-agents-that-spend-money-425m</link>
      <guid>https://dev.to/_76130e67067eab4c8510/autonomy-for-20-a-human-above-it-a-pattern-for-ai-agents-that-spend-money-425m</guid>
      <description>&lt;p&gt;At some point every "agentic" product roadmap runs into the same uncomfortable question: what happens when the agent needs to actually spend money, not just recommend an action to a human who then clicks a button? Recommending is easy to sandbox. Spending isn't. And the usual answers — "just require approval for everything" or "just trust the agent" — both fail in an obvious way. Approve-everything means the agent isn't really autonomous, it's a slow suggestion box. Trust-the-agent means the first bad prompt injection or hallucinated tool call has a direct line to your bank balance.&lt;/p&gt;

&lt;p&gt;I wanted to see what a middle position actually looks like in code, not in a slide. So I built a small agent, sera-gate-agent, that pays multi-currency invoices for you and enforces a rule most people would agree with intuitively: small payments go through on their own, larger ones stop and wait for a person.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfx8zy289pkghmp24v1y.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfx8zy289pkghmp24v1y.webp" alt=" " width="800" height="798"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dbih1v1oav36058pmpa.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dbih1v1oav36058pmpa.jpg" alt=" " width="800" height="629"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzrc7s4mvcm35ig9phx8w.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzrc7s4mvcm35ig9phx8w.jpg" alt=" " width="800" height="745"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The trigger: a protocol that treats a keypair as a login
&lt;/h3&gt;

&lt;p&gt;The settlement side runs on Sera, an on-chain FX protocol with an MCP server (sera-mcp) released under MIT specifically so agents can call it. What made me want to build on it wasn't the currency coverage, it was the auth model: there's no signup form. You generate a keypair, sign an EIP-712 message with it, and that signature is your API key request. A human finds this mildly inconvenient. An agent finds it exactly as convenient as an auth flow can be — no page to navigate, no CAPTCHA, no session cookie, just a key it already has.&lt;/p&gt;

&lt;p&gt;That detail is what made "give the agent a wallet" feel like a natural next step rather than a stunt. If the whole point of an agent is that it acts without a human driving each click, an auth system built around clicking a human through steps is already fighting the premise.&lt;/p&gt;

&lt;h3&gt;
  
  
  The design: one number, one gated function
&lt;/h3&gt;

&lt;p&gt;The rule I wanted was simple enough to say in one sentence: under $20, the agent executes on its own; over $20, it issues an approval card and a human has to click "approve" in a browser before anything moves. The part that took actual thought wasn't the rule, it was making sure the rule couldn't be talked around.&lt;/p&gt;

&lt;p&gt;The threshold check itself is almost embarrassingly small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decide_auto_or_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SERA_GATE_AUTO_THRESHOLD_USD&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;20&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AUTO&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REQUIRE_APPROVAL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What matters more is where it's called from. The underlying protocol exposes tools that actually move funds — execute a swap, send a transfer, pay an invoice directly. None of those are given to the language model. The only thing the model can call is a wrapper function, and that wrapper is the only code path that's allowed to reach the real execution tools. The threshold check lives inside that wrapper, not in a prompt instruction telling the model to "please ask before spending more than $20."&lt;/p&gt;

&lt;p&gt;That distinction is the whole point, and it's easy to gloss over. A system prompt is a request. A model that's very good at following instructions will follow it correctly almost all the time — but "almost all the time" is a bad security property for something that moves money, whether the failure mode is prompt injection, a weird edge case in tool output, or the model just deciding the instruction doesn't apply this time. Removing the tool from the model's reachable set removes the failure mode instead of making it rarer.&lt;/p&gt;

&lt;p&gt;I also didn't want the agent-side threshold to be the only backstop. Sera has its own server-side policy preset that caps what any single API key can move per transaction and per day, independent of anything my code does. I keep that ceiling well above my $20 agent-side threshold, so it's a true second layer rather than a duplicate of the first — if my code has a bug, or the signing key leaks, there's still a hard ceiling that doesn't route through my logic at all.&lt;/p&gt;

&lt;h3&gt;
  
  
  What this generalizes to
&lt;/h3&gt;

&lt;p&gt;None of this is specific to on-chain payments. The same shape applies to an agent that can issue refunds, send emails to customers, delete records, or push a deploy: pick a dimension the org already has intuitions about (dollar amount, blast radius, reversibility), pick a threshold, and make the boundary a property of which functions are reachable rather than a property of what the model has been told. The dollar threshold is just the easiest one to make legible in a demo, because everyone already has an intuition for what $20 versus $5,000 means.&lt;/p&gt;

&lt;p&gt;The uncomfortable part, and I don't think there's a clean answer to it, is that the threshold is still a judgment call, and it's a judgment call about how much you trust a system that doesn't get tired, doesn't get talked into things the way a person does, but also doesn't have the context a person has for "this specific invoice looks off." Setting it too low turns the agent back into a suggestion box. Setting it too high means the day it's wrong, it's wrong at a scale a human never got the chance to catch. I don't think that number should be static — the next version of this probably ties it to something like payee history or a running risk score instead of a flat constant — but getting the enforcement boundary right, structurally, felt like the part worth building first.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Q2A2epOxGrY"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;sera-gate-agent runs on Sepolia testnet today, built on AWS Bedrock AgentCore Runtime and Strands, with a Next.js front end for the approval cards a human actually clicks through.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>blockchain</category>
      <category>aws</category>
    </item>
    <item>
      <title>Amazon Leo vs Starlink: The 2026 Satellite Internet Race</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 20 Aug 2026 18:09:22 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/amazon-leo-vs-starlink-the-2026-satellite-internet-race-598b</link>
      <guid>https://dev.to/_76130e67067eab4c8510/amazon-leo-vs-starlink-the-2026-satellite-internet-race-598b</guid>
      <description>&lt;p&gt;Project Kuiper quietly became Amazon Leo in November 2025. Enterprise beta opened on April 8, 2026. So the natural question: how close is Amazon actually getting to Starlink?&lt;/p&gt;

&lt;p&gt;Short answer: not close at all. And Amazon just missed a federal deadline in the process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Amazon Leo&lt;/th&gt;
&lt;th&gt;Starlink&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Satellites in orbit&lt;/td&gt;
&lt;td&gt;~345-400 (Aug 2026)&lt;/td&gt;
&lt;td&gt;~10,020 (Apr 2026)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subscribers&lt;/td&gt;
&lt;td&gt;Not yet public (enterprise beta only)&lt;/td&gt;
&lt;td&gt;10M+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countries served&lt;/td&gt;
&lt;td&gt;5 targeted for mid/late 2026 launch&lt;/td&gt;
&lt;td&gt;150+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consumer pricing&lt;/td&gt;
&lt;td&gt;Not yet announced&lt;/td&gt;
&lt;td&gt;$50-120/mo (Residential)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard hardware&lt;/td&gt;
&lt;td&gt;Not yet announced (aiming smaller/cheaper)&lt;/td&gt;
&lt;td&gt;$499 ($199 for Mini)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Launch method&lt;/td&gt;
&lt;td&gt;Atlas V / Falcon 9 / Ariane 6 (outsourced)&lt;/td&gt;
&lt;td&gt;Falcon 9 (in-house, reusable)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FCC status&lt;/td&gt;
&lt;td&gt;Missed the 50% deployment target, waiver carries a spectrum-priority penalty&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The snapshot dates don't line up exactly (Starlink's count is from April, Amazon Leo's from August), but the order of magnitude tells the story either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The satellite count isn't even in the same order of magnitude
&lt;/h2&gt;

&lt;p&gt;As of April 2026, Starlink had &lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;10,020 satellites in orbit versus 241 for Amazon Leo&lt;/a&gt;. By August 2026 Amazon Leo had climbed to roughly &lt;a href="https://www.aboutamazon.com/news/innovation-at-amazon/project-kuiper-satellite-rocket-launch-progress-updates" rel="noopener noreferrer"&gt;345-400 satellites&lt;/a&gt; after a string of launches. Starlink is still ahead by more than 20x.&lt;/p&gt;

&lt;p&gt;Subscriber numbers tell the same story. Starlink serves &lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;over 10 million subscribers across 150+ countries&lt;/a&gt;. Amazon Leo isn't selling to consumers yet — it's still in enterprise beta, with &lt;a href="https://thenextweb.com/news/amazon-leo-satellite-internet-mid-2026" rel="noopener noreferrer"&gt;residential service targeted for mid-2026 in five countries&lt;/a&gt;: the US, Canada, the UK, France, and Germany.&lt;/p&gt;

&lt;h2&gt;
  
  
  Amazon actually missed its FCC deadline
&lt;/h2&gt;

&lt;p&gt;This is the part that surprised me most.&lt;/p&gt;

&lt;p&gt;Amazon's FCC authorization for its 3,236-satellite constellation came with a condition: launch 50% (1,616 satellites) by July 30, 2026, or risk losing priority status. Actual count at the deadline: &lt;a href="https://www.satellitetoday.com/connectivity/2026/06/05/fcc-gives-amazon-leo-50-deployment-waiver-with-conditions-on-spectrum-priority/" rel="noopener noreferrer"&gt;331 satellites&lt;/a&gt; — about 20% of target.&lt;/p&gt;

&lt;p&gt;The FCC granted a waiver, but not for free. Under the &lt;a href="https://www.geekwire.com/2026/fcc-gives-amazon-leo-more-leeway-on-its-satellite-deployment-schedule/" rel="noopener noreferrer"&gt;terms of the extension&lt;/a&gt;, any Gen1 satellite launched after July 30 temporarily loses the spectrum priority status Amazon earned in earlier FCC processing rounds, until March 30, 2028, or until it hits the 50% mark, whichever comes first. In a business where orbital slots and spectrum priority are genuinely scarce, that's a real cost, not just a paperwork inconvenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the gap exists: nobody owns Amazon's rocket
&lt;/h2&gt;

&lt;p&gt;SpaceX runs Starlink and also builds the rocket that launches it. Falcon 9 is reusable, flies constantly, and SpaceX controls its own launch cadence end to end. That vertical integration is the actual root of Starlink's lead.&lt;/p&gt;

&lt;p&gt;Amazon Leo has no equivalent. It buys launches from &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;Atlas V, Falcon 9, and Ariane 6&lt;/a&gt;, and depends on other companies' schedules. Vulcan Centaur and Blue Origin's New Glenn are coming online too — Blue Origin being Bezos-founded but organizationally separate from Amazon — with a target of &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;20+ missions in 2026 and 30+ in 2027&lt;/a&gt;. Still nowhere near Starlink's cadence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Amazon Leo could still win
&lt;/h2&gt;

&lt;p&gt;The one card Starlink genuinely can't match: native AWS integration. A company already running workloads on AWS can get satellite backhaul into its own AWS region as part of one coherent stack. Starlink has no cloud platform to offer alongside it.&lt;/p&gt;

&lt;p&gt;The engineering also reflects a later start. Amazon Leo uses &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;optical inter-satellite links and a custom baseband chip called Prometheus&lt;/a&gt;, and flies at a lower inclination (30-51°) that concentrates coverage on populated mid-latitudes rather than the wider polar coverage Starlink offers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Starlink, for reference
&lt;/h2&gt;

&lt;p&gt;Pricing as of 2026: &lt;a href="https://www.usmobile.com/blog/starlink-cost/" rel="noopener noreferrer"&gt;Residential $50-120/mo, Roam $50-165/mo, Mini $30/mo plus $199 hardware, Business from $250/mo&lt;/a&gt;. Standard hardware holds steady at $499.&lt;/p&gt;

&lt;p&gt;The more interesting move is T-Satellite, Starlink's direct-to-cell partnership with T-Mobile: &lt;a href="https://www.satelliteinternet.com/providers/starlink/starlink-direct-to-cell/" rel="noopener noreferrer"&gt;$10/month, free on higher-tier plans, works across 60+ phone models regardless of carrier&lt;/a&gt;. No military or first-responder discount currently exists on the core service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;On raw numbers, this isn't a race yet — it's a head start plus a company still finding its footing. Amazon has committed &lt;a href="https://www.datacenterdynamics.com/en/news/amazon-promises-invest-more-10bn-project-kuiper-satellite-internet-business/" rel="noopener noreferrer"&gt;more than $10 billion&lt;/a&gt; to closing the gap, and the AWS integration angle is real. Whether it matters depends entirely on whether Amazon Leo can fix its launch cadence before the spectrum penalty clock runs out in March 2028.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;Amazon LEO Vs Starlink: Price, Speed, Latency, and Fit — 5Gstore&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.aboutamazon.com/news/innovation-at-amazon/project-kuiper-satellite-rocket-launch-progress-updates" rel="noopener noreferrer"&gt;Amazon Leo mission updates — About Amazon&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenextweb.com/news/amazon-leo-satellite-internet-mid-2026" rel="noopener noreferrer"&gt;Amazon Leo targets mid-2026 commercial launch — The Next Web&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.satellitetoday.com/connectivity/2026/06/05/fcc-gives-amazon-leo-50-deployment-waiver-with-conditions-on-spectrum-priority/" rel="noopener noreferrer"&gt;FCC Gives Amazon Leo 50% Deployment Waiver, With Conditions on Spectrum Priority — Via Satellite&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.geekwire.com/2026/fcc-gives-amazon-leo-more-leeway-on-its-satellite-deployment-schedule/" rel="noopener noreferrer"&gt;FCC gives Amazon Leo more leeway for deploying satellites — GeekWire&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;Amazon Leo — Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.usmobile.com/blog/starlink-cost/" rel="noopener noreferrer"&gt;Starlink Plans &amp;amp; Pricing In 2026 — US Mobile&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.satelliteinternet.com/providers/starlink/starlink-direct-to-cell/" rel="noopener noreferrer"&gt;Starlink T-Satellite: Cost, Compatible Phones &amp;amp; Coverage — SatelliteInternet.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.navyweek.org/discount/starlink-military-discount/" rel="noopener noreferrer"&gt;Starlink Military Discount 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.datacenterdynamics.com/en/news/amazon-promises-invest-more-10bn-project-kuiper-satellite-internet-business/" rel="noopener noreferrer"&gt;Amazon promises to invest more than $10bn in Project Kuiper — DCD&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>space</category>
      <category>cloud</category>
      <category>satellite</category>
    </item>
    <item>
      <title>I Rebuilt YouTube on AWS Alone (and Hit Every Wall)</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 18 Aug 2026 15:54:36 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-rebuilt-youtube-on-aws-alone-and-hit-every-wall-3jh2</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-rebuilt-youtube-on-aws-alone-and-hit-every-wall-3jh2</guid>
      <description>&lt;p&gt;It started as a simple question: how is YouTube actually built? One thing led to another, and a few hours later I had a single-user, self-hosted video platform running in production on AWS — after redesigning the auth layer from scratch mid-build, chasing down an "exec format error," and discovering that avoiding a NAT Gateway didn't actually save me any money. Here's the whole story, including the parts that didn't work the first time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What YouTube is actually made of
&lt;/h2&gt;

&lt;p&gt;Before writing any code, I wanted to understand what I was copying. Roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Upload and transcoding&lt;/strong&gt;: uploaded video gets converted into 144p through 4K/8K across multiple codecs, processed by a huge fleet of parallel workers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CDN&lt;/strong&gt;: Google Global Cache — dedicated caching nodes placed directly inside ISP networks — plus adaptive bitrate streaming&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metadata&lt;/strong&gt;: Vitess (a sharding layer over MySQL) and Bigtable/Spanner&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recommendations&lt;/strong&gt;: a two-stage candidate generation + ranking ML system&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content ID&lt;/strong&gt;: audio/video fingerprinting to detect copyright infringement&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ads&lt;/strong&gt;: backed by Google Ad Manager&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9hxa0dg3t7n5gwqtvi8.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9hxa0dg3t7n5gwqtvi8.jpg" alt=" " width="800" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Can AWS alone reproduce it?
&lt;/h2&gt;

&lt;p&gt;Most of the functional skeleton maps cleanly onto managed AWS services: S3 for upload, MediaConvert for transcoding, CloudFront for delivery, DynamoDB for metadata, OpenSearch for search, Personalize for recommendations. That part is genuinely achievable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhyzp4pgvp6qppz65qh7o.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhyzp4pgvp6qppz65qh7o.jpg" alt=" " width="800" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few pieces aren't:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Content ID&lt;/strong&gt; has no AWS-managed equivalent. You'd need a third-party SaaS like Audible Magic or ACRCloud, or roll your own fingerprinting with something like Chromaprint against a reference database you'd also have to build&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ISP-embedded caching&lt;/strong&gt; — CloudFront has a global edge network, but nothing at the density of nodes sitting inside individual ISPs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ad auctions&lt;/strong&gt; at Google Ad Manager's scale aren't something you build yourself; you'd hand this off to an existing ad network&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a personal project, I decided not to build Content ID or ad serving at all. That decision turned out to be tied directly to a legal question I hadn't expected to spend time on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The legal research I didn't expect to do
&lt;/h2&gt;

&lt;p&gt;Building a video-sharing app in Japan, even a personal one, touches a surprising number of regulations, so I checked before writing any infrastructure code.&lt;/p&gt;

&lt;p&gt;First, the &lt;strong&gt;Telecommunications Business Act&lt;/strong&gt;. One-way video distribution generally doesn't require registration, since you're not "mediating someone else's communication." Add a comment section or DMs between users, though, and that changes.&lt;/p&gt;

&lt;p&gt;Second, the &lt;strong&gt;Act on the Limitation of Liability for Damages of Specified Telecommunications Service Providers&lt;/strong&gt; (Japan's provider-liability law, recently renamed to something closer to "platform accountability act"). Any platform accepting user-generated content is expected to run a takedown-request contact point; cross a large-user threshold (10M+ monthly users in Japan) and heavier obligations kick in.&lt;/p&gt;

&lt;p&gt;Third — and this is the one that actually shaped the design — &lt;strong&gt;Article 30 of the Copyright Act&lt;/strong&gt;, the private-use reproduction exception. Keep something fully private, accessible only to yourself, and it falls under private use. Make it public and it becomes "transmission to the public" (公衆送信), where that exception no longer applies. I also checked whether gating access behind a login would be enough to stay private if I let a few people in. It isn't automatically: under Japanese copyright law, "the public" includes "a specific but numerous group," so the real question isn't whether there's a login screen, it's &lt;em&gt;how many people, and how close a relationship&lt;/em&gt;. Family-sized is safe; a wider circle of friends risks crossing into "specific but numerous."&lt;/p&gt;

&lt;p&gt;Given all that, I decided the app would support exactly one user — me. No sign-up, no invite flow. That sidesteps the Telecommunications Business Act and the platform-liability law entirely, and keeps everything inside the private-use exception.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plan
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Single user only, no sign-up&lt;/li&gt;
&lt;li&gt;AWS only (I use Vercel for most other projects, but not this one)&lt;/li&gt;
&lt;li&gt;No Content ID, no ad serving&lt;/li&gt;
&lt;li&gt;Upload → S3 → MediaConvert (transcode to HLS) → CloudFront&lt;/li&gt;
&lt;li&gt;Web UI on ECS Fargate + ALB + CloudFront (App Runner was already off the table — AWS stopped accepting new App Runner services)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wrote the CDK for VPC, S3, DynamoDB, a Lambda to kick off MediaConvert jobs, ECS, CloudFront, and Cognito. So far, so normal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cognito's ALB integration needs HTTPS, and I didn't have a domain
&lt;/h2&gt;

&lt;p&gt;My first pass at auth used the ALB's native &lt;code&gt;authenticate-cognito&lt;/code&gt; listener action — no app code needed, ALB handles the redirect to Cognito's hosted UI for you. Clean, until deploy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resource handler returned message: "Actions of type 'authenticate-cognito' are supported only on HTTPS listeners"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That action only works on HTTPS listeners, which means an ACM certificate, which means a real, DNS-verifiable domain — something this project didn't have. Buying a domain just for this felt like the wrong trade, so I moved authentication into the app itself instead, using Next.js's &lt;code&gt;proxy.ts&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Moving auth to the app ran straight into "no NAT Gateway"
&lt;/h2&gt;

&lt;p&gt;I reconfigured the Cognito App Client as a public client (no secret) and switched to PKCE for the authorization code exchange, so the browser could talk to Cognito directly instead of routing through the ALB's constraints.&lt;/p&gt;

&lt;p&gt;Redeployed, logged in, and got a 500 on the callback:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;⨯ [TypeError: fetch failed] {

      at ignore-listed frames {
    code: 'ETIMEDOUT',
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ECS task had no path to the internet. To keep costs down I'd built the VPC with interface endpoints instead of a NAT Gateway — but there's no VPC endpoint for Cognito's Hosted UI/OAuth domain (&lt;code&gt;*.auth.&amp;lt;region&amp;gt;.amazoncognito.com&lt;/code&gt;). The container simply couldn't reach it.&lt;/p&gt;

&lt;p&gt;The fix was to move the token exchange itself into the browser. The only thing that actually needs to happen server-side is JWT verification (fetching the JWKS), which &lt;em&gt;is&lt;/em&gt; covered by the &lt;code&gt;cognito-idp&lt;/code&gt; VPC endpoint. The browser already has internet access, so it can talk to Cognito's token endpoint directly. That change got login working without ever adding a NAT Gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Forgot to pin the CPU architecture, container wouldn't start
&lt;/h2&gt;

&lt;p&gt;Redeployed again, and this time the ECS task crash-looped indefinitely. CloudWatch Logs had exactly one line to offer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exec /usr/local/bin/docker-entrypoint.sh: exec format error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Building the Docker image on an Apple Silicon Mac produces an arm64 image. Fargate defaults to x86_64. Nothing about the mismatch surfaces until the container tries to actually execute. Setting &lt;code&gt;runtimePlatform&lt;/code&gt; to ARM64 on the &lt;code&gt;FargateTaskDefinition&lt;/code&gt; fixed it. I also turned on the ECS deployment circuit breaker at the same time — without it, a failing deployment can take up to three hours to be reported as failed, and I'd already lost about 40 minutes not noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The health check was hitting the login redirect
&lt;/h2&gt;

&lt;p&gt;Next failure: the ALB health check was pointed at &lt;code&gt;/&lt;/code&gt;, which — like every other route — goes through the app's auth gate. An unauthenticated health check gets a 302, the ALB reads that as unhealthy, and the deployment fails outright. Added a dedicated &lt;code&gt;/api/health&lt;/code&gt; route that skips the auth check, and that was that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Video played, but the manifest path was broken
&lt;/h2&gt;

&lt;p&gt;Deployment finally succeeded, upload worked, MediaConvert finished the job — and the video was just a black rectangle.&lt;/p&gt;

&lt;p&gt;The cause: MediaConvert's completion event returns &lt;code&gt;outputGroupDetails.playlistFilePaths&lt;/code&gt; as a full &lt;code&gt;s3://bucket/key&lt;/code&gt; URI, not a bucket-relative key. I'd been storing that value directly as &lt;code&gt;manifestKey&lt;/code&gt;, so the app's &lt;code&gt;/${manifestKey}&lt;/code&gt; template produced a broken &lt;code&gt;/s3://bucket/...&lt;/code&gt; path. Since I already control the output prefix at job-creation time, I switched to deriving the key deterministically instead of trusting the event payload. Don't take an AWS event field at face value if you can compute the same thing yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  CloudFront's signed cookies ignored my wildcard
&lt;/h2&gt;

&lt;p&gt;To lock down &lt;code&gt;/renditions/*&lt;/code&gt; (the actual video files) behind CloudFront's Key Group, I used &lt;code&gt;@aws-sdk/cloudfront-signer&lt;/code&gt;'s &lt;code&gt;getSignedCookies&lt;/code&gt; with a &lt;code&gt;url&lt;/code&gt; + &lt;code&gt;dateLessThan&lt;/code&gt; — the "canned policy" form — expecting a wildcard path to cover everything under it. It didn't:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"AccessDenied"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Access denied"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Along the way I also discovered this AWS account already had a CloudFront Public Key from a different project, and my first debugging attempt had grabbed the wrong Key Pair ID entirely — worth checking &lt;code&gt;aws cloudfront list-public-keys&lt;/code&gt; before assuming there's only one.)&lt;/p&gt;

&lt;p&gt;The actual fix was switching to an explicit custom policy — passing &lt;code&gt;policy&lt;/code&gt; with a JSON statement whose &lt;code&gt;Resource&lt;/code&gt; includes the wildcard — rather than the canned &lt;code&gt;url&lt;/code&gt;/&lt;code&gt;dateLessThan&lt;/code&gt; shortcut. The SDK happily accepts a wildcard in the canned form; CloudFront just doesn't honor it the same way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deleting a video brought it back from the dead
&lt;/h2&gt;

&lt;p&gt;With everything working, I added delete. It's supposed to be a simple DynamoDB + S3 cleanup, but it hit two separate bugs.&lt;/p&gt;

&lt;p&gt;First, IAM: &lt;code&gt;grantWrite&lt;/code&gt;/&lt;code&gt;grantDelete&lt;/code&gt; only cover object-level actions (&lt;code&gt;s3:PutObject*&lt;/code&gt;, &lt;code&gt;s3:DeleteObject*&lt;/code&gt;), not the bucket-level &lt;code&gt;s3:ListBucket&lt;/code&gt; that &lt;code&gt;ListObjectsV2&lt;/code&gt; needs during cleanup. Adding &lt;code&gt;grantRead&lt;/code&gt; fixed it.&lt;/p&gt;

&lt;p&gt;Second, and more interesting: deleting a video that was still processing let the MediaConvert-completion Lambda fire &lt;em&gt;after&lt;/em&gt; deletion, calling &lt;code&gt;UpdateItem&lt;/code&gt; on a videoId that no longer existed. DynamoDB's &lt;code&gt;UpdateItem&lt;/code&gt; creates the item if it's missing — so the "deleted" video would silently reappear, partially populated. Adding &lt;code&gt;ConditionExpression: 'attribute_exists(videoId)'&lt;/code&gt; made that update a no-op instead of a resurrection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The security group I "locked down" wasn't actually locked down
&lt;/h2&gt;

&lt;p&gt;I wanted the ALB reachable only through CloudFront, so I restricted its security group to CloudFront's managed prefix list (&lt;code&gt;pl-58a04531&lt;/code&gt;). Deployed, checked the actual rule set, and found this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"IpRanges"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"CidrIp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.0.0.0/0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"Description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow from anyone on port 80"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"PrefixListIds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"PrefixListId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pl-58a04531"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both rules were live at once. The culprit was the ALB listener's &lt;code&gt;open&lt;/code&gt; property, which defaults to &lt;code&gt;true&lt;/code&gt; and silently adds its own 0.0.0.0/0 ingress rule regardless of what you've configured on the security group yourself. Setting &lt;code&gt;open: false&lt;/code&gt; on &lt;code&gt;addListener&lt;/code&gt; removed it. This is the kind of gap you only catch by actually reading the deployed state back from the AWS CLI — the CDK code alone looked correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  Avoiding a NAT Gateway didn't actually save money
&lt;/h2&gt;

&lt;p&gt;Once things were stable, I priced out the fixed monthly cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Monthly (approx.)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5x VPC interface endpoints&lt;/td&gt;
&lt;td&gt;~$50.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ALB&lt;/td&gt;
&lt;td&gt;~$20–23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECS Fargate (0.25 vCPU / 0.5GB, ARM64)&lt;/td&gt;
&lt;td&gt;~$8.90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secrets Manager&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$83&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I'd built five interface endpoints specifically to avoid a NAT Gateway (roughly $44.60/month in Tokyo, plus data processing). Adding them up, the endpoints cost about the same as the NAT Gateway would have — sometimes more. Consolidating to a single NAT Gateway would save maybe $5–6/month at the cost of a single point of failure, which is a fine trade for a personal, single-user app. I ended up leaving the endpoint-based setup as-is; the savings weren't worth the churn.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually running
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Login via Cognito with PKCE, one user account, no sign-up flow&lt;/li&gt;
&lt;li&gt;Upload → S3 → Lambda → MediaConvert → HLS&lt;/li&gt;
&lt;li&gt;CloudFront with signed cookies gating the video files themselves&lt;/li&gt;
&lt;li&gt;Delete, a Japanese/English toggle, and a dark, YouTube-ish UI&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is &lt;a href="https://github.com/yama3133/mytube" rel="noopener noreferrer"&gt;public on GitHub&lt;/a&gt;, including the README section explaining, in plain terms, why multi-user upload was never on the table.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furx449h8w70unbha5zgo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furx449h8w70unbha5zgo.png" alt=" " width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu7f31m46m7fbhryjy73.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu7f31m46m7fbhryjy73.png" alt=" " width="800" height="487"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mobile&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0a82kh7zvh38rt1imw3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0a82kh7zvh38rt1imw3.png" alt=" " width="624" height="1514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2d4h4iewczjvj0u9065.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2d4h4iewczjvj0u9065.png" alt=" " width="634" height="1514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that actually took the time
&lt;/h2&gt;

&lt;p&gt;None of the individual fixes here were hard once I knew what was wrong. What took the time was reading logs — &lt;code&gt;exec format error&lt;/code&gt;, &lt;code&gt;AccessDenied&lt;/code&gt;, &lt;code&gt;InvalidKey&lt;/code&gt; — each one terse, each one caused by something completely different. Past a certain point, building on managed AWS services stops being about writing code and starts being about getting fast at figuring out why something &lt;em&gt;isn't&lt;/em&gt; working.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cdk</category>
      <category>nextjs</category>
      <category>cognito</category>
    </item>
    <item>
      <title>AWS Instance Store: Built to Disappear, On Purpose</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Fri, 14 Aug 2026 05:32:58 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/aws-instance-store-built-to-disappear-on-purpose-59p7</link>
      <guid>https://dev.to/_76130e67067eab4c8510/aws-instance-store-built-to-disappear-on-purpose-59p7</guid>
      <description>&lt;p&gt;Instance Store gets introduced in almost every AWS storage comparison the same way: "NVMe SSD physically attached to the host, faster than EBS, but the data disappears when the instance stops." That last clause usually reads like a warning label. Most guides then walk you straight into the safe, well-worn use cases — Cassandra nodes that don't mind losing a replica, Spark shuffle space, a scratch disk for sorting temp files. All correct, all a little boring.&lt;/p&gt;

&lt;p&gt;What if the disappearing part isn't the catch, but the whole point?&lt;/p&gt;

&lt;p&gt;Two workloads make that case surprisingly well: blockchain nodes doing a fast state sync, and compute jobs that touch data you'd rather not still have lying around tomorrow. They don't look related at first. One is about speed, the other about disappearance. But they're actually the same trick told twice — you get a disk that works blazingly fast for a short, defined burst, and then erases itself as a side effect of you being done with it. You're not fighting the ephemerality. You're renting it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What instance store actually promises&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Worth being precise here, because the security argument later depends on it. Instance store data does not persist through a stop, a terminate, a hibernate, or an underlying host failure. It does persist through a plain reboot, since that keeps you on the same physical host. AWS documents that the storage is not accessible to whoever gets the host next, which is the property that makes both ideas below work at all.&lt;/p&gt;

&lt;p&gt;What instance store is not: a certified secure-erase mechanism. If your compliance framework requires a documented, auditable wipe procedure — HIPAA, PCI-DSS, that kind of thing — "the disk went away when I stopped the instance" is not a control you can point an auditor at. Keep that distinction in mind as you read the rest of this, because the two ideas here are architectural thought experiments, not a substitute for whatever your compliance team actually needs signed off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Idea one: the node that only exists to catch up&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Syncing a blockchain node from genesis, or even from a recent snapshot, is an I/O-bound slog. You're writing and reading state data continuously for hours, sometimes days, and once the node is caught up, most of that historical grind stops mattering — what you actually want going forward is a warm, synced node.&lt;/p&gt;

&lt;p&gt;The usual move is to provision an instance with a big EBS volume, let it sync, and keep paying for that volume indefinitely. Instance store flips the framing: treat the sync itself as the disposable part. Spin up an instance with local NVMe, let it rip through the sync at NVMe speeds instead of network-attached-storage speeds, and once it's caught up, snapshot the resulting state to S3 or EBS. The instance that did the syncing was never meant to be the long-term home for that data — it was a sprinter, not a warehouse. If it dies mid-sync, you weren't attached to it anyway; you just launch another one and let it catch up again, ideally from a recent checkpoint instead of genesis.&lt;/p&gt;

&lt;p&gt;This is basically the render-farm mentality applied to sync jobs: the compute is consumable, the output is what you keep. It also pairs naturally with Spot — losing a spot instance mid-sync is annoying, not catastrophic, precisely because you never treated its disk as the source of truth.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcqcmefzpdh3061u57ao.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcqcmefzpdh3061u57ao.jpg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Idea two: compute that isn't supposed to remember anything&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now the other direction. Imagine a batch job that has to touch something sensitive for a few minutes — decrypting a payload, running a one-off transformation on data you were only ever supposed to process, not retain. The usual anxiety with EBS-backed compute is the tail: did the volume get deleted on termination, did a snapshot get left behind by accident, is there a stray AMI somewhere with that data baked in.&lt;/p&gt;

&lt;p&gt;Instance store sidesteps most of that tail by construction. Launch the instance, do the job, terminate it. There's no volume to remember to delete, because there was never a persistent volume to begin with. The "forgetting" isn't a cleanup step you have to remember to run — it's what happens automatically when the job's done and you walk away. It's less "secure deletion" and more "the environment was never built to have a memory."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2cv86pqeo4eu3op7r7e.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2cv86pqeo4eu3op7r7e.jpg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'll be upfront that this is the idea I'd stress-test hardest before trusting it with anything actually regulated. It's a genuinely nice property for internal tooling, dev/test data that's sensitive but not audited, or a proof of concept where "the disk goes away" is a reasonable enough story. It is not, on its own, a story you'd want to tell a security auditor for anything under a real compliance regime — that needs KMS-backed encryption, documented key destruction, and probably a paper trail instance store just doesn't produce.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The thread connecting them&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both ideas lean on the same underlying shift: stop treating the disk's short lifespan as a constraint to architect around, and start treating it as the reason the architecture works. The blockchain sync node is fast because nobody's paying the tax of durable storage during the grind. The confidential job is simple because nobody has to remember to clean up after it. In both cases the disappearing act isn't a workaround — it's doing actual work.&lt;/p&gt;

&lt;p&gt;None of this replaces the orthodox use cases. Cassandra nodes, EMR clusters, and CI runners are still the bread and butter of instance store, and for good reason — they're proven, well-documented, and nobody's going to ask you hard questions about why you picked them. But it's worth remembering that "the data goes away" is a spec, not a bug report, and specs can be designed around on purpose. Sometimes the most interesting infrastructure decision is picking the tool that forgets on schedule, and building the rest of the system to expect exactly that.&lt;/p&gt;

</description>
      <category>instancestore</category>
      <category>ec2</category>
    </item>
    <item>
      <title>How to Pause an AI Agent for Human Approval Without a WebSocket</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 13 Aug 2026 15:54:18 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/how-to-pause-an-ai-agent-for-human-approval-without-a-websocket-19cm</link>
      <guid>https://dev.to/_76130e67067eab4c8510/how-to-pause-an-ai-agent-for-human-approval-without-a-websocket-19cm</guid>
      <description>&lt;p&gt;If an AI agent needs a human to approve something mid-task, the instinct is usually to reach for a websocket, a message queue, or some kind of push notification service to bridge the backend and the frontend. I ended up not needing any of that. One DynamoDB row, polled from both sides, does the whole job. I built this for &lt;a href="https://github.com/yama3133/sub-sentry" rel="noopener noreferrer"&gt;SubSentry&lt;/a&gt;, an agent for AWS's Agents for Humans Hackathon that renews clean subscriptions on its own and asks a human before touching anything that looks like a price hike, a duplicate charge, or an unrecognized merchant. This post is about the mechanism underneath that "asking," not the subscription-tracking part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the problem
&lt;/h2&gt;

&lt;p&gt;An agent tool call that needs human approval has to do two contradictory things at once. It has to actually block, because the agent's next step depends on the answer, and it has to somehow let something completely separate (a browser tab, a Slack bot, a CLI) deliver that answer whenever a human gets around to it, which could be five seconds or five minutes later. The backend process and the thing collecting the human's decision don't share memory, don't share a request, and in my case run on entirely different platforms (Bedrock AgentCore Runtime for the agent, Vercel serverless functions for the UI).&lt;/p&gt;

&lt;p&gt;The trick is to stop thinking of it as backend-talks-to-frontend at all. Neither side needs to know the other exists. They both just need to agree on one row.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbbsm46clw88s8ir3tr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbbsm46clw88s8ir3tr.png" alt=" " width="800" height="834"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The row as a mailbox
&lt;/h2&gt;

&lt;p&gt;Here's the actual store, trimmed slightly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;suggested_action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ttl_seconds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;approval_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approval_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subscription_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PENDING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;suggested_action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;suggested_action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasons&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;ttl_seconds&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;decision&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nf"&gt;_dynamo_put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wait_for_decision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;poll_sec&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PENDING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="nf"&gt;_dynamo_put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;poll_sec&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent's tool calls &lt;code&gt;request_approval&lt;/code&gt;, gets back an &lt;code&gt;approval_id&lt;/code&gt;, and immediately calls &lt;code&gt;wait_for_decision&lt;/code&gt; on it, which just sits there polling DynamoDB once a second. That's the entire "block" side. It's a plain Python &lt;code&gt;while True&lt;/code&gt; loop, nothing fancier, because AgentCore Runtime is already paying for a long-running invocation, so there's no reason to make the waiting clever.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other side never has to know it's being waited on
&lt;/h2&gt;

&lt;p&gt;The frontend's job is smaller than it sounds: read rows where &lt;code&gt;status = PENDING&lt;/code&gt;, render them as cards, and when a human clicks Approve or Reject, write the decision back. Here's the write, as a DynamoDB &lt;code&gt;UpdateItem&lt;/code&gt; call from a Next.js API route:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;ddb&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;UpdateCommand&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;TableName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;TABLES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;approvals&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;Key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;UpdateExpression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SET #s = :d, decision = :d, #r = :r, decided_at = :t&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ExpressionAttributeNames&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#s&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;status&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#r&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;reason&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;ExpressionAttributeValues&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:d&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:r&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:t&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:pending&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PENDING&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;ConditionExpression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;attribute_exists(approval_id) AND #s = :pending&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ReturnValues&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ALL_NEW&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;ConditionExpression&lt;/code&gt; is doing more work than it looks like. It means two people can't both approve the same card and have it silently double-apply, and it means a decision can't land on a row that already expired. If the condition fails, DynamoDB throws &lt;code&gt;ConditionalCheckFailedException&lt;/code&gt;, which the route turns into a 409. No locking, no transactions, just a condition on a single-item write.&lt;/p&gt;

&lt;p&gt;And that's the whole contract. The agent doesn't call an API on the frontend. The frontend doesn't call an API on the agent. A CLI can write the same &lt;code&gt;UpdateItem&lt;/code&gt; and it works identically, which is why &lt;code&gt;agent.py approve &amp;lt;id&amp;gt;&lt;/code&gt; from a terminal resolves the exact same pending card as clicking Approve in the browser. Neither side was written with the other in mind, they just both read and write the same table with the same status field.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks if you're not careful
&lt;/h2&gt;

&lt;p&gt;The failure mode I actually hit wasn't in this mechanism, it was one level down. Strands dispatches multiple tool calls from the same agent turn concurrently, and my first pass at local storage (before I had a real DynamoDB table wired up) was a plain read-JSON-modify-write with no locking. Two tool calls landing at nearly the same instant would both read the file, both append their own entry in memory, and whichever one wrote last won, silently dropping the other's write. I found it because a local &lt;code&gt;.approvals.json&lt;/code&gt; file had two writes visibly tangled together mid-file, not because a test failed cleanly.&lt;/p&gt;

&lt;p&gt;The fix was five lines, a &lt;code&gt;threading.Lock&lt;/code&gt; around the read-modify-write section:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;_LOCK&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_local_load&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
    &lt;span class="nf"&gt;_local_save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;DynamoDB's per-item &lt;code&gt;UpdateItem&lt;/code&gt; doesn't have this problem at all, since each write targets one item atomically. The bug only existed because my local dev fallback was reinventing a worse version of what DynamoDB gives you for free. Worth remembering next time a "just write it to a JSON file for now" shortcut feels harmless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just use a websocket
&lt;/h2&gt;

&lt;p&gt;I did consider it, mostly out of habit. But a websocket needs a persistent connection on both ends, which means something has to stay alive to hold it, and AgentCore Runtime invocations and Vercel serverless functions are both built around not staying alive longer than they have to. Polling a table every one to three seconds costs nothing worth optimizing at this scale, and it means either side of the system can restart, redeploy, or die completely mid-wait and the other side won't even notice, because the DynamoDB row is the only thing that has to survive.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/WkiH9TUdEYw"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;If you want to see the whole thing running, live demo and code are here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Live demo: &lt;a href="https://sub-sentry-xi.vercel.app" rel="noopener noreferrer"&gt;https://sub-sentry-xi.vercel.app&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/yama3133/sub-sentry" rel="noopener noreferrer"&gt;https://github.com/yama3133/sub-sentry&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Built for AWS's Agents for Humans Hackathon. #AgentsforHumans&lt;/p&gt;

</description>
      <category>aws</category>
      <category>agentsforhumans</category>
      <category>dynamodb</category>
      <category>architecture</category>
    </item>
    <item>
      <title>Giving Claude Code a Voice (and a Face)</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 13 Aug 2026 03:44:12 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/giving-claude-code-a-voice-and-a-face-30o8</link>
      <guid>https://dev.to/_76130e67067eab4c8510/giving-claude-code-a-voice-and-a-face-30o8</guid>
      <description>&lt;p&gt;I spend a lot of time waiting on Claude Code. Not idle waiting — I tab away, work on something else, and come back later. The problem is "later" is a guess. A run that finishes in 40 seconds and one that's still going 20 minutes later look identical from another window: nothing happens until I check.&lt;/p&gt;

&lt;p&gt;So I gave it a voice. Specifically, Zundamon's voice — a free, widely used Japanese character voice from VOICEVOX — plus a small character that shows up in the corner of my screen when there's something to say.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdqv6l7ibwca593d3menj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdqv6l7ibwca593d3menj.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What it actually does
&lt;/h2&gt;

&lt;p&gt;Claude Code has hooks that fire on specific events: when a turn ends (&lt;code&gt;Stop&lt;/code&gt;), when a subagent finishes (&lt;code&gt;SubagentStop&lt;/code&gt;), and when the CLI is waiting on me for longer than a few seconds (&lt;code&gt;Notification&lt;/code&gt;). Each of those now runs a small Python script that sends a short line of text to a background daemon sitting in my menu bar. The daemon asks VOICEVOX to synthesize it, plays the result, and pops up a borderless little window with Zundamon and a speech bubble.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi4s9p8brcd9pd5j2dvq1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fi4s9p8brcd9pd5j2dvq1.png" alt=" " width="734" height="292"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The daemon also remembers which of four display modes I picked: always on screen, only visible while actually talking, voice-only with no window at all, or fully muted. That choice is saved to a config file so it survives a restart.&lt;/p&gt;

&lt;p&gt;None of this touches the network beyond localhost — the hook script, the daemon, and VOICEVOX Engine all talk to each other over &lt;code&gt;127.0.0.1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8luo0csrwztt3737y1jf.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8luo0csrwztt3737y1jf.jpg" alt=" " width="800" height="1067"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that actually took the time
&lt;/h2&gt;

&lt;p&gt;Wiring the pipeline together was the easy afternoon. Making it not sound broken took a lot longer, and most of the bugs were the kind you only notice by actually listening or actually looking.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;English text came out mangled.&lt;/strong&gt; The first version just piped my last message straight into VOICEVOX. It's a Japanese-only engine, so anything in English got sounded out phonetically and wrong — &lt;code&gt;SubagentStop&lt;/code&gt; came out as something closer to "sa-ba-jento-sutoppu" than anything recognizable. I ended up stripping Markdown, keeping only the first sentence of a message (Claude's replies tend to open with a clean summary), and running the rest through a small hand-built glossary that swaps known English terms for their katakana reading before anything unknown gets dropped. That second part — silently dropping unknown English rather than mangling it — was a deliberate trade-off. It's less informative than reading everything, but it sounds like a voice instead of a glitch.&lt;/p&gt;

&lt;p&gt;That fix had a side effect I didn't catch right away: the same regex that stripped stray English letters was also eating plain digits, since &lt;code&gt;[A-Za-z0-9]&lt;/code&gt; was too broad. Task counts and port numbers just silently disappeared from what got spoken. Tightening the pattern to only match tokens that &lt;em&gt;start&lt;/em&gt; with a letter fixed it without touching how English gets filtered.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A single kanji read wrong.&lt;/strong&gt; The word 角 (corner) is genuinely ambiguous in isolation — it can be read &lt;em&gt;kado&lt;/em&gt; or &lt;em&gt;kaku&lt;/em&gt; depending on context, and VOICEVOX's default analysis picked the wrong one for how I was using it. VOICEVOX Engine exposes a real dictionary API for exactly this (&lt;code&gt;/user_dict_word&lt;/code&gt;), so I registered the correct reading. It didn't take effect. Turned out the default word type is "proper noun," and the built-in dictionary entry for a plain corner apparently wins over a proper-noun override in that grammatical position. Re-registering it as a common noun with a higher priority fixed it immediately.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Failures were completely silent.&lt;/strong&gt; At some point during testing, VOICEVOX had quietly crashed. The hook still fired, the bubble still popped up with the right caption — and nothing played. No error anywhere, because the exception was being caught and swallowed. I added a log line and a recovery path: if synthesis fails, relaunch the engine, wait for it to come back up, and retry once. A crash now costs a few seconds of silence instead of going unnoticed for the rest of the day.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The window was flush with the corner. The character wasn't.&lt;/strong&gt; I pinned the floating window to (0, 0) — the literal bottom-right corner of the screen — and the character's feet still hovered above the edge with a visible gap. The window position was correct; I checked. What wasn't correct was the source artwork: the PNG had transparent padding baked in below the feet, so the visible pixels stopped short of the canvas edge. Cropping the artwork to its actual bounding box before resizing fixed it — no code change to the window logic at all.&lt;/p&gt;


&lt;div&gt;
    &lt;iframe src="https://www.youtube.com/embed/LLP_Z5a_O7s"&gt;
    &lt;/iframe&gt;
  &lt;/div&gt;


&lt;h2&gt;
  
  
  Where it ended up
&lt;/h2&gt;

&lt;p&gt;A menu-bar daemon, four display modes, a self-healing connection to VOICEVOX, a growing glossary for English terms, and a corrected dictionary entry for one very specific kanji. It's a small tool, and I don't think any single piece of it was hard. What was hard was that every failure mode was invisible by default — wrong pronunciation, dropped digits, a dead engine, a mispositioned character — and the only way to catch any of it was to actually sit there and listen, or take a screenshot and zoom in.&lt;/p&gt;

&lt;p&gt;If you're building something similar: budget real time for the boring verification loop, not just the integration. That's where all of this actually lived.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>macos</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Building a Real-Time Earthquake Alert PWA with AI Agents</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Wed, 12 Aug 2026 06:44:20 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/building-a-real-time-earthquake-alert-pwa-with-ai-agents-46c7</link>
      <guid>https://dev.to/_76130e67067eab4c8510/building-a-real-time-earthquake-alert-pwa-with-ai-agents-46c7</guid>
      <description>&lt;p&gt;Japan sits on four tectonic plates. Most people who live here carry that fact quietly in the back of their mind, and every so often it stops being background noise. On July 28, 2026, an earthquake hit Kumamoto and people lost their lives. I want to say that plainly before anything else in this post: my thoughts are with the people who were killed and the people still living with what that day did to their homes and their sense of safety. Everything technical that follows exists because of days like that one.&lt;/p&gt;

&lt;p&gt;I build small AI-agent side projects most weekends, and this time I wanted to build something that actually mattered to me personally rather than something clever for its own sake: a real-time earthquake alert app. Text only, no map to stare at, just "something is happening, here's where, here's how strong." I built it end to end — WebSocket ingestion on AWS, a Strands Agent on Amazon Bedrock AgentCore Runtime writing the alert copy in six languages, a PWA on Vercel that plays a different alert sound depending on severity. Then, while I was still testing it, an actual earthquake swarm started in Kumamoto and the app got tested by real data whether I liked it or not. More on that below.&lt;/p&gt;

&lt;h2&gt;
  
  
  What it does
&lt;/h2&gt;

&lt;p&gt;The rules are simple on purpose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Shows anything of shindo (Japanese seismic intensity) 1 or higher&lt;/li&gt;
&lt;li&gt;Shindo 5-weak and above, Earthquake Early Warning (EEW), and tsunami warnings get a red, emphasized card&lt;/li&gt;
&lt;li&gt;A map of Japan appears only for those emphasized cases, with the affected prefectures highlighted — no map at all for routine shindo 1-2 reports, because most of the time a map adds noise, not signal&lt;/li&gt;
&lt;li&gt;Everything is offered in Japanese, English, Chinese, Korean, Russian, and Arabic, with right-to-left layout for Arabic&lt;/li&gt;
&lt;li&gt;Push notifications work even when the browser tab is closed, using a real audio file rather than the browser's default ping&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this is exotic technology. What made it interesting to build was how many small, real-world details turned out to matter.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyiqmwfzxkmtimyxzngu2.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyiqmwfzxkmtimyxzngu2.jpg" alt=" " width="800" height="767"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture
&lt;/h2&gt;

&lt;p&gt;A Fargate task in ap-northeast-1 holds a persistent WebSocket connection to &lt;a href="https://www.p2pquake.net/" rel="noopener noreferrer"&gt;P2P Earthquake Information&lt;/a&gt;, a free, community-run relay of Japan Meteorological Agency (JMA) data. When a message comes in, the connector filters it, normalizes it into a common shape, and hands the structured data to a Strands Agent running on Amazon Bedrock AgentCore Runtime. The agent's only job is to turn "maxScale: 45, areas: [Wajima, Suzu], tsunami: advisory" into a short, calm headline and summary — in all six languages, in a single call. The result gets written to DynamoDB and pushed to every subscribed browser over Web Push. On the other side, a Next.js PWA on Vercel reads events through a small API Gateway/Lambda layer and renders the live feed.&lt;/p&gt;

&lt;p&gt;I picked Fargate for the connector specifically because a WebSocket listener needs to stay alive, and Lambda's execution model doesn't fit that. Everything else — the API routes, the event storage, the translation — is about as serverless as I could make it.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9n84ec0g9b9wtxjqss2t.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F9n84ec0g9b9wtxjqss2t.jpg" alt=" " width="800" height="469"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The parts that didn't work on the first try
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;The EEW payload lied to me by omission.&lt;/strong&gt; I assumed the &lt;code&gt;areas&lt;/code&gt; array in a P2P quake "555" message (Earthquake Early Warning) would contain municipality-level shindo estimates, the way the finalized "551" earthquake info does. It doesn't. It's peer relay counts inside the P2P network itself — completely unrelated to the earthquake. My first build rendered a wall of "undefined (shindo 1)" entries in production before I caught it and rewrote that part to only pull real detail from the confirmed 551 message that (usually) follows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A SigV4 signature bug cost me most of an evening.&lt;/strong&gt; I'm calling AWS API Gateway from a Vercel function using OIDC federation — no static AWS keys, just a role Vercel assumes per request. Requests without a query string worked fine. Requests with one (&lt;code&gt;?limit=30&lt;/code&gt;) came back 403 every time. IAM policy simulation said "allowed." Assuming the role manually from my own machine and signing the exact same request worked. The difference turned out to be embarrassingly simple: I'd built the query string directly into the request path instead of passing it through the signer's dedicated &lt;code&gt;query&lt;/code&gt; field, so the signature covered a URL that didn't match what actually went out over the wire.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;ARM64 vs x86_64.&lt;/strong&gt; I build the connector's Docker image on an Apple Silicon Mac. Fargate defaults to x86_64. The container built fine, pushed fine, and then died on startup with &lt;code&gt;exec format error&lt;/code&gt;. One line (&lt;code&gt;runtimePlatform: { cpuArchitecture: ARM64 }&lt;/code&gt;) fixed it, but it took a failed deployment to notice.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A chicken-and-egg secret.&lt;/strong&gt; The Fargate task reads its VAPID (Web Push) keys from Secrets Manager. If you let CDK create that secret fresh, it initializes with a random placeholder value, and the container crashes trying to parse it as JSON before you ever get a chance to set the real keys. The fix was to create the secret with real values first, then have CDK import it instead of creating it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;iOS wouldn't let audio play.&lt;/strong&gt; Mobile Safari and Chrome (both WebKit under the hood on iOS) block &lt;code&gt;audio.play()&lt;/code&gt; unless it's called synchronously inside a real user gesture. A push notification arriving is not a user gesture, no matter how you slice it. The fix is a small trick: play a near-silent sound the moment the user taps "enable notifications," which unlocks the audio element for later programmatic playback in that same tab.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;A timezone bug that hid inside a feature I was proud of.&lt;/strong&gt; I added a safety feature: if an EEW alert isn't followed by confirmed earthquake details within 10 minutes, quietly relabel it "unconfirmed" instead of leaving it looking urgent forever. It shipped, and it did nothing. The bug: P2P's timestamps look like &lt;code&gt;2026/08/12 10:40:03&lt;/code&gt;, already in JST, with no timezone marker. Passed straight into &lt;code&gt;new Date()&lt;/code&gt; inside a UTC-timezone container, that string gets interpreted as UTC — silently shifting every event nine hours into the future. &lt;code&gt;now - eventTime&lt;/code&gt; was permanently negative, so the 10-minute check never once fired. I only found it because real alerts sat "urgent" for over an hour and someone (me) asked "why hasn't this changed."&lt;/p&gt;

&lt;h2&gt;
  
  
  Then a real swarm showed up
&lt;/h2&gt;

&lt;p&gt;While I was mid-build, Kumamoto started experiencing a real earthquake swarm — dozens of quakes over more than sixteen hours, several strong enough to trigger EEW. My test app, still very much a work in progress, started receiving genuine alerts back to back. It was uncomfortable in a way a synthetic load test never is: I was watching a tool I'd built for exactly this situation get exercised by the real thing while I was still finding bugs in it. It's also the reason the "unconfirmed" feature and the timezone bug above both got fixed in the same afternoon — nothing motivates a fix like watching your own notification fire for the fourth time in twenty minutes with no idea if it's real.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyoo4qkjgq6mv8lgqrb7m.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyoo4qkjgq6mv8lgqrb7m.png" alt=" " width="800" height="936"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd tell past me
&lt;/h2&gt;

&lt;p&gt;Test with real, messy, live data as early as possible. Every bug above — the EEW payload shape, the SigV4 query string, the timezone parsing — was invisible in my own mocked test data and obvious the moment real traffic hit it. Mocks are useful for exercising code paths, but they can't tell you that your assumptions about a third-party API's shape are wrong, and they definitely can't tell you your timestamp parsing silently breaks in a specific timezone.&lt;/p&gt;

&lt;p&gt;The stack, if it's useful to anyone building something similar: AWS Fargate, Amazon DynamoDB, Amazon API Gateway, AWS Lambda, Amazon Bedrock AgentCore Runtime running a Strands Agent, Next.js on Vercel with OIDC federation for keyless AWS access, and Web Push with VAPID for notifications.&lt;/p&gt;

&lt;p&gt;It's a small app. It won't stop an earthquake. But if it gets one person a few extra seconds of warning, or tells someone in a language other than Japanese that the shaking they just felt was real and roughly how strong, that's the whole point of having built it.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>webdev</category>
      <category>javascript</category>
    </item>
    <item>
      <title>Schrödinger's Payment: An AI Agent for Duplicate Charges</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 11 Aug 2026 15:13:24 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/schrodingers-payment-an-ai-agent-for-duplicate-charges-12o8</link>
      <guid>https://dev.to/_76130e67067eab4c8510/schrodingers-payment-an-ai-agent-for-duplicate-charges-12o8</guid>
      <description>&lt;p&gt;Amazon Web Services published a piece a few weeks ago called &lt;em&gt;&lt;a href="https://aws.amazon.com/blogs/industries/from-connected-to-resilient-cloud-native-payment-connectivity-on-aws/" rel="noopener noreferrer"&gt;From Connected to Resilient: Cloud-Native Payment Connectivity on AWS&lt;/a&gt;&lt;/em&gt;. It's a deep dive into hardening AWS PrivateLink and Resource Gateway for payment networks that speak ISO 8583 over long-lived TCP sessions. Not exactly light reading, but one detail stuck with me: a Network Load Balancer's idle timeout for TLS listeners is fixed at 350 seconds, and if your client's TCP keepalive interval is longer than that, the NLB will quietly kill the connection. Neither side gets an error. The client just... stops hearing back.&lt;/p&gt;

&lt;p&gt;The article frames this as a resilience problem, and it is one. But I kept thinking about it from the other side. If a payment request goes out and the connection dies before the response comes back, what does the &lt;em&gt;client&lt;/em&gt; actually know? Nothing. The charge might have gone through. It might not have. Until someone checks, both are simultaneously true.&lt;/p&gt;

&lt;p&gt;That's a physics joke I couldn't resist building.&lt;/p&gt;

&lt;h2&gt;
  
  
  Schrödinger's Payment
&lt;/h2&gt;

&lt;p&gt;I built a small app that reproduces this exact failure mode and then makes you live in it. You set an idle timeout and a keepalive interval, send a payment, and watch a heartbeat animation play out in real time. If the keepalive pulses arrive often enough, the connection survives and you get a clean ACK. If they don't, the connection resets right on schedule, and the payment falls into superposition. The backend, unaware anyone left, keeps processing and commits the charge anyway. The client just never finds out.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1o3gtfkmptbxt11i3g6t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1o3gtfkmptbxt11i3g6t.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Turn the keepalive interval up past the idle timeout and you get to watch the reset happen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwu52cimmiarbqxi9u1wp.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwu52cimmiarbqxi9u1wp.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh41qperrgjed1glahdxy.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh41qperrgjed1glahdxy.png" alt=" " width="800" height="512"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfdxnjjxloeq0p6iorko.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftfdxnjjxloeq0p6iorko.png" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The interesting part isn't the outage
&lt;/h2&gt;

&lt;p&gt;Anyone who's built a retry mechanism knows the fix for "did my request actually go through": idempotency keys. Retry with the same key, and a well-built server returns the original result instead of charging twice. I built that path first, and it works exactly as expected.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhcjgg8yq74wzpl9j9no3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhcjgg8yq74wzpl9j9no3.png" alt=" " width="800" height="731"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The part I actually wanted to explore was what happens when the &lt;em&gt;client&lt;/em&gt; doesn't retry cleanly. Say the mobile app restarted, or the request got routed through a different channel, and it generates a brand-new idempotency key for what is, to a human, obviously the same purchase. Now the server sees two different keys, the same order, and has to guess whether this is a legitimate second attempt or someone about to get charged twice.&lt;/p&gt;

&lt;p&gt;I didn't want to hardcode that guess as a rule. So instead of a simple "same amount within N seconds = safe," the app hands the situation to Claude through Amazon Bedrock's Converse API. It gets the full record: both attempts, their amounts, timestamps, and any note the retrying party attached. Then it has to decide whether to block it as a duplicate, allow it as a distinct transaction, or, if it isn't confident either way, hand the decision to a human instead of guessing.&lt;/p&gt;

&lt;p&gt;That last option turned out to be the whole point. When I retried with the same amount and no explanation, the agent blocked it outright at 85% confidence, citing the identical amount and a two-second gap. But when I changed the retry amount and added a note like &lt;em&gt;"customer isn't sure if the amount was right, asked to double check,"&lt;/em&gt; it stopped short and asked for a human instead. The amount mismatch could mean a legitimate correction, but the short elapsed time and unacknowledged first attempt kept it from being sure. That's roughly the judgment call a fraud analyst makes, minus the human.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuq6yaw147575jj69uzy1.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fuq6yaw147575jj69uzy1.png" alt=" " width="800" height="832"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ugnw3r30crzjbvpzfrl.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F7ugnw3r30crzjbvpzfrl.png" alt=" " width="800" height="804"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F99iyauur9sngzf0uzqev.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F99iyauur9sngzf0uzqev.jpg" alt=" " width="800" height="500"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where it actually broke
&lt;/h2&gt;

&lt;p&gt;I deployed the first version to Vercel and ran through the same flow with curl to sanity-check it end to end. Sending a payment and retrying with a new key worked as a pair of requests. Then I called the approve endpoint right after, and got back "no record pending approval found." The record I'd just created, seconds earlier, was gone.&lt;/p&gt;

&lt;p&gt;The app's ledger, standing in for the payment backend's source of truth, lives as an in-memory &lt;code&gt;Map&lt;/code&gt; in a Node process. That's fine locally, where &lt;code&gt;next dev&lt;/code&gt; runs everything as one long-lived process. On Vercel, &lt;code&gt;/api/pay&lt;/code&gt;, &lt;code&gt;/api/resolve&lt;/code&gt;, &lt;code&gt;/api/approve&lt;/code&gt;, and &lt;code&gt;/api/ledger&lt;/code&gt; each compile into their own serverless function. They don't share memory at all; I'd built four backends that couldn't talk to each other and only noticed because I happened to test the full sequence with curl instead of clicking through the UI once and calling it done.&lt;/p&gt;

&lt;p&gt;I merged all four into a single dynamic route, &lt;code&gt;/api/action/[type]&lt;/code&gt;, so at least they'd be the same function. That closes most of the gap, since a single function's warm instance does hold state between requests, but it isn't a real fix. Vercel doesn't guarantee that every request lands on the same warm instance. A genuinely production-grade version of this would need actual persistent storage: Vercel KV, DynamoDB, anything that isn't a &lt;code&gt;Map&lt;/code&gt; in Lambda memory. For a demo meant to show off &lt;em&gt;reasoning&lt;/em&gt;, not &lt;em&gt;infrastructure&lt;/em&gt;, I decided that was an honest place to stop. The live version is a &lt;em&gt;simulation&lt;/em&gt;, not a payment backend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try it
&lt;/h2&gt;

&lt;p&gt;The demo runs in English and Japanese, and it's live at &lt;a href="https://schrodinger-payment.vercel.app" rel="noopener noreferrer"&gt;schrodinger-payment.vercel.app&lt;/a&gt;. It's a Next.js app with API routes calling Claude through Amazon Bedrock's Converse API, deployed on Vercel with a scoped IAM user that can only invoke Bedrock models and nothing else. Every screenshot above was captured against the real deployment, agent reasoning included. Nothing here is scripted.&lt;/p&gt;

&lt;p&gt;If you've read the AWS article this is riffing on, Pattern A (three-layer keepalive alignment) is the one this whole thing reproduces. Patterns B through D, covering graceful maintenance windows and tenant isolation, are still on my list to turn into something equally over-engineered.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>nextjs</category>
      <category>bedrock</category>
    </item>
  </channel>
</rss>
