<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Yuuki Yamashita</title>
    <description>The latest articles on DEV Community by Yuuki Yamashita (@_76130e67067eab4c8510).</description>
    <link>https://dev.to/_76130e67067eab4c8510</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3963934%2Ff567e490-409e-4254-8600-f596ed5e7e99.png</url>
      <title>DEV Community: Yuuki Yamashita</title>
      <link>https://dev.to/_76130e67067eab4c8510</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/_76130e67067eab4c8510"/>
    <language>en</language>
    <item>
      <title>Generating electrode stimulation patterns for a simulated cortical prosthesis</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Sat, 03 Oct 2026 08:03:34 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/generating-electrode-stimulation-patterns-for-a-simulated-cortical-prosthesis-2712</link>
      <guid>https://dev.to/_76130e67067eab4c8510/generating-electrode-stimulation-patterns-for-a-simulated-cortical-prosthesis-2712</guid>
      <description>&lt;p&gt;Stimulating the visual cortex makes a person see small spots of light called phosphenes. A visual prosthesis built on that has to decide, for every image, which electrodes get how much current. I wrote that step in PyTorch, with a differentiable simulator on the other end that shows what the stimulation would look like. It runs on a laptop CPU, with 100, 1,000 or 10,000 electrodes.&lt;/p&gt;

&lt;p&gt;This is a simulation. I did not stimulate anything. The only literature numbers in it are the electrode thresholds, though the cortical map and the simulator's design come from papers too. Here is how it is put together, and the evaluation, which surprised me twice.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh3drjhvml9nlrr354ii1.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh3drjhvml9nlrr354ii1.jpg" alt=" " width="800" height="325"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Placing electrodes and phosphenes
&lt;/h2&gt;

&lt;p&gt;Vision is distorted on the cortex: the center of the visual field takes far more area than the periphery. The complex logarithm w = log(z + a) approximates that map (Schwartz, Vision Research 1980). I put electrodes on a jittered grid in cortical coordinates, one patch per hemisphere, and mapped them back to the visual field. The value a = 0.75 degrees is my own choice.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;z&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;np&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uu&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mf"&gt;1j&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;vv&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;A_DEG&lt;/span&gt;            &lt;span class="c1"&gt;# right visual field, in degrees
&lt;/span&gt;&lt;span class="n"&gt;keep&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;z&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;real&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.0&lt;/span&gt;                        &lt;span class="c1"&gt;# drop points that cross the vertical meridian
&lt;/span&gt;&lt;span class="n"&gt;sigma&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;size_scale&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;du&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;ecc&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;A_DEG&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;     &lt;span class="c1"&gt;# phosphene size grows with eccentricity
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Phosphene size is the cortical spread divided by magnification, so it scales with eccentricity plus a. A search over grid sizes gives exactly 100, 1,000 or 10,000 electrodes.&lt;/p&gt;

&lt;h2&gt;
  
  
  From current to percept
&lt;/h2&gt;

&lt;p&gt;Each electrode has a threshold. Brightness is a sigmoid around it, and the percept is the saturated sum of Gaussian blobs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;brightness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current_ua&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sigmoid&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="n"&gt;current_ua&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;thr&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;SLOPE&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;thr&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;forward&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;current_ua&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;total&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;brightness&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;current_ua&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;@&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;basis&lt;/span&gt;   &lt;span class="c1"&gt;# basis: (electrodes, pixels)
&lt;/span&gt;    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;total&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Fernández et al. (2021) reported single-electrode thresholds of 66.8 ± 36.5 µA with a 96-electrode Utah array. I drew thresholds from a lognormal with that mean and SD, clipped to 15 to 140 µA. After clipping the mean is 65.2 and the SD 31.2 µA, so the SD is lower than the paper's. Brightness slope, phosphene size and the 150 µA current limit are my assumptions. The simulator has no timing and no interaction between neighbouring electrodes, and it is a minimal reimplementation inspired by Dynaphos (van der Grinten et al., eLife 2024), not a port.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three encoders
&lt;/h2&gt;

&lt;p&gt;The encoder decides a target brightness for each electrode, and a mapper turns that into current using estimated thresholds:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;brightness_to_current&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;thr_hat&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;clamp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.02&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.98&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;logit&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;b&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="n"&gt;b&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="nf"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;thr_hat&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;SLOPE&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="n"&gt;logit&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;clamp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;I_MAX_UA&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;Simple sampling: each electrode gets the phosphene-weighted mean of the image&lt;/li&gt;
&lt;li&gt;Edges: the same, after a Sobel filter&lt;/li&gt;
&lt;li&gt;Learned: a CNN with 107,017 parameters. It looks at the image at three scales and samples its feature maps at each phosphene's position in the visual field. A small head then combines those features with eccentricity, phosphene size, the estimated threshold and the current budget&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because the learned encoder samples features per electrode, one set of weights serves all three electrode counts. A mean-current limit is enforced by scaling all currents down together.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calibration
&lt;/h2&gt;

&lt;p&gt;A real patient's thresholds are unknown. I simulated the clinic's procedure: for each electrode, a binary search with 7 steps, asking "did you see a spot?" three times per step and taking the majority, against a noisy psychometric response. That is 21 questions per electrode, 21,000 for a 1,000-electrode array. The estimate lands within about 6% of the true threshold, against about 51% when every electrode is assumed to sit at the population mean.&lt;/p&gt;

&lt;h2&gt;
  
  
  Training
&lt;/h2&gt;

&lt;p&gt;I trained end to end through the simulator for 3,000 steps on synthetic scenes (shapes, lines, letters), taking about 10 minutes on a laptop CPU. Each step picks an electrode count, a layout and threshold set, a calibration state (40% uncalibrated, otherwise noisy estimates up to 35%) and a current limit of none, 40, 25 or 15 µA. The loss is mean squared error plus 0.5 × (1 − SSIM) against a slightly blurred scene, and a small penalty on mean current.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluation
&lt;/h2&gt;

&lt;p&gt;I scored 120 held-out scenes with a simulated patient whose layout jitter and thresholds were not seen in training. Two metrics: SSIM against the blurred scene (large regions), and the correlation between the edge maps of percept and scene (fine detail).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Electrodes&lt;/th&gt;
&lt;th&gt;Sampling, uncalibrated&lt;/th&gt;
&lt;th&gt;Sampling, calibrated&lt;/th&gt;
&lt;th&gt;Learned, calibrated&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;0.030 / 0.351&lt;/td&gt;
&lt;td&gt;0.077 / 0.471&lt;/td&gt;
&lt;td&gt;0.274 / 0.603&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000&lt;/td&gt;
&lt;td&gt;0.085 / 0.229&lt;/td&gt;
&lt;td&gt;0.270 / 0.452&lt;/td&gt;
&lt;td&gt;0.540 / 0.568&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10,000&lt;/td&gt;
&lt;td&gt;0.150 / 0.157&lt;/td&gt;
&lt;td&gt;0.470 / 0.431&lt;/td&gt;
&lt;td&gt;0.658 / 0.512&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Each cell is edge correlation / SSIM. Two results stand out:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Calibration is worth about as much as learning. At 1,000 electrodes, simple sampling goes from 0.085 to 0.270 with calibration, and the learned encoder adds another 0.27&lt;/li&gt;
&lt;li&gt;With a 25 µA mean-current limit, the learned encoder at 1,000 electrodes keeps an edge correlation of 0.512, against 0.307 for simple sampling&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The two surprises came from looking at things, not from the table. A comparison figure showed a bright vertical stripe down the middle of the simple-sampling percept. The log map overshoots the vertical meridian near the fovea, so left and right electrodes overlapped. Dropping those points and retraining removed it.&lt;/p&gt;

&lt;p&gt;The second was SSIM. It did not improve with electrode count and even fell (0.603, 0.568, 0.512 for the learned encoder), which contradicted what the images showed. SSIM against a blurred scene mostly rewards large regions, so a smooth percept from 100 electrodes scores well. I added the edge correlation, which does rise (0.274, 0.540, 0.658). SSIM and MSE are also the training objective, so they favor the learned encoder by construction. Edge correlation was not optimized directly.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flumd2bwpc5eqfsydsvh3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flumd2bwpc5eqfsydsvh3.png" alt=" " width="800" height="478"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;p&gt;The evaluation uses synthetic scenes, and the simulated patient comes from the same model that trained the encoder, so transfer to a different patient model is unshown. Phosphenes are fixed to the gaze in a real prosthesis, which would need eye tracking that I left out. These are image metrics, not a test with a person.&lt;/p&gt;

&lt;p&gt;I ran the demo with my laptop camera, locally and on an EC2 instance behind CloudFront. In both, the image got better as I raised the electrode count. I deleted the AWS resources afterwards.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2rkmxlp8xgl1oxl9d5h.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fs2rkmxlp8xgl1oxl9d5h.jpg" alt=" " width="800" height="420"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;This covers the output side. For the input side, my earlier post rebuilds a brain-to-text decoder from public neural data.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Schwartz, Vision Research 1980: &lt;a href="https://www.sciencedirect.com/science/article/abs/pii/0042698980900905" rel="noopener noreferrer"&gt;https://www.sciencedirect.com/science/article/abs/pii/0042698980900905&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Fernández et al., J Clin Invest 2021: &lt;a href="https://www.jci.org/articles/view/151331" rel="noopener noreferrer"&gt;https://www.jci.org/articles/view/151331&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dynaphos (eLife 2024): &lt;a href="https://elifesciences.org/articles/85812" rel="noopener noreferrer"&gt;https://elifesciences.org/articles/85812&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>pytorch</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>Rebuilding a brain-to-text decoder: from 50% to 23.5% word errors</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Sat, 03 Oct 2026 06:49:28 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/rebuilding-a-brain-to-text-decoder-from-50-to-235-word-errors-291n</link>
      <guid>https://dev.to/_76130e67067eab4c8510/rebuilding-a-brain-to-text-decoder-from-50-to-235-word-errors-291n</guid>
      <description>&lt;p&gt;People who can no longer speak can still try to, and the motor cortex activity from that attempt can be decoded into text. Willett et al. (Nature 2023) published the neural data behind their speech BCI, so I rebuilt the decoder from it. The model trained on CPUs for about five hours, and the word error rate on 880 test sentences went from 50.0% to 23.5% over four changes (23.0% on a held-out half at the end).&lt;/p&gt;

&lt;p&gt;This is an offline reanalysis of recorded data. Here is what each change did and what did not help.&lt;/p&gt;

&lt;h2&gt;
  
  
  Data
&lt;/h2&gt;

&lt;p&gt;The Dryad release (competitionData, 3.67 GB, CC0) holds 8,800 training and 880 test sentences from one participant with ALS, recorded on 24 days. Following the baseline, I used area 6v only: threshold crossings and spike-band power on 128 electrodes, 256 features per 20 ms bin, z-scored per recording block.&lt;/p&gt;

&lt;p&gt;Targets are phonemes, 39 plus silence, from the first CMU dictionary pronunciation. Words missing from the dictionary go through g2p_en.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model
&lt;/h2&gt;

&lt;p&gt;The model follows the PyTorch baseline (cffan/neural_seq_decoder):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;conv1d&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;smooth&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;padding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;same&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;groups&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;shape&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# Gaussian smoothing
&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;einsum&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;btd,bdk-&amp;gt;btk&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day_w&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day_b&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;                            &lt;span class="c1"&gt;# per-day input layer
&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;F&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;softsign&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;unfold&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;unsqueeze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;transpose&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                    &lt;span class="c1"&gt;# 32 bins, stride 4
&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;gru&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                                                                 &lt;span class="c1"&gt;# bidirectional GRU
&lt;/span&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;out&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;log_softmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;                                                                 &lt;span class="c1"&gt;# 40 classes + CTC blank
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The per-day input layer absorbs day-to-day drift in what the electrodes pick up. I shrank the GRU from 5 layers and 1,024 units to 3 layers and 512 units (37.8 million parameters) and grouped sentences by length to cut padding. On a c7i.4xlarge with 8 threads that took 1.5 s per batch, against 11.7 s for the baseline size with random batches. 10,000 batches finished in 5 hours 12 minutes.&lt;/p&gt;

&lt;p&gt;Greedy phoneme error rate on the test set: 22.2%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: nearest words, 50.0%
&lt;/h2&gt;

&lt;p&gt;Splitting the greedy phonemes at silences and replacing each group with the closest word from the training vocabulary gives 50.0% word errors. One wrong phoneme changes the word, and a missed silence merges two.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: let an LLM write, 34.6%
&lt;/h2&gt;

&lt;p&gt;I gave Claude Sonnet 4.6 (Amazon Bedrock, temperature 0) the top five phoneme strings from a CTC prefix beam search and asked for the sentence. Word errors dropped to 39.2%. The results were uneven, though. With a single candidate, the model rebuilt "do you hear the sleigh bells ringing" perfectly from "dew you he the stable shells using". It also turned the almost-correct "quite a you movie are base off of that" into "why do you think they are based off of that".&lt;/p&gt;

&lt;p&gt;When the phonemes are bad, the LLM writes a plausible new sentence. A cheap guard helped. If the LLM's word count differs from the decoder's silence-delimited word count by more than about 11%, keep the no-LLM output. I chose the threshold on one half of the test set and applied it to the other half, which gave 34.6%. Decoder confidence (mean max frame probability, beam score) made a weaker switch, at 38.7 to 39.1%.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: decode words, let the LLM pick, 26.0%
&lt;/h2&gt;

&lt;p&gt;The guard treats a symptom. The fix was to stop the LLM from writing. I decoded words directly from the CTC output with a lexicon and a 3-gram, using torchaudio's flashlight decoder, and gave the LLM the 10 best sequences to choose from by number:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;dec&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;ctc_decoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;lexicon&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lexicon.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;tokens.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lm&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lm.arpa&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;nbest&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                  &lt;span class="n"&gt;beam_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;lm_weight&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;word_score&lt;/span&gt;&lt;span class="o"&gt;=-&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;unk_score&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-inf&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
                  &lt;span class="n"&gt;blank_token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;-&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;sil_token&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SIL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;unk_word&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;&amp;lt;unk&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Each lexicon entry ends with SIL, because the training targets put silence after every word. The 3-gram came from the 8,800 training sentences only, written as an ARPA file with absolute discounting and Katz-style backoff. LM weight and word score were picked by 2-fold cross-validation on the test set.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First candidate, no LLM: 29.7%&lt;/li&gt;
&lt;li&gt;LLM picks from the top 10: 26.0%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Two traps on the way:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;With the full CMU dictionary (about 130,000 words) the error rate jumped to around 90%. My LM gave &lt;code&gt;&amp;lt;unk&amp;gt;&lt;/code&gt; about 7% of the probability mass, so every rare dictionary word scored as well as an unknown word and won often. Restricting the lexicon to the LM's 6,569 words fixed it&lt;/li&gt;
&lt;li&gt;The decoder's lexicon trie keeps at most 6 words per pronunciation. I write the lexicon in LM frequency order so the rare homophones are the ones dropped&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Step 4: a bigger language model, 23.5%
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Language model&lt;/th&gt;
&lt;th&gt;First candidate&lt;/th&gt;
&lt;th&gt;LLM picks from 10&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Training sentences only&lt;/td&gt;
&lt;td&gt;29.7%&lt;/td&gt;
&lt;td&gt;26.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LibriSpeech 3-gram, pruned (openslr SLR11)&lt;/td&gt;
&lt;td&gt;29.8%&lt;/td&gt;
&lt;td&gt;not run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tatoeba English (about 2.04 million sentences) + training text counted twice, singletons dropped&lt;/td&gt;
&lt;td&gt;25.7%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23.5%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;LibriSpeech's LM comes from books and did not help, probably because the test sentences are conversational. Tatoeba's short sentences match better. I removed the 39 Tatoeba sentences that were identical to test sentences before building the LM. With it, 297 of 880 sentences came out exactly right, and the oracle (best of the 10 candidates, chosen with the answer) is 19.4%.&lt;/p&gt;

&lt;p&gt;Examples where the LLM picked a better candidate than the decoder's first:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;they close sometime after eat → they close sometime after eight&lt;/li&gt;
&lt;li&gt;i had the bake done on it → i had the brakes done on it&lt;/li&gt;
&lt;li&gt;solving the variables in the occasion → solving the variables in the equation&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It swapped a correct first candidate for a wrong one in 3 sentences.&lt;/p&gt;

&lt;h2&gt;
  
  
  What did not help
&lt;/h2&gt;

&lt;p&gt;By now I had made many choices while looking at the test set, so I split it: even sentences for development, odd sentences for evaluation. I compared methods on development only, then ran two of them once each on evaluation: the Sonnet baseline and Opus, the best on development.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Method&lt;/th&gt;
&lt;th&gt;Dev (440)&lt;/th&gt;
&lt;th&gt;Eval (440)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No LLM&lt;/td&gt;
&lt;td&gt;26.1%&lt;/td&gt;
&lt;td&gt;25.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.6, top 10&lt;/td&gt;
&lt;td&gt;23.8%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23.0%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.6, top 20&lt;/td&gt;
&lt;td&gt;23.9%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 4.6, top 10 + phoneme string as a hint&lt;/td&gt;
&lt;td&gt;25.8%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.6, top 10&lt;/td&gt;
&lt;td&gt;23.4%&lt;/td&gt;
&lt;td&gt;23.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Opus won on development but tied on evaluation (bootstrap 95% interval of the difference: −18 to +11 words), so I report the Sonnet baseline, 23.0%, as the held-out number. Twenty candidates raised the oracle from 19.6% to 18.7% on development, but the model could not find the extra correct ones. With the phoneme hint, Sonnet explained instead of answering in 401 of 440 calls, and the picks got worse.&lt;/p&gt;

&lt;p&gt;The ceiling is the candidate list. Better candidates come from the language model and the acoustic model, not from a smarter selector.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limits
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;The model is smaller than the baseline. Training PER was 4.9% against 22.2% on the test set, so generalization is the first thing to fix&lt;/li&gt;
&lt;li&gt;I did not use the official large n-gram, which needs around 60 GB of memory to build the decoder graph&lt;/li&gt;
&lt;li&gt;Most method choices used cross-validation on the test set. Only the final dev/eval split is a clean held-out evaluation&lt;/li&gt;
&lt;li&gt;The paper's 23.8% comes from a different model, language model, vocabulary and evaluation, so it is not a direct comparison&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftn5rwemwjjddid35a544.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftn5rwemwjjddid35a544.jpg" alt=" " width="800" height="430"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Willett et al., Nature 2023: &lt;a href="https://www.nature.com/articles/s41586-023-06377-x" rel="noopener noreferrer"&gt;https://www.nature.com/articles/s41586-023-06377-x&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dataset (Dryad, CC0): &lt;a href="https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czpq" rel="noopener noreferrer"&gt;https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czpq&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;PyTorch baseline: &lt;a href="https://github.com/cffan/neural_seq_decoder" rel="noopener noreferrer"&gt;https://github.com/cffan/neural_seq_decoder&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Card et al., NEJM 2024: &lt;a href="https://www.nejm.org/doi/full/10.1056/NEJMoa2314132" rel="noopener noreferrer"&gt;https://www.nejm.org/doi/full/10.1056/NEJMoa2314132&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;LibriSpeech LM (openslr SLR11): &lt;a href="https://www.openslr.org/11/" rel="noopener noreferrer"&gt;https://www.openslr.org/11/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Tatoeba: &lt;a href="https://tatoeba.org" rel="noopener noreferrer"&gt;https://tatoeba.org&lt;/a&gt; (English sentences, CC BY 2.0 FR)&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>python</category>
      <category>pytorch</category>
      <category>machinelearning</category>
    </item>
    <item>
      <title>What it takes to remove the monitor and keyboard: decoders, projection and phosphenes</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 29 Sep 2026 07:37:26 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/what-it-takes-to-remove-the-monitor-and-keyboard-decoders-projection-and-phosphenes-1og1</link>
      <guid>https://dev.to/_76130e67067eab4c8510/what-it-takes-to-remove-the-monitor-and-keyboard-decoders-projection-and-phosphenes-1og1</guid>
      <description>&lt;p&gt;The question that started this: can AI send signals into a brain so that an image shows up on someone's retina? I wanted to see AI output with my own eyes and skip the keyboard. I have not built anything yet, so this post is a technical survey of the pieces and a plan for what can be reproduced in software. Figures come from papers and official announcements, and I mark the ones I could only confirm through press reports.&lt;/p&gt;

&lt;p&gt;One correction to the question first. The retina is the input side of vision, and I found no path that sends an image back to it. An image delivered through the brain forms in the visual cortex instead. That gives two independent problems: getting intent into the AI (input) and getting the AI's output into my vision (output).&lt;/p&gt;

&lt;h2&gt;
  
  
  Input: how a brain-to-text decoder works
&lt;/h2&gt;

&lt;p&gt;The systems that work today read attempted speech from motor areas, not free thought. The pipeline in Willett et al. (Nature 2023) is a good baseline:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;256 intracortical electrodes in total, recorded from a person with ALS&lt;/li&gt;
&lt;li&gt;Neural features binned into short time windows&lt;/li&gt;
&lt;li&gt;A recurrent network trained with CTC loss that outputs phoneme probabilities&lt;/li&gt;
&lt;li&gt;A 3-gram language model over a 125,000-word vocabulary to turn phonemes into words&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Result: 62 words per minute, with a 9.1% word error rate on a 50-word vocabulary and 23.8% on 125,000 words. Card et al. (NEJM 2024) later added large language model rescoring, and held 97.5% accuracy for 8.4 months at about 32 words per minute. Metzger et al. (Nature 2023) used 253-channel surface electrodes instead and reported 78 words per minute with a 25.5% error rate on general sentences from a 1,024-word vocabulary.&lt;/p&gt;

&lt;p&gt;The hard part is not the model size. Signals drift from day to day as electrodes shift, so these decoders recalibrate. A common trick is a separate input layer per recording day. Here is a sketch in PyTorch, with placeholder sizes, that I have not run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;torch.nn&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;nn&lt;/span&gt;

&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;SpeechDecoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;Module&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_days&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_feat&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;256&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hidden&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;512&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_phonemes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="nf"&gt;super&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;__init__&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="c1"&gt;# one input layer per recording day absorbs electrode drift
&lt;/span&gt;        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day_in&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;ModuleList&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Linear&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_feat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_feat&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_days&lt;/span&gt;&lt;span class="p"&gt;)])&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;gru&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;GRU&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;n_feat&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;hidden&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;num_layers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;batch_first&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;dropout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Linear&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;hidden&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;n_phonemes&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# +1 for the CTC blank
&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;forward&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;x&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;tanh&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;day_in&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;day&lt;/span&gt;&lt;span class="p"&gt;](&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
        &lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;gru&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;x&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;out&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;log_softmax&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;ctc&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;CTCLoss&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;blank&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;zero_infinity&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Latency has a floor too. Wairagkar et al. (Nature 2025) synthesized voice from a 256-electrode implant with neural processing inside 10 ms. Listeners transcribing the result had a median word error rate of 43.75%, against 96.43% for the participant's own unaided speech.&lt;/p&gt;

&lt;p&gt;Inner speech is an open question. Kunz et al. (Cell 2025) decoded some inner speech in four participants, with word error rates of 26 to 54% at 125,000 words. They also tested a safeguard: the decoder unlocks only when the person imagines a keyword (the paper uses "Chitty Chitty Bang Bang"), and detection exceeded 98%. Without electrodes the numbers drop. The Brain2Qwerty preprint reports character error rates of 32% with MEG and 67% with EEG.&lt;/p&gt;

&lt;p&gt;The practical non-invasive route is muscle activity. The Nature 2025 paper behind Meta's Neural Band trained a surface EMG model that works without per-person calibration, at 20.9 words per minute for handwriting in the air. It shipped on September 30, 2025, bundled with the Ray-Ban Display.&lt;/p&gt;

&lt;h2&gt;
  
  
  Output: two ways to put AI output into my vision
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Retinal projection is an optics problem
&lt;/h3&gt;

&lt;p&gt;A beam focused at the pupil center, sometimes called a Maxwellian view, keeps the image sharp regardless of the eye's focus. QD Laser's retinal scanning display works this way, with RGB lasers and a MEMS mirror. Meta's Orion (a 70 degree prototype, per press reports) and Ray-Ban Display (one eye, reported as 600 by 600 pixels) use waveguides, which is a different approach.&lt;/p&gt;

&lt;p&gt;The constraints come from the conservation of etendue. Field of view and eyebox trade off, so a wide view with a tolerant eyebox needs bigger optics. Two things follow:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The exit pupil has to track the gaze, so eye tracking latency decides whether the image survives a saccade&lt;/li&gt;
&lt;li&gt;At 60 pixels per degree, a 100 by 100 degree view is about 36 million pixels per eye (my arithmetic), so foveated rendering is close to mandatory&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Software can help in the loop. Neural holography (Peng et al., SIGGRAPH Asia 2020, code on GitHub) trains the hologram computation with the camera in the loop, which absorbs real hardware error. The point for a cloud design is that gaze tracking and pupil steering must stay on the device. The generated image can come from the cloud.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cortical stimulation is an encoding problem
&lt;/h3&gt;

&lt;p&gt;A stimulation pipeline needs to answer: for a target image, which electrodes fire, and at what current? A workable structure has four layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;A map from each electrode to the phosphene it produces (position, size, brightness)&lt;/li&gt;
&lt;li&gt;A differentiable simulator that predicts the percept from the stimulation&lt;/li&gt;
&lt;li&gt;An encoder network trained end to end through that simulator&lt;/li&gt;
&lt;li&gt;Per-patient tuning from the patient's reports
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;stim&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;encoder&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;frame&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;             &lt;span class="c1"&gt;# (batch, n_electrodes) currents
&lt;/span&gt;&lt;span class="n"&gt;percept&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;simulator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;stim&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;         &lt;span class="c1"&gt;# differentiable phosphene image
&lt;/span&gt;&lt;span class="n"&gt;loss&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;torch&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;nn&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;functional&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mse_loss&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;percept&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;simplified_scene&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Dynaphos (van der Grinten et al., eLife 2024) is a PyTorch simulator for step 2. The evidence for the underlying effect is small so far. Fernández et al. (2021) implanted a 96-electrode Utah array in a fully blind 57-year-old participant for six months, and the participant identified some letters and object outlines. Beauchamp et al. (Cell 2020) traced letters by stimulating electrodes in sequence, and four sighted and two blind participants recognized them. Chen et al. (Science 2020) got monkeys to perceive shapes through 1,024 channels.&lt;/p&gt;

&lt;p&gt;The bandwidth gap explains why this is slow. The retina's output is estimated at about 10 Mbps, extrapolated from guinea pig recordings of about 100,000 ganglion cells (Koch et al., 2006). Even counting one bit per pulse, a 1,000-electrode implant stimulated 10 times a second carries about 10 kbps, three orders of magnitude below that (my own estimate). Neuralink's Blindsight got an FDA Breakthrough Device designation in September 2024, but I could not confirm a human implant.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F522zse105gepnblcidhp.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F522zse105gepnblcidhp.jpg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What can be reproduced in software
&lt;/h2&gt;

&lt;p&gt;Datasets exist for each piece:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Piece&lt;/th&gt;
&lt;th&gt;Data or tool&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Speech decoder&lt;/td&gt;
&lt;td&gt;Willett 2023 data on Dryad, and Kaggle's Brain-to-Text '25 built on the data from Card et al.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;EMG typing&lt;/td&gt;
&lt;td&gt;emg2qwerty (NeurIPS 2024, 108 people, 346 hours, CC BY-NC-SA 4.0)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Phosphene simulation&lt;/td&gt;
&lt;td&gt;Dynaphos, and pulse2percept for retinal implants&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hologram computation&lt;/td&gt;
&lt;td&gt;The neural holography code from Stanford&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On AWS, a replay of recorded data can run through AWS IoT Core or Amazon Kinesis Data Streams into an Amazon EC2 GPU instance for the decoder, with Amazon Bedrock for language model rescoring. A second EC2 GPU instance can run the phosphene simulator, and Amazon DCV can stream the result to a browser. When I checked in September 2026, my account's Amazon SageMaker AI quota for GPU training jobs was 0, so training would use EC2 GPU instances, or CPU for a small GRU.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38lu9jg51n87jjtod5ss.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F38lu9jg51n87jjtod5ss.jpg" alt=" " width="800" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I would start
&lt;/h2&gt;

&lt;p&gt;A simulated cortical vision demo: a webcam frame becomes the phosphene image that 1,000 electrodes could produce. It needs no data collection and runs entirely in software. Replaying a public brain-to-text dataset through the decoder comes next. Electrodes and optics stay out of reach for now, and I would not claim otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  References
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Willett et al., Nature 2023: &lt;a href="https://pmc.ncbi.nlm.nih.gov/articles/PMC10468393/" rel="noopener noreferrer"&gt;https://pmc.ncbi.nlm.nih.gov/articles/PMC10468393/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Metzger et al., Nature 2023: &lt;a href="https://www.nature.com/articles/s41586-023-06443-4" rel="noopener noreferrer"&gt;https://www.nature.com/articles/s41586-023-06443-4&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Card et al., NEJM 2024: &lt;a href="https://doi.org/10.1056/nejmoa2314132" rel="noopener noreferrer"&gt;https://doi.org/10.1056/nejmoa2314132&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Wairagkar et al., Nature 2025 (press release): &lt;a href="https://www.eurekalert.org/news-releases/1087025" rel="noopener noreferrer"&gt;https://www.eurekalert.org/news-releases/1087025&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Kunz et al., Cell 2025: &lt;a href="https://pubmed.ncbi.nlm.nih.gov/40816265/" rel="noopener noreferrer"&gt;https://pubmed.ncbi.nlm.nih.gov/40816265/&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Brain2Qwerty (arXiv): &lt;a href="https://arxiv.org/abs/2502.17480" rel="noopener noreferrer"&gt;https://arxiv.org/abs/2502.17480&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Meta sEMG, Nature 2025: &lt;a href="https://www.nature.com/articles/s41586-025-09255-w" rel="noopener noreferrer"&gt;https://www.nature.com/articles/s41586-025-09255-w&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Fernández et al., J Clin Invest 2021: &lt;a href="https://www.jci.org/articles/view/151331" rel="noopener noreferrer"&gt;https://www.jci.org/articles/view/151331&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Beauchamp et al., Cell 2020: &lt;a href="https://www.cell.com/cell/fulltext/S0092-8674(20)30496-7" rel="noopener noreferrer"&gt;https://www.cell.com/cell/fulltext/S0092-8674(20)30496-7&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Chen et al., Science 2020: &lt;a href="https://www.science.org/doi/10.1126/science.abd7435" rel="noopener noreferrer"&gt;https://www.science.org/doi/10.1126/science.abd7435&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Koch et al., Current Biology 2006: &lt;a href="https://www.cell.com/current-biology/fulltext/S0960-9822(06)01639-3" rel="noopener noreferrer"&gt;https://www.cell.com/current-biology/fulltext/S0960-9822(06)01639-3&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Dynaphos (eLife 2024): &lt;a href="https://elifesciences.org/articles/85812" rel="noopener noreferrer"&gt;https://elifesciences.org/articles/85812&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Neural holography: &lt;a href="https://github.com/computational-imaging/neural-holography" rel="noopener noreferrer"&gt;https://github.com/computational-imaging/neural-holography&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Willett 2023 dataset (Dryad): &lt;a href="https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czpq" rel="noopener noreferrer"&gt;https://datadryad.org/dataset/doi:10.5061/dryad.x69p8czpq&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Brain-to-Text '25 (Kaggle): &lt;a href="https://www.kaggle.com/competitions/brain-to-text-25" rel="noopener noreferrer"&gt;https://www.kaggle.com/competitions/brain-to-text-25&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;emg2qwerty: &lt;a href="https://github.com/facebookresearch/emg2qwerty" rel="noopener noreferrer"&gt;https://github.com/facebookresearch/emg2qwerty&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Neuralink Blindsight (FDA designation): &lt;a href="https://neuralink.com/updates/neuralink-receives-breakthrough-device-designation-for-blindsight/" rel="noopener noreferrer"&gt;https://neuralink.com/updates/neuralink-receives-breakthrough-device-designation-for-blindsight/&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>ai</category>
      <category>machinelearning</category>
      <category>aws</category>
      <category>hardware</category>
    </item>
    <item>
      <title>I Asked Claude Code Why AI Labs Suddenly Want to Slow AI Down</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Mon, 14 Sep 2026 01:38:18 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-asked-claude-code-why-ai-labs-suddenly-want-to-slow-ai-down-gmg</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-asked-claude-code-why-ai-labs-suddenly-want-to-slow-ai-down-gmg</guid>
      <description>&lt;h2&gt;
  
  
  The trigger
&lt;/h2&gt;

&lt;p&gt;On September 12, 2026, three competitors, Anthropic, OpenAI, and xAI, all said roughly the same thing within hours of each other: AI development needs to slow down. Companies that fight over the same enterprise contracts and the same researchers do not usually agree on anything in public. So I opened Claude Code and asked it to walk through what was actually going on.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round one: the nuclear weapons comparison
&lt;/h2&gt;

&lt;p&gt;My first question was whether this looked like the nuclear non-proliferation playbook: the people who built the weapon are the ones who later argue loudest for controlling it.&lt;/p&gt;

&lt;p&gt;Claude Code's answer split the comparison into what holds and what does not. What holds: the leading developer warning about the danger of what it built, and a first mover shaping the resulting regulation in a way that locks in its own position, is a real structural echo of the NPT era. What breaks down: nuclear proliferation has a physical bottleneck, uranium enrichment, that can be detected from orbit. AI capability spreads at the speed of compute and data, with no equivalent tripwire. Nuclear weapons are held specifically to never be used; frontier AI models are built specifically to be used, constantly, for profit. The incentive structures point in opposite directions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round two: fact-checking the essay itself
&lt;/h2&gt;

&lt;p&gt;Rather than take the "slow down" framing at face value, I had Claude Code pull the actual source: Dario Amodei's essay "We Must Pace the Frontier." Worth doing, because most coverage ("AI CEOs want to slow AI down") undersells what he is actually proposing. He is explicit that pacing "does not mean halting model training." The real ask is third-party testing across four risk categories, permanent embedded evaluators inside Anthropic, and an antitrust carve-out so labs can coordinate on safety without violating competition law.&lt;/p&gt;

&lt;p&gt;That last item is the one critics keyed in on. Chamath Palihapitiya's response on X: "Dario makes the case to stop open source and concentrate enormous technological and economic power with Anthropic." Whether or not you buy that read, it is a materially different claim than "Anthropic wants AI to be safer," and it is worth knowing both versions exist before writing a hot take.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round three: the detail I almost missed
&lt;/h2&gt;

&lt;p&gt;I asked Claude Code to check whether anything else happened around that week that might be relevant. It surfaced a threat intelligence report Anthropic had published two days earlier, on September 10: seven Chinese AI labs, including Moonshot AI (maker of the Kimi models), accused of large-scale distillation of Claude, GPT, Gemini, and Grok. The specific claim against Moonshot: roughly 300,000 requests routed to Claude over ten days through a network of over 5,000 fraudulent accounts, with Claude's answers served back to users labeled as Kimi's own. US intelligence agencies backed the concern.&lt;/p&gt;

&lt;p&gt;Two days between that report and the "slow the frontier" essay is a short gap. Claude Code was careful not to claim a confirmed causal link, correctly, since none of the primary sources draw one. But it flagged the sequence as worth including, because both stories run on the same anxiety: capability spreading somewhere it should not.&lt;/p&gt;

&lt;h2&gt;
  
  
  Round four: the rumor that did not survive fact-checking
&lt;/h2&gt;

&lt;p&gt;The same week, a rumor was circulating that Huawei's founder and his family had fled China. I asked Claude Code to fold that into the piece too. It refused to treat it as established, and pushed back before writing anything: the only source was a single screenshot posted to Chinese social media, the original poster had called it unverified, Huawei and Chinese authorities had not commented, and Chinese social media reaction skewed skeptical. It is in this piece as an unconfirmed rumor that circulated the same week, not as evidence of anything.&lt;/p&gt;

&lt;p&gt;It would have been easy to let an interesting-sounding rumor slide into a paragraph about geopolitical pressure without that check. Separating what is verified from what is circulating on social media, before it goes into a draft, is the step that gets skipped most often when writing under deadline.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgi29dk8bnf1x85bommb.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwgi29dk8bnf1x85bommb.jpg" alt=" " width="799" height="386"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Takeaway
&lt;/h2&gt;

&lt;p&gt;No single motive explains three competitors agreeing in public. Some of the fear about near-term AGI timelines is probably genuine. The specific regulatory ask looks a lot like regulatory capture. And the timing next to the distillation report suggests a geopolitical containment story running underneath the safety framing. All three can be true at once, which is a less satisfying headline than any one of them alone, but it is the one the primary sources actually support.&lt;/p&gt;

&lt;p&gt;Sources: &lt;a href="https://darioamodei.com/post/we-must-pace-the-frontier" rel="noopener noreferrer"&gt;Dario Amodei, "We Must Pace the Frontier"&lt;/a&gt;, &lt;a href="https://www.axios.com/2026/09/12/anthropic-ai-amodei-pacing" rel="noopener noreferrer"&gt;Axios&lt;/a&gt;, &lt;a href="https://x.com/chamath/status/2098780471966802037" rel="noopener noreferrer"&gt;Chamath Palihapitiya on X&lt;/a&gt;, &lt;a href="https://www.coindesk.com/tech/2026/09/09/moonshot-s-kimi-rattled-markets-u-s-agencies-now-say-it-was-trained-on-american-models" rel="noopener noreferrer"&gt;CoinDesk on the Moonshot/Kimi distillation report&lt;/a&gt;. The Huawei rumor traces to a single unverified screenshot with no mainstream corroboration; see &lt;a href="https://hidamaricolumn.com/huawei-ren-escape-rumor/" rel="noopener noreferrer"&gt;this Japanese-language writeup&lt;/a&gt; for the closest thing to a source trail.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>anthropic</category>
      <category>claudecode</category>
      <category>opensource</category>
    </item>
    <item>
      <title>I Tested Whether cdkd Really Deploys Faster Than cdk deploy</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 03 Sep 2026 15:58:25 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-tested-whether-cdkd-really-deploys-faster-than-cdk-deploy-25i4</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-tested-whether-cdkd-really-deploys-faster-than-cdk-deploy-25i4</guid>
      <description>&lt;p&gt;A tool claiming "up to 15x faster than cdk deploy" showed up in my feed a while back. Drop-in replacement, it said: keep your CDK app exactly as it is, just swap &lt;code&gt;cdk deploy&lt;/code&gt; for &lt;code&gt;cdkd deploy&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I've learned to be skeptical of "Nx faster" claims. So I actually deployed something real to AWS with both tools and timed it. Short version: it really is that fast.&lt;/p&gt;

&lt;h2&gt;
  
  
  What cdkd actually is
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/go-to-k/cdkd" rel="noopener noreferrer"&gt;cdkd&lt;/a&gt; deploys an existing AWS CDK app without going through CloudFormation. It calls the AWS SDK directly instead. It's built by &lt;a href="https://github.com/go-to-k" rel="noopener noreferrer"&gt;go-to-k&lt;/a&gt; (Kenta Goto), an AWS DevTools Hero and CDK top contributor who also maintains &lt;code&gt;cls3&lt;/code&gt; (a fast S3 bucket emptier) and &lt;code&gt;delstack&lt;/code&gt; (for cleaning up stuck CloudFormation/CDK stacks) — tools that quietly fix the annoying parts of working with AWS. cdkd feels like the biggest one yet, and I mean that as a compliment grounded in actually using it, not a throwaway one.&lt;/p&gt;

&lt;p&gt;The mechanism is straightforward. cdkd runs the exact same CDK synth step as the CDK CLI, producing the same CloudFormation template. What changes is everything after that: instead of handing the template to CloudFormation, cdkd's own engine reads the resource dependency graph (&lt;code&gt;Ref&lt;/code&gt;, &lt;code&gt;Fn::GetAtt&lt;/code&gt;), builds a DAG, and fires AWS SDK / Cloud Control API calls directly, in parallel, as soon as each resource's dependencies are satisfied.&lt;/p&gt;

&lt;p&gt;Worth saying up front: cdkd calls itself not production-ready, dev/test only. This isn't a "replace CloudFormation in prod" pitch.&lt;/p&gt;

&lt;h2&gt;
  
  
  I actually ran both, on real AWS
&lt;/h2&gt;

&lt;p&gt;cdkd's own README backs up the 15x number with a VPC + Lambda + SQS + CloudFront benchmark. So I wrote that same stack as a CDK app and deployed it twice — &lt;code&gt;DeployRaceCfn&lt;/code&gt; via &lt;code&gt;cdk deploy&lt;/code&gt;, &lt;code&gt;DeployRaceCdkd&lt;/code&gt; via &lt;code&gt;cdkd deploy&lt;/code&gt; — to the same AWS account, same region (ap-northeast-1).&lt;/p&gt;

&lt;p&gt;The stack:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;VPC (2 AZ + NAT Gateway) with a Lambda inside it, fronted by a Function URL&lt;/li&gt;
&lt;li&gt;CloudFront, origin set to that Function URL&lt;/li&gt;
&lt;li&gt;SQS + EventSourceMapping + a consumer Lambda&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;First attempt failed. The account had hit its VPC limit (five, the default):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resource handler returned message: "The maximum number of VPCs has been reached.
(Service: Ec2, Status Code: 400, ...)"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Other test stacks in the same account were sitting on VPCs I'd forgotten about. I tore down the cdkd stack (already measured, no longer needed) to free a slot and reran. cdkd writes a structured event log to S3 on every run (&lt;code&gt;cdkd events&lt;/code&gt;), so even the failed attempt was easy to diagnose after the fact.&lt;/p&gt;

&lt;p&gt;Timing compares the deploy phase only. Synth is identical work either way (same &lt;code&gt;aws-cdk-lib&lt;/code&gt;), so I excluded it, matching how cdkd's own benchmarks are measured. The &lt;code&gt;cdk deploy&lt;/code&gt; timeline comes from CloudFormation's &lt;code&gt;DescribeStackEvents&lt;/code&gt;; the &lt;code&gt;cdkd deploy&lt;/code&gt; timeline comes from &lt;code&gt;cdkd events &amp;lt;stack&amp;gt; --run &amp;lt;id&amp;gt; --format json&lt;/code&gt;, reading the &lt;code&gt;RESOURCE_STARTED&lt;/code&gt; / &lt;code&gt;RESOURCE_SUCCEEDED&lt;/code&gt; events it records itself.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;cdk deploy&lt;/th&gt;
&lt;th&gt;cdkd deploy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Time&lt;/td&gt;
&lt;td&gt;479.3s&lt;/td&gt;
&lt;td&gt;95.0s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resources created&lt;/td&gt;
&lt;td&gt;34&lt;/td&gt;
&lt;td&gt;33&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;5.0x.&lt;/strong&gt; The one extra resource on the CloudFormation side is &lt;code&gt;AWS::CDK::Metadata&lt;/code&gt;, a bookkeeping resource that only exists there — both sides build the same 33 real resources.&lt;/p&gt;

&lt;p&gt;Watching &lt;code&gt;cdkd deploy&lt;/code&gt; run, IAM roles and route tables land in a burst right at the start:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[1/33] ✓ RaceQueueE818AC65 (AWS::SQS::Queue) created
[2/33] ✓ RaceVpcIGW94C1C01D (AWS::EC2::InternetGateway) created
[3/33] ✓ RaceVpcPublicSubnet1EIPC3B3497E (AWS::EC2::EIP) created
[4/33] ✓ ConsumerFunctionServiceRole68E8FEB1 (AWS::IAM::Role) created
[5/33] ✓ MainFunctionServiceRole8C918DF0 (AWS::IAM::Role) created
...
CloudFront Distribution Distribution830FAC52 accepted (not waiting for Deployed; pass --full-wait to wait)
Deployment Summary:
  Created: 33 / Duration: 94.99s
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Meanwhile &lt;code&gt;cdk deploy&lt;/code&gt; takes 16.6 seconds just to get its first VPC. By that point cdkd already has the Lambda wired up to SQS.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the time actually goes
&lt;/h2&gt;

&lt;p&gt;cdk deploy hands the template to CloudFormation, which creates resources one at a time. cdkd reads the same template's dependency graph itself and calls the AWS SDK / Cloud Control API directly, in parallel, as soon as dependencies clear. Cutting out the CloudFormation middleman is most of the story, but two specific waits explain most of the 384-second gap.&lt;/p&gt;

&lt;p&gt;The NAT Gateway is the first one. cdk deploy works through the SQS/Lambda side of the graph before it gets around to waiting on the NAT Gateway to come up. cdkd hits that wait much earlier. Same AWS-side wait either way — what differs is how early in the run you eat it.&lt;/p&gt;

&lt;p&gt;CloudFront is the bigger one. CloudFormation's default behavior waits until the distribution reaches &lt;code&gt;Deployed&lt;/code&gt; — full global propagation, three-plus minutes. cdkd's default returns as soon as &lt;code&gt;CreateDistribution&lt;/code&gt; is accepted (there's a &lt;code&gt;--full-wait&lt;/code&gt; flag if you want CloudFormation's behavior instead). Of the 479.3 seconds cdk deploy took, 183 of them — over three minutes — are spent solely on that one wait.&lt;/p&gt;

&lt;h2&gt;
  
  
  This might actually change how people deploy
&lt;/h2&gt;

&lt;p&gt;I'll admit it: the slowness of iterating on a real AWS resource is part of why people reach for Vercel or Amplify instead. Build a VPC + Lambda + CloudFront stack in CDK and every check-your-work loop costs minutes, sometimes double digits of them.&lt;/p&gt;

&lt;p&gt;cdkd doesn't replace CloudFormation's state management, drift detection, or rollback handling, and it says so itself. But "CloudFormation in prod, cdkd while iterating" is now a real option, not a hypothetical. A CI pipeline that rebuilds a PR environment on every push, or just the apply-then-check loop on your own machine, running five times faster is enough to make "AWS-native is slower to iterate on than Vercel or Amplify" stop being true for a chunk of use cases.&lt;/p&gt;

&lt;p&gt;I also built an interactive replay of the actual measured timeline, plus a narrated demo video, if you want to see the two runs side by side.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00agsh7vxmj5wpz8sdvz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F00agsh7vxmj5wpz8sdvz.png" alt=" " width="800" height="382"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cpxd8adrvnftf04goeq.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8cpxd8adrvnftf04goeq.png" alt=" " width="800" height="320"&gt;&lt;/a&gt;&lt;br&gt;
  &lt;iframe src="https://www.youtube.com/embed/F9_ewhE9zY8" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;cdkd is &lt;a href="https://github.com/go-to-k/cdkd" rel="noopener noreferrer"&gt;go-to-k/cdkd&lt;/a&gt; (Apache-2.0). Worth a star if this was useful.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cdk</category>
      <category>cloudformation</category>
      <category>devops</category>
    </item>
    <item>
      <title>Detecting Shadow AI Agents with AWS Agent Registry</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 01 Sep 2026 02:56:07 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/detecting-shadow-ai-agents-with-aws-agent-registry-o4c</link>
      <guid>https://dev.to/_76130e67067eab4c8510/detecting-shadow-ai-agents-with-aws-agent-registry-o4c</guid>
      <description>&lt;p&gt;AWS Agent Registry (GA August 2026) gives a team a private catalog for AI agents, MCP servers, and tools — semantic search, approval workflows, CloudTrail audit trails. What it doesn't give you out of the box is a way to find the agents that were &lt;em&gt;never registered in the first place&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;This is a pattern for closing that gap: scan the AWS account for AgentCore runtimes, diff them against what's actually in the registry, and route anything unregistered through a real approval flow before it becomes a permanent record. I built it as &lt;strong&gt;Shadow Agent Hunter&lt;/strong&gt;; this post is about the Agent Registry API details that made it work (and the ones that didn't, at first).&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpiivmq47ddts83oh45h.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdpiivmq47ddts83oh45h.jpg" alt=" " width="800" height="508"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The three API calls that matter
&lt;/h2&gt;

&lt;p&gt;Agent Registry splits into a control plane (&lt;code&gt;agent-registry-control&lt;/code&gt;) for managing registries and records, and a data plane (&lt;code&gt;agent-registry&lt;/code&gt;) for searching approved ones.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Finding what's already registered.&lt;/strong&gt; &lt;code&gt;list_registry_records&lt;/code&gt; returns every record regardless of status, so pending/draft records don't get re-flagged on a rescan either:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;

&lt;span class="n"&gt;control&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent-registry-control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;registered_names&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;set&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;kwargs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;registryId&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nextToken&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;
    &lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_registry_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;registered_names&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;registryRecords&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;next_token&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;nextToken&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;next_token&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I originally reached for the &lt;code&gt;provenance&lt;/code&gt; field here, expecting to link records back to their source AgentCore runtime ARN. Don't — &lt;code&gt;create_registry_record&lt;/code&gt; rejects a caller-supplied &lt;code&gt;provenance&lt;/code&gt; with &lt;code&gt;ValidationException: provenance cannot be set by the caller&lt;/code&gt;. It's populated only by the service's own auto-detection/sync integrations, not by a plain API call. Matching on record name is simpler anyway, since names are unique within a registry.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Searching for duplicates.&lt;/strong&gt; &lt;code&gt;search_discoverable_registry_records&lt;/code&gt; does hybrid semantic + keyword search, ordered by relevance:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;registry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agent-registry&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;registry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;search_discoverable_registry_records&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;searchQuery&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runtime_name&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;runtime_description&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;registryIds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;registry_arn&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;maxResults&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There's no numeric score in the response — just an ordered list. With a small registry (say, two or three approved records), that means it always returns &lt;em&gt;something&lt;/em&gt;, even for genuinely unrelated runtimes, because there's nothing better to rank against. This isn't a bug so much as a reminder that semantic search needs a reasonably sized corpus to actually discriminate. It's also a decent argument for keeping a human in the approval loop rather than auto-rejecting on any hit.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Writing the approval.&lt;/strong&gt; Three sequential calls: create, submit, approve.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;created&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create_registry_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;runtime_name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;recordType&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CUSTOM&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;recordVersion&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;1.0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;descriptors&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;custom&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;runtimeArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;runtime_arn&lt;/span&gt;&lt;span class="p"&gt;})}},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;record_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;created&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;recordArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/record/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;  &lt;span class="c1"&gt;# not returned directly
&lt;/span&gt;
&lt;span class="c1"&gt;# record starts CREATING and must reach DRAFT before you can submit it
&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;15&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get_registry_record&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;DRAFT&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;submit_registry_record_for_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;control&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update_registry_record_status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;registryId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;registry_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;recordId&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;record_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;APPROVED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;statusReason&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reviewed and approved&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Two things worth flagging: &lt;code&gt;CreateRegistryRecordResponse&lt;/code&gt; only returns &lt;code&gt;recordArn&lt;/code&gt; and &lt;code&gt;status&lt;/code&gt; — no &lt;code&gt;recordId&lt;/code&gt; field, so you extract it from the ARN — and creation is asynchronous, so a record submitted for approval before it leaves &lt;code&gt;CREATING&lt;/code&gt; will fail.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cross-referencing with AgentCore Runtime and CloudTrail
&lt;/h2&gt;

&lt;p&gt;The other half of "shadow agent" detection is enumerating what's actually running, independent of the registry:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;agentcore&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bedrock-agentcore-control&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;runtimes&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;agentcore&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;list_agent_runtimes&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agentRuntimes&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The IAM action for this is &lt;code&gt;bedrock-agentcore:ListAgentRuntimes&lt;/code&gt; — note the namespace is &lt;code&gt;bedrock-agentcore&lt;/code&gt;, not &lt;code&gt;bedrock-agentcore-control&lt;/code&gt; like the SDK package name would suggest. Getting this wrong produces a plain &lt;code&gt;AccessDeniedException&lt;/code&gt; with no hint about the namespace mismatch.&lt;/p&gt;

&lt;p&gt;For attribution — who deployed an unregistered runtime — CloudTrail's &lt;code&gt;LookupEvents&lt;/code&gt; over &lt;code&gt;CreateAgentRuntime&lt;/code&gt; gets you there, with the usual 90-day retention caveat:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;cloudtrail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;boto3&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;client&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cloudtrail&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;region_name&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;us-east-1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;cloudtrail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;lookup_events&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;LookupAttributes&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AttributeKey&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EventName&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AttributeValue&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CreateAgentRuntime&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Events&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]:&lt;/span&gt;
    &lt;span class="n"&gt;detail&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;CloudTrailEvent&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;])&lt;/span&gt;
    &lt;span class="n"&gt;arn&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;responseElements&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;agentRuntimeArn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;deployer&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;userIdentity&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;arn&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Infra notes
&lt;/h2&gt;

&lt;p&gt;Agent Registry has no CDK construct as of September 2026 — no &lt;code&gt;AWS::AgentRegistry::Registry&lt;/code&gt; CloudFormation resource type exists yet. I provisioned the registry and its seed records with the boto3 script above rather than CDK. AgentCore Runtime, by contrast, has a stable L2 construct (&lt;code&gt;aws_bedrockagentcore.Runtime&lt;/code&gt;), and &lt;code&gt;AgentRuntimeArtifact.fromCodeAsset()&lt;/code&gt; deploys straight from a local Python directory with no Docker step.&lt;/p&gt;

&lt;p&gt;If you're running this on Vercel with OIDC federation to AWS: the OIDC provider is scoped per Vercel &lt;em&gt;team&lt;/em&gt;, not per project. A second project under the same team hitting &lt;code&gt;new iam.OpenIdConnectProvider(...)&lt;/code&gt; will fail deployment, since an AWS account only accepts one provider per issuer URL. Import the existing one with &lt;code&gt;iam.OpenIdConnectProvider.fromOpenIdConnectProviderArn()&lt;/code&gt; instead of creating a second.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/oO9YQVkyDjY" width="710" height="399"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;h2&gt;
  
  
  Result
&lt;/h2&gt;

&lt;p&gt;Scanning an account with a couple of intentionally-similar and intentionally-unrelated AgentCore runtimes: the similar one surfaces a real "possible duplicate" match against an existing approved record, and approving the unrelated one produces a genuine &lt;code&gt;APPROVED&lt;/code&gt; record you can see with &lt;code&gt;list-registry-records&lt;/code&gt; afterward — not a mock, the actual governance workflow AWS Agent Registry ships.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>ai</category>
      <category>cloud</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>The Legal Wall I Hit Building a YouTube Clone on AWS</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Wed, 26 Aug 2026 16:31:29 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/the-legal-wall-i-hit-building-a-youtube-clone-on-aws-1ocf</link>
      <guid>https://dev.to/_76130e67067eab4c8510/the-legal-wall-i-hit-building-a-youtube-clone-on-aws-1ocf</guid>
      <description>&lt;p&gt;Right now I'm testing how far you can push a YouTube-style video platform built entirely on AWS. Upload, transcode, stream — EC2, S3, and CloudFront handle all of that without much drama. Then I started sketching out a comment section and a direct-message feature between users, and I stopped typing.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;Wait. Does this need a government license now?&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;That one question sent me down a rabbit hole through Japan's Telecommunications Business Act, Copyright Act, and a law with an even longer name that regulates online platforms. None of this is legal advice — I'm not a lawyer, just a developer who got nervous enough to read the actual statutes. If you're building something for real users at scale, talk to an actual lawyer. But if you're a solo developer wondering whether your side project is quietly illegal, this is the walkthrough I wish existed.&lt;/p&gt;

&lt;h2&gt;
  
  
  Japan's Telecommunications Business Act: the "are you a phone company" test
&lt;/h2&gt;

&lt;p&gt;Japan has a law called the &lt;a href="https://www.japaneselawtranslation.go.jp/en/laws/view/3648/en" rel="noopener noreferrer"&gt;Telecommunications Business Act&lt;/a&gt; (電気通信事業法). In the simplest possible terms: it's the law that decides whether your app is legally acting like a phone company, and if so, whether you need to tell the government about it.&lt;/p&gt;

&lt;p&gt;There are two tracks. Article 9 requires full registration, and it only kicks in if you own and operate your own transmission lines — think an actual telecom carrier laying fiber. Article 16 requires a lighter-weight notification, and it applies to almost everyone else who runs a communications service without owning that physical infrastructure. Since your app sits on AWS rather than your own fiber network, Article 16 (notification) is the one that could apply to you, not Article 9.&lt;/p&gt;

&lt;p&gt;Whether it actually applies comes down to one test: are you mediating communication between other people? A one-way video stream — you upload, others watch — is your own communication going out to viewers. It's not you relaying messages between two other people, so it generally falls outside the notification requirement. Add a DM feature, a live chat relay, or real-time comment broadcasting between users, though, and you start looking a lot more like a company that carries other people's messages, which is exactly what the law is watching for.&lt;/p&gt;

&lt;h2&gt;
  
  
  If you're the only one who can upload, you're basically fine
&lt;/h2&gt;

&lt;p&gt;Here's the scorecard for a video app where you're the only content creator — think a personal portfolio site that happens to look like YouTube.&lt;/p&gt;

&lt;p&gt;Telecommunications Business Act notification isn't required, because you're not relaying anyone else's communication. The Information Distribution Platform Act (more on that below) doesn't apply either, since there's no user-generated content for anyone to complain about. Japan's Act on the Protection of Personal Information technically applies the moment you add user accounts or store watch history, but for a small non-commercial project, a basic privacy policy covers you in practice.&lt;/p&gt;

&lt;p&gt;If this describes your project, you can build it, ship it, and move on — the same way you'd treat any other side project.&lt;/p&gt;

&lt;h2&gt;
  
  
  The moment strangers can upload, the rulebook gets thicker
&lt;/h2&gt;

&lt;p&gt;Turn on user uploads for everyone, and you've built a real user-generated-content platform. That changes things.&lt;/p&gt;

&lt;p&gt;The Telecommunications Business Act question gets murkier once you add DMs or chat, but realistically, the enforcement risk for a small side project is close to zero — call it a legal gray zone rather than a hard stop.&lt;/p&gt;

&lt;p&gt;The law that actually matters here is Japan's Information Distribution Platform Act (情報流通プラットフォーム対処法), which took effect on April 1, 2025. It's the successor to what used to be called the Provider Liability Limitation Act, and at its core it requires any platform hosting user content to have a way for people to report defamatory or copyright-infringing material and get it taken down. If you're running a UGC service, you need this regardless of size — but in practice, for a solo project, that requirement is satisfiable with a single email address people can send takedown requests to.&lt;/p&gt;

&lt;p&gt;The obligations scale up hard once you're huge. If your platform averages more than 10 million senders a month, or 20 million total, the government designates you a "large-scale specified telecommunications service provider" and you owe additional obligations like publishing your takedown response record. As of April 2025, the companies actually carrying that designation are Google, LY Corporation (LINE Yahoo), Meta, and TikTok. A solo developer isn't getting anywhere near that threshold, which is why the lightweight version — one inbox, checked occasionally — is a realistic bar to clear.&lt;/p&gt;

&lt;p&gt;There's also a quieter risk: if your platform lets anyone upload anything, someone eventually will upload something they don't own the rights to. If your code is open source on GitHub, that's a separate reputational problem worth thinking about even before the legal one.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbtd0ukvhaerfrjcwtpj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fwbtd0ukvhaerfrjcwtpj.png" alt=" " width="800" height="600"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The real trap was copyright law
&lt;/h2&gt;

&lt;p&gt;After all that, the thing that actually surprised me wasn't the telecom law or the platform law. It was copyright.&lt;/p&gt;

&lt;p&gt;Japan's &lt;a href="https://www.branche-ip.jp/2014/02/02/%E8%91%97%E4%BD%9C%E6%A8%A9%E6%B3%95%EF%BC%9A%E3%80%8C%E5%85%AC%E8%A1%86%E3%80%8D%E3%81%AB%E5%90%AB%E3%81%BE%E3%82%8C%E3%82%8B%E3%80%8C%E7%89%B9%E5%AE%9A%E3%81%8B%E3%81%A4%E5%A4%9A%E6%95%B0%E3%81%AE/" rel="noopener noreferrer"&gt;Copyright Act defines "the public"&lt;/a&gt; in a way that's broader than it sounds. Article 2, paragraph 5 states that "the public," for purposes of this law, includes "specific and numerous persons" — not just strangers off the street. In plain English: even if everyone who can see your content is someone you personally know and approved individually, if that group gets large enough, the law can treat it the same as posting it publicly. There's no hard headcount in the statute, but the widely cited rule of thumb among Japanese IP lawyers is that once you're past roughly 50 people, you're squarely in "many" territory, and past legal disputes have used numbers like 300+ as clearly qualifying.&lt;/p&gt;

&lt;p&gt;Then there's &lt;a href="https://note.com/copyrights/n/n2a269608a662" rel="noopener noreferrer"&gt;Article 23&lt;/a&gt;, covering the public transmission right. For anything automatically deliverable over a network — which covers basically all web and app content — this right also covers something Japanese law calls "making transmittable" (送信可能化). That means the right is triggered the moment content becomes available for someone to access, whether or not anyone has actually clicked play yet. "Nobody's actually watched it, so I'm fine" doesn't hold up as a legal argument under this framework.&lt;/p&gt;

&lt;p&gt;The counterweight is Article 30, the private-use exception. Using copyrighted material within a genuinely private circle — yourself, your family, people you live with — is exempt. Courts have described the boundary as "an extremely limited circle of personal relationships." A &lt;a href="https://www.businesslawyers.jp/articles/1247" rel="noopener noreferrer"&gt;2022 Supreme Court case&lt;/a&gt; involving JASRAC (Japan's music licensing body) and music schools is a good real-world illustration: the court found that a teacher performing music for students, one at a time, still counted as a "public" performance, because the audience rotated through an effectively unlimited stream of students over time. The lesson generalizes: an audience made of individually-approved people can still add up to "public" if it's large enough or churns enough.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxbmrf4fqyfk5da0bldz.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fdxbmrf4fqyfk5da0bldz.png" alt=" " width="800" height="440"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Passwords don't automatically make something private
&lt;/h2&gt;

&lt;p&gt;I went into this assuming that if I gated content behind a login and personally approved every viewer, I was safely inside the private-use exception. That turned out to be only half true.&lt;/p&gt;

&lt;p&gt;What decides the boundary isn't whether there's a password. It's who's on the other side of it, and how many of them there are. Family and people you live with land solidly inside the private-use exception — courts have consistently protected that "extremely limited circle." Once you're individually approving friends and acquaintances, the calculus shifts: the more people you add, and the less close the relationship, the more likely a court would call that group "specific and numerous" rather than private, regardless of whether you technically required a password to get in. A login screen controls access technically. It says nothing about whether the underlying use is legally private.&lt;/p&gt;

&lt;p&gt;The uncomfortable implication is that "approve anyone who asks" as a growth strategy — the design pattern, not any specific headcount — reads legally closer to "public" than "private," because the pool has no real ceiling.&lt;/p&gt;

&lt;p&gt;One important caveat, since a reader flagged this after the Japanese version of this post went semi-viral on X: all of this only matters if the content isn't fully yours to begin with. If you personally shot, wrote, and own every frame of what you're distributing, the public-transmission and private-use analysis above is close to irrelevant — copyright belongs to the creator, and you're free to distribute your own original work to as many people as you want. Where this actually bites is content that includes someone else's copyrighted material without permission: a recording that happens to pick up background music, a clip of broadcast TV, a screen recording that captures someone else's copyrighted app or footage. The moment third-party material is baked into what you're distributing, the "how many people, how close are you to them" analysis above is what determines your exposure.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add payments or age gates, and yet more laws show up
&lt;/h2&gt;

&lt;p&gt;Everything above covers just the core of a video platform. Add features, and you pick up more regulatory surface.&lt;/p&gt;

&lt;p&gt;The moment you're storing watch history or account data, Japan's Act on the Protection of Personal Information kicks in — a privacy policy that states your purpose of use is the baseline expectation. If minors might realistically use your service, the Act on Development of an Environment that Provides Safe and Secure Internet Use for Young People brings in expectations around age verification and filtering. And if you add tipping, subscriptions, or ad revenue sharing — anything that moves money — the Payment Services Act and the Act on Specified Commercial Transactions apply separately.&lt;/p&gt;

&lt;p&gt;For a personal, non-commercial project, this tier is "know it exists, revisit it if you add the feature" rather than something to solve up front.&lt;/p&gt;

&lt;h2&gt;
  
  
  So what does a solo developer actually need to do
&lt;/h2&gt;

&lt;p&gt;The pattern that emerged after all this reading: the real dividing line isn't "can other people upload," it's "is this actually public." A fully private deployment — access-controlled, URL not shared anywhere — falls inside the private-use exception with real room to spare. Even if a video you personally recorded happens to capture copyrighted material in the background, keeping it to private viewing is generally fine.&lt;/p&gt;

&lt;p&gt;The moment you deploy publicly, or post the URL on a blog where anyone can find it, you've crossed into public transmission, and the private-use exception no longer covers you — even if you're the only person who ever uploaded anything, placing unlicensed third-party material there can be infringing.&lt;/p&gt;

&lt;p&gt;Three things made this manageable for a solo project. First, keep access genuinely locked down — authentication plus a URL you don't publish — if you want to stay inside the private-use exception. Second, for anything you do deploy publicly or show in a GitHub README, use only material you made yourself or that's explicitly license-free; no copyrighted broadcast footage, no commercial music tracks. Third, add a line to your README or terms of service stating the app is intended for the creator's personal use and isn't designed for uploading third-party copyrighted material — it won't prevent a determined bad actor, but it's a reasonable statement of intent if the question ever comes up.&lt;/p&gt;

&lt;p&gt;Building something like automated audio fingerprinting to detect infringing uploads is overkill for a single-user, access-controlled personal project. Locking down access and being deliberate about demo content covers the realistic risk at this scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Closing thought
&lt;/h2&gt;

&lt;p&gt;I genuinely expected video streaming to be a solved, boring technical problem by now. I did not expect to spend this much time in statute text. The good news is that a personal, single-user version of this project needs essentially no legal paperwork. The part that actually changed my mental model was realizing that "I put a password on it" doesn't automatically mean "this is private" under Japanese copyright law — that one took a re-read to sink in.&lt;/p&gt;

&lt;p&gt;Still testing how far this goes on AWS. More to come if there's more to find.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>japan</category>
      <category>legal</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Autonomy for $20, a Human Above It: A Pattern for AI Agents That Spend Money</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Wed, 26 Aug 2026 08:27:21 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/autonomy-for-20-a-human-above-it-a-pattern-for-ai-agents-that-spend-money-425m</link>
      <guid>https://dev.to/_76130e67067eab4c8510/autonomy-for-20-a-human-above-it-a-pattern-for-ai-agents-that-spend-money-425m</guid>
      <description>&lt;p&gt;At some point every "agentic" product roadmap runs into the same uncomfortable question: what happens when the agent needs to actually spend money, not just recommend an action to a human who then clicks a button? Recommending is easy to sandbox. Spending isn't. And the usual answers — "just require approval for everything" or "just trust the agent" — both fail in an obvious way. Approve-everything means the agent isn't really autonomous, it's a slow suggestion box. Trust-the-agent means the first bad prompt injection or hallucinated tool call has a direct line to your bank balance.&lt;/p&gt;

&lt;p&gt;I wanted to see what a middle position actually looks like in code, not in a slide. So I built a small agent, sera-gate-agent, that pays multi-currency invoices for you and enforces a rule most people would agree with intuitively: small payments go through on their own, larger ones stop and wait for a person.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfx8zy289pkghmp24v1y.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Flfx8zy289pkghmp24v1y.webp" alt=" " width="800" height="798"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dbih1v1oav36058pmpa.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4dbih1v1oav36058pmpa.jpg" alt=" " width="800" height="629"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzrc7s4mvcm35ig9phx8w.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fzrc7s4mvcm35ig9phx8w.jpg" alt=" " width="800" height="745"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  The trigger: a protocol that treats a keypair as a login
&lt;/h3&gt;

&lt;p&gt;The settlement side runs on Sera, an on-chain FX protocol with an MCP server (sera-mcp) released under MIT specifically so agents can call it. What made me want to build on it wasn't the currency coverage, it was the auth model: there's no signup form. You generate a keypair, sign an EIP-712 message with it, and that signature is your API key request. A human finds this mildly inconvenient. An agent finds it exactly as convenient as an auth flow can be — no page to navigate, no CAPTCHA, no session cookie, just a key it already has.&lt;/p&gt;

&lt;p&gt;That detail is what made "give the agent a wallet" feel like a natural next step rather than a stunt. If the whole point of an agent is that it acts without a human driving each click, an auth system built around clicking a human through steps is already fighting the premise.&lt;/p&gt;

&lt;h3&gt;
  
  
  The design: one number, one gated function
&lt;/h3&gt;

&lt;p&gt;The rule I wanted was simple enough to say in one sentence: under $20, the agent executes on its own; over $20, it issues an approval card and a human has to click "approve" in a browser before anything moves. The part that took actual thought wasn't the rule, it was making sure the rule couldn't be talked around.&lt;/p&gt;

&lt;p&gt;The threshold check itself is almost embarrassingly small:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;decide_auto_or_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Decimal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;SERA_GATE_AUTO_THRESHOLD_USD&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;20&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AUTO&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="n"&gt;threshold&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;REQUIRE_APPROVAL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;What matters more is where it's called from. The underlying protocol exposes tools that actually move funds — execute a swap, send a transfer, pay an invoice directly. None of those are given to the language model. The only thing the model can call is a wrapper function, and that wrapper is the only code path that's allowed to reach the real execution tools. The threshold check lives inside that wrapper, not in a prompt instruction telling the model to "please ask before spending more than $20."&lt;/p&gt;

&lt;p&gt;That distinction is the whole point, and it's easy to gloss over. A system prompt is a request. A model that's very good at following instructions will follow it correctly almost all the time — but "almost all the time" is a bad security property for something that moves money, whether the failure mode is prompt injection, a weird edge case in tool output, or the model just deciding the instruction doesn't apply this time. Removing the tool from the model's reachable set removes the failure mode instead of making it rarer.&lt;/p&gt;

&lt;p&gt;I also didn't want the agent-side threshold to be the only backstop. Sera has its own server-side policy preset that caps what any single API key can move per transaction and per day, independent of anything my code does. I keep that ceiling well above my $20 agent-side threshold, so it's a true second layer rather than a duplicate of the first — if my code has a bug, or the signing key leaks, there's still a hard ceiling that doesn't route through my logic at all.&lt;/p&gt;

&lt;h3&gt;
  
  
  What this generalizes to
&lt;/h3&gt;

&lt;p&gt;None of this is specific to on-chain payments. The same shape applies to an agent that can issue refunds, send emails to customers, delete records, or push a deploy: pick a dimension the org already has intuitions about (dollar amount, blast radius, reversibility), pick a threshold, and make the boundary a property of which functions are reachable rather than a property of what the model has been told. The dollar threshold is just the easiest one to make legible in a demo, because everyone already has an intuition for what $20 versus $5,000 means.&lt;/p&gt;

&lt;p&gt;The uncomfortable part, and I don't think there's a clean answer to it, is that the threshold is still a judgment call, and it's a judgment call about how much you trust a system that doesn't get tired, doesn't get talked into things the way a person does, but also doesn't have the context a person has for "this specific invoice looks off." Setting it too low turns the agent back into a suggestion box. Setting it too high means the day it's wrong, it's wrong at a scale a human never got the chance to catch. I don't think that number should be static — the next version of this probably ties it to something like payee history or a running risk score instead of a flat constant — but getting the enforcement boundary right, structurally, felt like the part worth building first.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/Q2A2epOxGrY"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;sera-gate-agent runs on Sepolia testnet today, built on AWS Bedrock AgentCore Runtime and Strands, with a Next.js front end for the approval cards a human actually clicks through.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>blockchain</category>
      <category>aws</category>
    </item>
    <item>
      <title>Amazon Leo vs Starlink: The 2026 Satellite Internet Race</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 20 Aug 2026 18:09:22 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/amazon-leo-vs-starlink-the-2026-satellite-internet-race-598b</link>
      <guid>https://dev.to/_76130e67067eab4c8510/amazon-leo-vs-starlink-the-2026-satellite-internet-race-598b</guid>
      <description>&lt;p&gt;Project Kuiper quietly became Amazon Leo in November 2025. Enterprise beta opened on April 8, 2026. So the natural question: how close is Amazon actually getting to Starlink?&lt;/p&gt;

&lt;p&gt;Short answer: not close at all. And Amazon just missed a federal deadline in the process.&lt;/p&gt;

&lt;h2&gt;
  
  
  The numbers, side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Amazon Leo&lt;/th&gt;
&lt;th&gt;Starlink&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Satellites in orbit&lt;/td&gt;
&lt;td&gt;~345-400 (Aug 2026)&lt;/td&gt;
&lt;td&gt;~10,020 (Apr 2026)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subscribers&lt;/td&gt;
&lt;td&gt;Not yet public (enterprise beta only)&lt;/td&gt;
&lt;td&gt;10M+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Countries served&lt;/td&gt;
&lt;td&gt;5 targeted for mid/late 2026 launch&lt;/td&gt;
&lt;td&gt;150+&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consumer pricing&lt;/td&gt;
&lt;td&gt;Not yet announced&lt;/td&gt;
&lt;td&gt;$50-120/mo (Residential)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard hardware&lt;/td&gt;
&lt;td&gt;Not yet announced (aiming smaller/cheaper)&lt;/td&gt;
&lt;td&gt;$499 ($199 for Mini)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Launch method&lt;/td&gt;
&lt;td&gt;Atlas V / Falcon 9 / Ariane 6 (outsourced)&lt;/td&gt;
&lt;td&gt;Falcon 9 (in-house, reusable)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FCC status&lt;/td&gt;
&lt;td&gt;Missed the 50% deployment target, waiver carries a spectrum-priority penalty&lt;/td&gt;
&lt;td&gt;N/A&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The snapshot dates don't line up exactly (Starlink's count is from April, Amazon Leo's from August), but the order of magnitude tells the story either way.&lt;/p&gt;

&lt;h2&gt;
  
  
  The satellite count isn't even in the same order of magnitude
&lt;/h2&gt;

&lt;p&gt;As of April 2026, Starlink had &lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;10,020 satellites in orbit versus 241 for Amazon Leo&lt;/a&gt;. By August 2026 Amazon Leo had climbed to roughly &lt;a href="https://www.aboutamazon.com/news/innovation-at-amazon/project-kuiper-satellite-rocket-launch-progress-updates" rel="noopener noreferrer"&gt;345-400 satellites&lt;/a&gt; after a string of launches. Starlink is still ahead by more than 20x.&lt;/p&gt;

&lt;p&gt;Subscriber numbers tell the same story. Starlink serves &lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;over 10 million subscribers across 150+ countries&lt;/a&gt;. Amazon Leo isn't selling to consumers yet — it's still in enterprise beta, with &lt;a href="https://thenextweb.com/news/amazon-leo-satellite-internet-mid-2026" rel="noopener noreferrer"&gt;residential service targeted for mid-2026 in five countries&lt;/a&gt;: the US, Canada, the UK, France, and Germany.&lt;/p&gt;

&lt;h2&gt;
  
  
  Amazon actually missed its FCC deadline
&lt;/h2&gt;

&lt;p&gt;This is the part that surprised me most.&lt;/p&gt;

&lt;p&gt;Amazon's FCC authorization for its 3,236-satellite constellation came with a condition: launch 50% (1,616 satellites) by July 30, 2026, or risk losing priority status. Actual count at the deadline: &lt;a href="https://www.satellitetoday.com/connectivity/2026/06/05/fcc-gives-amazon-leo-50-deployment-waiver-with-conditions-on-spectrum-priority/" rel="noopener noreferrer"&gt;331 satellites&lt;/a&gt; — about 20% of target.&lt;/p&gt;

&lt;p&gt;The FCC granted a waiver, but not for free. Under the &lt;a href="https://www.geekwire.com/2026/fcc-gives-amazon-leo-more-leeway-on-its-satellite-deployment-schedule/" rel="noopener noreferrer"&gt;terms of the extension&lt;/a&gt;, any Gen1 satellite launched after July 30 temporarily loses the spectrum priority status Amazon earned in earlier FCC processing rounds, until March 30, 2028, or until it hits the 50% mark, whichever comes first. In a business where orbital slots and spectrum priority are genuinely scarce, that's a real cost, not just a paperwork inconvenience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the gap exists: nobody owns Amazon's rocket
&lt;/h2&gt;

&lt;p&gt;SpaceX runs Starlink and also builds the rocket that launches it. Falcon 9 is reusable, flies constantly, and SpaceX controls its own launch cadence end to end. That vertical integration is the actual root of Starlink's lead.&lt;/p&gt;

&lt;p&gt;Amazon Leo has no equivalent. It buys launches from &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;Atlas V, Falcon 9, and Ariane 6&lt;/a&gt;, and depends on other companies' schedules. Vulcan Centaur and Blue Origin's New Glenn are coming online too — Blue Origin being Bezos-founded but organizationally separate from Amazon — with a target of &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;20+ missions in 2026 and 30+ in 2027&lt;/a&gt;. Still nowhere near Starlink's cadence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Amazon Leo could still win
&lt;/h2&gt;

&lt;p&gt;The one card Starlink genuinely can't match: native AWS integration. A company already running workloads on AWS can get satellite backhaul into its own AWS region as part of one coherent stack. Starlink has no cloud platform to offer alongside it.&lt;/p&gt;

&lt;p&gt;The engineering also reflects a later start. Amazon Leo uses &lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;optical inter-satellite links and a custom baseband chip called Prometheus&lt;/a&gt;, and flies at a lower inclination (30-51°) that concentrates coverage on populated mid-latitudes rather than the wider polar coverage Starlink offers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Starlink, for reference
&lt;/h2&gt;

&lt;p&gt;Pricing as of 2026: &lt;a href="https://www.usmobile.com/blog/starlink-cost/" rel="noopener noreferrer"&gt;Residential $50-120/mo, Roam $50-165/mo, Mini $30/mo plus $199 hardware, Business from $250/mo&lt;/a&gt;. Standard hardware holds steady at $499.&lt;/p&gt;

&lt;p&gt;The more interesting move is T-Satellite, Starlink's direct-to-cell partnership with T-Mobile: &lt;a href="https://www.satelliteinternet.com/providers/starlink/starlink-direct-to-cell/" rel="noopener noreferrer"&gt;$10/month, free on higher-tier plans, works across 60+ phone models regardless of carrier&lt;/a&gt;. No military or first-responder discount currently exists on the core service.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;On raw numbers, this isn't a race yet — it's a head start plus a company still finding its footing. Amazon has committed &lt;a href="https://www.datacenterdynamics.com/en/news/amazon-promises-invest-more-10bn-project-kuiper-satellite-internet-business/" rel="noopener noreferrer"&gt;more than $10 billion&lt;/a&gt; to closing the gap, and the AWS integration angle is real. Whether it matters depends entirely on whether Amazon Leo can fix its launch cadence before the spectrum penalty clock runs out in March 2028.&lt;/p&gt;




&lt;p&gt;&lt;strong&gt;Sources&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;a href="https://5gstore.com/blog/2026/06/21/amazon-leo-starlink/" rel="noopener noreferrer"&gt;Amazon LEO Vs Starlink: Price, Speed, Latency, and Fit — 5Gstore&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.aboutamazon.com/news/innovation-at-amazon/project-kuiper-satellite-rocket-launch-progress-updates" rel="noopener noreferrer"&gt;Amazon Leo mission updates — About Amazon&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://thenextweb.com/news/amazon-leo-satellite-internet-mid-2026" rel="noopener noreferrer"&gt;Amazon Leo targets mid-2026 commercial launch — The Next Web&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.satellitetoday.com/connectivity/2026/06/05/fcc-gives-amazon-leo-50-deployment-waiver-with-conditions-on-spectrum-priority/" rel="noopener noreferrer"&gt;FCC Gives Amazon Leo 50% Deployment Waiver, With Conditions on Spectrum Priority — Via Satellite&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.geekwire.com/2026/fcc-gives-amazon-leo-more-leeway-on-its-satellite-deployment-schedule/" rel="noopener noreferrer"&gt;FCC gives Amazon Leo more leeway for deploying satellites — GeekWire&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://en.wikipedia.org/wiki/Amazon_Leo" rel="noopener noreferrer"&gt;Amazon Leo — Wikipedia&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.usmobile.com/blog/starlink-cost/" rel="noopener noreferrer"&gt;Starlink Plans &amp;amp; Pricing In 2026 — US Mobile&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.satelliteinternet.com/providers/starlink/starlink-direct-to-cell/" rel="noopener noreferrer"&gt;Starlink T-Satellite: Cost, Compatible Phones &amp;amp; Coverage — SatelliteInternet.com&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.navyweek.org/discount/starlink-military-discount/" rel="noopener noreferrer"&gt;Starlink Military Discount 2026&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.datacenterdynamics.com/en/news/amazon-promises-invest-more-10bn-project-kuiper-satellite-internet-business/" rel="noopener noreferrer"&gt;Amazon promises to invest more than $10bn in Project Kuiper — DCD&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;

</description>
      <category>aws</category>
      <category>space</category>
      <category>cloud</category>
      <category>satellite</category>
    </item>
    <item>
      <title>I Rebuilt YouTube on AWS Alone (and Hit Every Wall)</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Tue, 18 Aug 2026 15:54:36 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/i-rebuilt-youtube-on-aws-alone-and-hit-every-wall-3jh2</link>
      <guid>https://dev.to/_76130e67067eab4c8510/i-rebuilt-youtube-on-aws-alone-and-hit-every-wall-3jh2</guid>
      <description>&lt;p&gt;It started as a simple question: how is YouTube actually built? One thing led to another, and a few hours later I had a single-user, self-hosted video platform running in production on AWS — after redesigning the auth layer from scratch mid-build, chasing down an "exec format error," and discovering that avoiding a NAT Gateway didn't actually save me any money. Here's the whole story, including the parts that didn't work the first time.&lt;/p&gt;

&lt;h2&gt;
  
  
  What YouTube is actually made of
&lt;/h2&gt;

&lt;p&gt;Before writing any code, I wanted to understand what I was copying. Roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Upload and transcoding&lt;/strong&gt;: uploaded video gets converted into 144p through 4K/8K across multiple codecs, processed by a huge fleet of parallel workers&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CDN&lt;/strong&gt;: Google Global Cache — dedicated caching nodes placed directly inside ISP networks — plus adaptive bitrate streaming&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Metadata&lt;/strong&gt;: Vitess (a sharding layer over MySQL) and Bigtable/Spanner&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Recommendations&lt;/strong&gt;: a two-stage candidate generation + ranking ML system&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content ID&lt;/strong&gt;: audio/video fingerprinting to detect copyright infringement&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ads&lt;/strong&gt;: backed by Google Ad Manager&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9hxa0dg3t7n5gwqtvi8.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fl9hxa0dg3t7n5gwqtvi8.jpg" alt=" " width="800" height="439"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Can AWS alone reproduce it?
&lt;/h2&gt;

&lt;p&gt;Most of the functional skeleton maps cleanly onto managed AWS services: S3 for upload, MediaConvert for transcoding, CloudFront for delivery, DynamoDB for metadata, OpenSearch for search, Personalize for recommendations. That part is genuinely achievable.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhyzp4pgvp6qppz65qh7o.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fhyzp4pgvp6qppz65qh7o.jpg" alt=" " width="800" height="541"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;A few pieces aren't:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Content ID&lt;/strong&gt; has no AWS-managed equivalent. You'd need a third-party SaaS like Audible Magic or ACRCloud, or roll your own fingerprinting with something like Chromaprint against a reference database you'd also have to build&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ISP-embedded caching&lt;/strong&gt; — CloudFront has a global edge network, but nothing at the density of nodes sitting inside individual ISPs&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ad auctions&lt;/strong&gt; at Google Ad Manager's scale aren't something you build yourself; you'd hand this off to an existing ad network&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a personal project, I decided not to build Content ID or ad serving at all. That decision turned out to be tied directly to a legal question I hadn't expected to spend time on.&lt;/p&gt;

&lt;h2&gt;
  
  
  The legal research I didn't expect to do
&lt;/h2&gt;

&lt;p&gt;Building a video-sharing app in Japan, even a personal one, touches a surprising number of regulations, so I checked before writing any infrastructure code.&lt;/p&gt;

&lt;p&gt;First, the &lt;strong&gt;Telecommunications Business Act&lt;/strong&gt;. One-way video distribution generally doesn't require registration, since you're not "mediating someone else's communication." Add a comment section or DMs between users, though, and that changes.&lt;/p&gt;

&lt;p&gt;Second, the &lt;strong&gt;Act on the Limitation of Liability for Damages of Specified Telecommunications Service Providers&lt;/strong&gt; (Japan's provider-liability law, recently renamed to something closer to "platform accountability act"). Any platform accepting user-generated content is expected to run a takedown-request contact point; cross a large-user threshold (10M+ monthly users in Japan) and heavier obligations kick in.&lt;/p&gt;

&lt;p&gt;Third — and this is the one that actually shaped the design — &lt;strong&gt;Article 30 of the Copyright Act&lt;/strong&gt;, the private-use reproduction exception. Keep something fully private, accessible only to yourself, and it falls under private use. Make it public and it becomes "transmission to the public" (公衆送信), where that exception no longer applies. I also checked whether gating access behind a login would be enough to stay private if I let a few people in. It isn't automatically: under Japanese copyright law, "the public" includes "a specific but numerous group," so the real question isn't whether there's a login screen, it's &lt;em&gt;how many people, and how close a relationship&lt;/em&gt;. Family-sized is safe; a wider circle of friends risks crossing into "specific but numerous."&lt;/p&gt;

&lt;p&gt;Given all that, I decided the app would support exactly one user — me. No sign-up, no invite flow. That sidesteps the Telecommunications Business Act and the platform-liability law entirely, and keeps everything inside the private-use exception.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plan
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Single user only, no sign-up&lt;/li&gt;
&lt;li&gt;AWS only (I use Vercel for most other projects, but not this one)&lt;/li&gt;
&lt;li&gt;No Content ID, no ad serving&lt;/li&gt;
&lt;li&gt;Upload → S3 → MediaConvert (transcode to HLS) → CloudFront&lt;/li&gt;
&lt;li&gt;Web UI on ECS Fargate + ALB + CloudFront (App Runner was already off the table — AWS stopped accepting new App Runner services)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I wrote the CDK for VPC, S3, DynamoDB, a Lambda to kick off MediaConvert jobs, ECS, CloudFront, and Cognito. So far, so normal.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cognito's ALB integration needs HTTPS, and I didn't have a domain
&lt;/h2&gt;

&lt;p&gt;My first pass at auth used the ALB's native &lt;code&gt;authenticate-cognito&lt;/code&gt; listener action — no app code needed, ALB handles the redirect to Cognito's hosted UI for you. Clean, until deploy:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Resource handler returned message: "Actions of type 'authenticate-cognito' are supported only on HTTPS listeners"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That action only works on HTTPS listeners, which means an ACM certificate, which means a real, DNS-verifiable domain — something this project didn't have. Buying a domain just for this felt like the wrong trade, so I moved authentication into the app itself instead, using Next.js's &lt;code&gt;proxy.ts&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Moving auth to the app ran straight into "no NAT Gateway"
&lt;/h2&gt;

&lt;p&gt;I reconfigured the Cognito App Client as a public client (no secret) and switched to PKCE for the authorization code exchange, so the browser could talk to Cognito directly instead of routing through the ALB's constraints.&lt;/p&gt;

&lt;p&gt;Redeployed, logged in, and got a 500 on the callback:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;⨯ [TypeError: fetch failed] {

      at ignore-listed frames {
    code: 'ETIMEDOUT',
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The ECS task had no path to the internet. To keep costs down I'd built the VPC with interface endpoints instead of a NAT Gateway — but there's no VPC endpoint for Cognito's Hosted UI/OAuth domain (&lt;code&gt;*.auth.&amp;lt;region&amp;gt;.amazoncognito.com&lt;/code&gt;). The container simply couldn't reach it.&lt;/p&gt;

&lt;p&gt;The fix was to move the token exchange itself into the browser. The only thing that actually needs to happen server-side is JWT verification (fetching the JWKS), which &lt;em&gt;is&lt;/em&gt; covered by the &lt;code&gt;cognito-idp&lt;/code&gt; VPC endpoint. The browser already has internet access, so it can talk to Cognito's token endpoint directly. That change got login working without ever adding a NAT Gateway.&lt;/p&gt;

&lt;h2&gt;
  
  
  Forgot to pin the CPU architecture, container wouldn't start
&lt;/h2&gt;

&lt;p&gt;Redeployed again, and this time the ECS task crash-looped indefinitely. CloudWatch Logs had exactly one line to offer:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;exec /usr/local/bin/docker-entrypoint.sh: exec format error
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Building the Docker image on an Apple Silicon Mac produces an arm64 image. Fargate defaults to x86_64. Nothing about the mismatch surfaces until the container tries to actually execute. Setting &lt;code&gt;runtimePlatform&lt;/code&gt; to ARM64 on the &lt;code&gt;FargateTaskDefinition&lt;/code&gt; fixed it. I also turned on the ECS deployment circuit breaker at the same time — without it, a failing deployment can take up to three hours to be reported as failed, and I'd already lost about 40 minutes not noticing.&lt;/p&gt;

&lt;h2&gt;
  
  
  The health check was hitting the login redirect
&lt;/h2&gt;

&lt;p&gt;Next failure: the ALB health check was pointed at &lt;code&gt;/&lt;/code&gt;, which — like every other route — goes through the app's auth gate. An unauthenticated health check gets a 302, the ALB reads that as unhealthy, and the deployment fails outright. Added a dedicated &lt;code&gt;/api/health&lt;/code&gt; route that skips the auth check, and that was that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Video played, but the manifest path was broken
&lt;/h2&gt;

&lt;p&gt;Deployment finally succeeded, upload worked, MediaConvert finished the job — and the video was just a black rectangle.&lt;/p&gt;

&lt;p&gt;The cause: MediaConvert's completion event returns &lt;code&gt;outputGroupDetails.playlistFilePaths&lt;/code&gt; as a full &lt;code&gt;s3://bucket/key&lt;/code&gt; URI, not a bucket-relative key. I'd been storing that value directly as &lt;code&gt;manifestKey&lt;/code&gt;, so the app's &lt;code&gt;/${manifestKey}&lt;/code&gt; template produced a broken &lt;code&gt;/s3://bucket/...&lt;/code&gt; path. Since I already control the output prefix at job-creation time, I switched to deriving the key deterministically instead of trusting the event payload. Don't take an AWS event field at face value if you can compute the same thing yourself.&lt;/p&gt;

&lt;h2&gt;
  
  
  CloudFront's signed cookies ignored my wildcard
&lt;/h2&gt;

&lt;p&gt;To lock down &lt;code&gt;/renditions/*&lt;/code&gt; (the actual video files) behind CloudFront's Key Group, I used &lt;code&gt;@aws-sdk/cloudfront-signer&lt;/code&gt;'s &lt;code&gt;getSignedCookies&lt;/code&gt; with a &lt;code&gt;url&lt;/code&gt; + &lt;code&gt;dateLessThan&lt;/code&gt; — the "canned policy" form — expecting a wildcard path to cover everything under it. It didn't:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nl"&gt;"error"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"AccessDenied"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="nl"&gt;"message"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="s2"&gt;"Access denied"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;(Along the way I also discovered this AWS account already had a CloudFront Public Key from a different project, and my first debugging attempt had grabbed the wrong Key Pair ID entirely — worth checking &lt;code&gt;aws cloudfront list-public-keys&lt;/code&gt; before assuming there's only one.)&lt;/p&gt;

&lt;p&gt;The actual fix was switching to an explicit custom policy — passing &lt;code&gt;policy&lt;/code&gt; with a JSON statement whose &lt;code&gt;Resource&lt;/code&gt; includes the wildcard — rather than the canned &lt;code&gt;url&lt;/code&gt;/&lt;code&gt;dateLessThan&lt;/code&gt; shortcut. The SDK happily accepts a wildcard in the canned form; CloudFront just doesn't honor it the same way.&lt;/p&gt;

&lt;h2&gt;
  
  
  Deleting a video brought it back from the dead
&lt;/h2&gt;

&lt;p&gt;With everything working, I added delete. It's supposed to be a simple DynamoDB + S3 cleanup, but it hit two separate bugs.&lt;/p&gt;

&lt;p&gt;First, IAM: &lt;code&gt;grantWrite&lt;/code&gt;/&lt;code&gt;grantDelete&lt;/code&gt; only cover object-level actions (&lt;code&gt;s3:PutObject*&lt;/code&gt;, &lt;code&gt;s3:DeleteObject*&lt;/code&gt;), not the bucket-level &lt;code&gt;s3:ListBucket&lt;/code&gt; that &lt;code&gt;ListObjectsV2&lt;/code&gt; needs during cleanup. Adding &lt;code&gt;grantRead&lt;/code&gt; fixed it.&lt;/p&gt;

&lt;p&gt;Second, and more interesting: deleting a video that was still processing let the MediaConvert-completion Lambda fire &lt;em&gt;after&lt;/em&gt; deletion, calling &lt;code&gt;UpdateItem&lt;/code&gt; on a videoId that no longer existed. DynamoDB's &lt;code&gt;UpdateItem&lt;/code&gt; creates the item if it's missing — so the "deleted" video would silently reappear, partially populated. Adding &lt;code&gt;ConditionExpression: 'attribute_exists(videoId)'&lt;/code&gt; made that update a no-op instead of a resurrection.&lt;/p&gt;

&lt;h2&gt;
  
  
  The security group I "locked down" wasn't actually locked down
&lt;/h2&gt;

&lt;p&gt;I wanted the ALB reachable only through CloudFront, so I restricted its security group to CloudFront's managed prefix list (&lt;code&gt;pl-58a04531&lt;/code&gt;). Deployed, checked the actual rule set, and found this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"IpRanges"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"CidrIp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"0.0.0.0/0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"Description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Allow from anyone on port 80"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"PrefixListIds"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="nl"&gt;"PrefixListId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pl-58a04531"&lt;/span&gt;&lt;span class="p"&gt;}]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both rules were live at once. The culprit was the ALB listener's &lt;code&gt;open&lt;/code&gt; property, which defaults to &lt;code&gt;true&lt;/code&gt; and silently adds its own 0.0.0.0/0 ingress rule regardless of what you've configured on the security group yourself. Setting &lt;code&gt;open: false&lt;/code&gt; on &lt;code&gt;addListener&lt;/code&gt; removed it. This is the kind of gap you only catch by actually reading the deployed state back from the AWS CLI — the CDK code alone looked correct.&lt;/p&gt;

&lt;h2&gt;
  
  
  Avoiding a NAT Gateway didn't actually save money
&lt;/h2&gt;

&lt;p&gt;Once things were stable, I priced out the fixed monthly cost:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Monthly (approx.)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;5x VPC interface endpoints&lt;/td&gt;
&lt;td&gt;~$50.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ALB&lt;/td&gt;
&lt;td&gt;~$20–23&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ECS Fargate (0.25 vCPU / 0.5GB, ARM64)&lt;/td&gt;
&lt;td&gt;~$8.90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secrets Manager&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;~$83&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I'd built five interface endpoints specifically to avoid a NAT Gateway (roughly $44.60/month in Tokyo, plus data processing). Adding them up, the endpoints cost about the same as the NAT Gateway would have — sometimes more. Consolidating to a single NAT Gateway would save maybe $5–6/month at the cost of a single point of failure, which is a fine trade for a personal, single-user app. I ended up leaving the endpoint-based setup as-is; the savings weren't worth the churn.&lt;/p&gt;

&lt;h2&gt;
  
  
  What's actually running
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Login via Cognito with PKCE, one user account, no sign-up flow&lt;/li&gt;
&lt;li&gt;Upload → S3 → Lambda → MediaConvert → HLS&lt;/li&gt;
&lt;li&gt;CloudFront with signed cookies gating the video files themselves&lt;/li&gt;
&lt;li&gt;Delete, a Japanese/English toggle, and a dark, YouTube-ish UI&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The code is &lt;a href="https://github.com/yama3133/mytube" rel="noopener noreferrer"&gt;public on GitHub&lt;/a&gt;, including the README section explaining, in plain terms, why multi-user upload was never on the table.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furx449h8w70unbha5zgo.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Furx449h8w70unbha5zgo.png" alt=" " width="" height=""&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu7f31m46m7fbhryjy73.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fiu7f31m46m7fbhryjy73.png" alt=" " width="800" height="487"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Mobile&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0a82kh7zvh38rt1imw3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fx0a82kh7zvh38rt1imw3.png" alt=" " width="624" height="1514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2d4h4iewczjvj0u9065.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fg2d4h4iewczjvj0u9065.png" alt=" " width="634" height="1514"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that actually took the time
&lt;/h2&gt;

&lt;p&gt;None of the individual fixes here were hard once I knew what was wrong. What took the time was reading logs — &lt;code&gt;exec format error&lt;/code&gt;, &lt;code&gt;AccessDenied&lt;/code&gt;, &lt;code&gt;InvalidKey&lt;/code&gt; — each one terse, each one caused by something completely different. Past a certain point, building on managed AWS services stops being about writing code and starts being about getting fast at figuring out why something &lt;em&gt;isn't&lt;/em&gt; working.&lt;/p&gt;

</description>
      <category>aws</category>
      <category>cdk</category>
      <category>nextjs</category>
      <category>cognito</category>
    </item>
    <item>
      <title>AWS Instance Store: Built to Disappear, On Purpose</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Fri, 14 Aug 2026 05:32:58 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/aws-instance-store-built-to-disappear-on-purpose-59p7</link>
      <guid>https://dev.to/_76130e67067eab4c8510/aws-instance-store-built-to-disappear-on-purpose-59p7</guid>
      <description>&lt;p&gt;Instance Store gets introduced in almost every AWS storage comparison the same way: "NVMe SSD physically attached to the host, faster than EBS, but the data disappears when the instance stops." That last clause usually reads like a warning label. Most guides then walk you straight into the safe, well-worn use cases — Cassandra nodes that don't mind losing a replica, Spark shuffle space, a scratch disk for sorting temp files. All correct, all a little boring.&lt;/p&gt;

&lt;p&gt;What if the disappearing part isn't the catch, but the whole point?&lt;/p&gt;

&lt;p&gt;Two workloads make that case surprisingly well: blockchain nodes doing a fast state sync, and compute jobs that touch data you'd rather not still have lying around tomorrow. They don't look related at first. One is about speed, the other about disappearance. But they're actually the same trick told twice — you get a disk that works blazingly fast for a short, defined burst, and then erases itself as a side effect of you being done with it. You're not fighting the ephemerality. You're renting it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What instance store actually promises&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Worth being precise here, because the security argument later depends on it. Instance store data does not persist through a stop, a terminate, a hibernate, or an underlying host failure. It does persist through a plain reboot, since that keeps you on the same physical host. AWS documents that the storage is not accessible to whoever gets the host next, which is the property that makes both ideas below work at all.&lt;/p&gt;

&lt;p&gt;What instance store is not: a certified secure-erase mechanism. If your compliance framework requires a documented, auditable wipe procedure — HIPAA, PCI-DSS, that kind of thing — "the disk went away when I stopped the instance" is not a control you can point an auditor at. Keep that distinction in mind as you read the rest of this, because the two ideas here are architectural thought experiments, not a substitute for whatever your compliance team actually needs signed off.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Idea one: the node that only exists to catch up&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Syncing a blockchain node from genesis, or even from a recent snapshot, is an I/O-bound slog. You're writing and reading state data continuously for hours, sometimes days, and once the node is caught up, most of that historical grind stops mattering — what you actually want going forward is a warm, synced node.&lt;/p&gt;

&lt;p&gt;The usual move is to provision an instance with a big EBS volume, let it sync, and keep paying for that volume indefinitely. Instance store flips the framing: treat the sync itself as the disposable part. Spin up an instance with local NVMe, let it rip through the sync at NVMe speeds instead of network-attached-storage speeds, and once it's caught up, snapshot the resulting state to S3 or EBS. The instance that did the syncing was never meant to be the long-term home for that data — it was a sprinter, not a warehouse. If it dies mid-sync, you weren't attached to it anyway; you just launch another one and let it catch up again, ideally from a recent checkpoint instead of genesis.&lt;/p&gt;

&lt;p&gt;This is basically the render-farm mentality applied to sync jobs: the compute is consumable, the output is what you keep. It also pairs naturally with Spot — losing a spot instance mid-sync is annoying, not catastrophic, precisely because you never treated its disk as the source of truth.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcqcmefzpdh3061u57ao.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fvcqcmefzpdh3061u57ao.jpg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;br&gt;
&lt;strong&gt;Idea two: compute that isn't supposed to remember anything&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Now the other direction. Imagine a batch job that has to touch something sensitive for a few minutes — decrypting a payload, running a one-off transformation on data you were only ever supposed to process, not retain. The usual anxiety with EBS-backed compute is the tail: did the volume get deleted on termination, did a snapshot get left behind by accident, is there a stray AMI somewhere with that data baked in.&lt;/p&gt;

&lt;p&gt;Instance store sidesteps most of that tail by construction. Launch the instance, do the job, terminate it. There's no volume to remember to delete, because there was never a persistent volume to begin with. The "forgetting" isn't a cleanup step you have to remember to run — it's what happens automatically when the job's done and you walk away. It's less "secure deletion" and more "the environment was never built to have a memory."&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2cv86pqeo4eu3op7r7e.jpg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fb2cv86pqeo4eu3op7r7e.jpg" alt=" " width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I'll be upfront that this is the idea I'd stress-test hardest before trusting it with anything actually regulated. It's a genuinely nice property for internal tooling, dev/test data that's sensitive but not audited, or a proof of concept where "the disk goes away" is a reasonable enough story. It is not, on its own, a story you'd want to tell a security auditor for anything under a real compliance regime — that needs KMS-backed encryption, documented key destruction, and probably a paper trail instance store just doesn't produce.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;The thread connecting them&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Both ideas lean on the same underlying shift: stop treating the disk's short lifespan as a constraint to architect around, and start treating it as the reason the architecture works. The blockchain sync node is fast because nobody's paying the tax of durable storage during the grind. The confidential job is simple because nobody has to remember to clean up after it. In both cases the disappearing act isn't a workaround — it's doing actual work.&lt;/p&gt;

&lt;p&gt;None of this replaces the orthodox use cases. Cassandra nodes, EMR clusters, and CI runners are still the bread and butter of instance store, and for good reason — they're proven, well-documented, and nobody's going to ask you hard questions about why you picked them. But it's worth remembering that "the data goes away" is a spec, not a bug report, and specs can be designed around on purpose. Sometimes the most interesting infrastructure decision is picking the tool that forgets on schedule, and building the rest of the system to expect exactly that.&lt;/p&gt;

</description>
      <category>instancestore</category>
      <category>ec2</category>
    </item>
    <item>
      <title>How to Pause an AI Agent for Human Approval Without a WebSocket</title>
      <dc:creator>Yuuki Yamashita</dc:creator>
      <pubDate>Thu, 13 Aug 2026 15:54:18 +0000</pubDate>
      <link>https://dev.to/_76130e67067eab4c8510/how-to-pause-an-ai-agent-for-human-approval-without-a-websocket-19cm</link>
      <guid>https://dev.to/_76130e67067eab4c8510/how-to-pause-an-ai-agent-for-human-approval-without-a-websocket-19cm</guid>
      <description>&lt;p&gt;If an AI agent needs a human to approve something mid-task, the instinct is usually to reach for a websocket, a message queue, or some kind of push notification service to bridge the backend and the frontend. I ended up not needing any of that. One DynamoDB row, polled from both sides, does the whole job. I built this for &lt;a href="https://github.com/yama3133/sub-sentry" rel="noopener noreferrer"&gt;SubSentry&lt;/a&gt;, an agent for AWS's Agents for Humans Hackathon that renews clean subscriptions on its own and asks a human before touching anything that looks like a price hike, a duplicate charge, or an unrecognized merchant. This post is about the mechanism underneath that "asking," not the subscription-tracking part.&lt;/p&gt;

&lt;h2&gt;
  
  
  The shape of the problem
&lt;/h2&gt;

&lt;p&gt;An agent tool call that needs human approval has to do two contradictory things at once. It has to actually block, because the agent's next step depends on the answer, and it has to somehow let something completely separate (a browser tab, a Slack bot, a CLI) deliver that answer whenever a human gets around to it, which could be five seconds or five minutes later. The backend process and the thing collecting the human's decision don't share memory, don't share a request, and in my case run on entirely different platforms (Bedrock AgentCore Runtime for the agent, Vercel serverless functions for the UI).&lt;/p&gt;

&lt;p&gt;The trick is to stop thinking of it as backend-talks-to-frontend at all. Neither side needs to know the other exists. They both just need to agree on one row.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbbsm46clw88s8ir3tr.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fetbbsm46clw88s8ir3tr.png" alt=" " width="800" height="834"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The row as a mailbox
&lt;/h2&gt;

&lt;p&gt;Here's the actual store, trimmed slightly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;request_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;suggested_action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ttl_seconds&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="n"&gt;approval_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;uuid&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uuid4&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
    &lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;approval_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;subscription_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;subscription_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PENDING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;suggested_action&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;suggested_action&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;reasons&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;reasons&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;amount_usd&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;amount_usd&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;created_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;now&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;ttl_seconds&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;decision&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nf"&gt;_dynamo_put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;


&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;wait_for_decision&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;poll_sec&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;1.0&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;entry&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;get_approval&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;PENDING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;time&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;float&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;expires_at&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]):&lt;/span&gt;
            &lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;EXPIRED&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="nf"&gt;_dynamo_put&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;entry&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
        &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;poll_sec&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent's tool calls &lt;code&gt;request_approval&lt;/code&gt;, gets back an &lt;code&gt;approval_id&lt;/code&gt;, and immediately calls &lt;code&gt;wait_for_decision&lt;/code&gt; on it, which just sits there polling DynamoDB once a second. That's the entire "block" side. It's a plain Python &lt;code&gt;while True&lt;/code&gt; loop, nothing fancier, because AgentCore Runtime is already paying for a long-running invocation, so there's no reason to make the waiting clever.&lt;/p&gt;

&lt;h2&gt;
  
  
  The other side never has to know it's being waited on
&lt;/h2&gt;

&lt;p&gt;The frontend's job is smaller than it sounds: read rows where &lt;code&gt;status = PENDING&lt;/code&gt;, render them as cards, and when a human clicks Approve or Reject, write the decision back. Here's the write, as a DynamoDB &lt;code&gt;UpdateItem&lt;/code&gt; call from a Next.js API route:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;ddb&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;UpdateCommand&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
    &lt;span class="na"&gt;TableName&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;TABLES&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;approvals&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;Key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;UpdateExpression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SET #s = :d, decision = :d, #r = :r, decided_at = :t&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ExpressionAttributeNames&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#s&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;status&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;#r&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;reason&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;ExpressionAttributeValues&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:d&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;decision&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:r&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;reason&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:t&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;:pending&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PENDING&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;ConditionExpression&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;attribute_exists(approval_id) AND #s = :pending&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;ReturnValues&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ALL_NEW&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;})&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;ConditionExpression&lt;/code&gt; is doing more work than it looks like. It means two people can't both approve the same card and have it silently double-apply, and it means a decision can't land on a row that already expired. If the condition fails, DynamoDB throws &lt;code&gt;ConditionalCheckFailedException&lt;/code&gt;, which the route turns into a 409. No locking, no transactions, just a condition on a single-item write.&lt;/p&gt;

&lt;p&gt;And that's the whole contract. The agent doesn't call an API on the frontend. The frontend doesn't call an API on the agent. A CLI can write the same &lt;code&gt;UpdateItem&lt;/code&gt; and it works identically, which is why &lt;code&gt;agent.py approve &amp;lt;id&amp;gt;&lt;/code&gt; from a terminal resolves the exact same pending card as clicking Approve in the browser. Neither side was written with the other in mind, they just both read and write the same table with the same status field.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this breaks if you're not careful
&lt;/h2&gt;

&lt;p&gt;The failure mode I actually hit wasn't in this mechanism, it was one level down. Strands dispatches multiple tool calls from the same agent turn concurrently, and my first pass at local storage (before I had a real DynamoDB table wired up) was a plain read-JSON-modify-write with no locking. Two tool calls landing at nearly the same instant would both read the file, both append their own entry in memory, and whichever one wrote last won, silently dropping the other's write. I found it because a local &lt;code&gt;.approvals.json&lt;/code&gt; file had two writes visibly tangled together mid-file, not because a test failed cleanly.&lt;/p&gt;

&lt;p&gt;The fix was five lines, a &lt;code&gt;threading.Lock&lt;/code&gt; around the read-modify-write section:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;_LOCK&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;_local_load&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="n"&gt;approval_id&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;entry&lt;/span&gt;
    &lt;span class="nf"&gt;_local_save&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;DynamoDB's per-item &lt;code&gt;UpdateItem&lt;/code&gt; doesn't have this problem at all, since each write targets one item atomically. The bug only existed because my local dev fallback was reinventing a worse version of what DynamoDB gives you for free. Worth remembering next time a "just write it to a JSON file for now" shortcut feels harmless.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just use a websocket
&lt;/h2&gt;

&lt;p&gt;I did consider it, mostly out of habit. But a websocket needs a persistent connection on both ends, which means something has to stay alive to hold it, and AgentCore Runtime invocations and Vercel serverless functions are both built around not staying alive longer than they have to. Polling a table every one to three seconds costs nothing worth optimizing at this scale, and it means either side of the system can restart, redeploy, or die completely mid-wait and the other side won't even notice, because the DynamoDB row is the only thing that has to survive.&lt;/p&gt;

&lt;p&gt;  &lt;iframe src="https://www.youtube.com/embed/WkiH9TUdEYw"&gt;
  &lt;/iframe&gt;
&lt;/p&gt;

&lt;p&gt;If you want to see the whole thing running, live demo and code are here:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Live demo: &lt;a href="https://sub-sentry-xi.vercel.app" rel="noopener noreferrer"&gt;https://sub-sentry-xi.vercel.app&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Code: &lt;a href="https://github.com/yama3133/sub-sentry" rel="noopener noreferrer"&gt;https://github.com/yama3133/sub-sentry&lt;/a&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Built for AWS's Agents for Humans Hackathon. #AgentsforHumans&lt;/p&gt;

</description>
      <category>aws</category>
      <category>agentsforhumans</category>
      <category>dynamodb</category>
      <category>architecture</category>
    </item>
  </channel>
</rss>
