<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Daniel Nwaneri</title>
    <description>The latest articles on DEV Community by Daniel Nwaneri (@dannwaneri).</description>
    <link>https://dev.to/dannwaneri</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3606168%2F7684e1e1-b986-4ee3-ae5b-56db2b97d286.jpg</url>
      <title>DEV Community: Daniel Nwaneri</title>
      <link>https://dev.to/dannwaneri</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dannwaneri"/>
    <language>en</language>
    <item>
      <title>i used to think in code. now i think in prompts.</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 12 Aug 2026 07:57:25 +0000</pubDate>
      <link>https://dev.to/dannwaneri/i-used-to-think-in-code-now-i-think-in-prompts-6h5</link>
      <guid>https://dev.to/dannwaneri/i-used-to-think-in-code-now-i-think-in-prompts-6h5</guid>
      <description>&lt;p&gt;used to have a habit of thinking in code. it'd be like walking down the street and immediately seeing what kind of code would make the next paragraph.&lt;/p&gt;

&lt;p&gt;now it's thinking in prompts. half a recipe for making pasta is in the middle of my brain. i just need to make the prompt to add that.&lt;/p&gt;

&lt;p&gt;i learned to code when i was a geophysicist with no computer science education. i was forced to read documentation. i built a lot. it took me years to realize that everyone who had a similar skill set was doing things that made it difficult to be effective. the world was full of idiots. now there are fewer idiots because of an LLM which knows everything. i'm pretty sure i wouldn't have thought it necessary to write code in the first place if someone could do it for me.&lt;/p&gt;

&lt;p&gt;no one talks about how to preserve the skills that will make future generations smarter. we're getting worse at thinking deep and dealing with humans because our AI tools don't encourage it.&lt;/p&gt;

&lt;p&gt;it's really shameful to think about the fact that now it's more efficient to think about things in terms of "prompt" rather than code.&lt;/p&gt;

&lt;p&gt;but this isn't an anti-LLM post. i use it all the time to save time on writing. i'm optimistic that it'll help me be a better programmer. and i've been doing three hackathons at the same time, and i couldn't have done it using only one brain at 100% usage capacity. it makes life so much easier.&lt;/p&gt;

&lt;p&gt;being behind is always an unpleasant feeling. i'm on the edge of a bug in the system:&lt;/p&gt;

&lt;p&gt;it tells me what the bug is, and suggests a fix.&lt;/p&gt;

&lt;p&gt;"what does it look like?" i ask.&lt;/p&gt;

&lt;p&gt;the answer has too many technical details. most of them are correct, but i'm too far behind on context.&lt;/p&gt;

&lt;p&gt;so i'll let it take care of it. it fixes the bug, commits the changes, and starts a pull request.&lt;/p&gt;

&lt;p&gt;"wait a minute… is there a reason for this commit?"&lt;/p&gt;

&lt;p&gt;"you're asking me a question before i even got started!"&lt;/p&gt;

&lt;p&gt;"this isn't good. i'll check the git log…"&lt;/p&gt;

&lt;p&gt;this sounds like an opportunity to build a mental model of what's happening. this is a perfect situation to be an agentless developer. but i can be busy with the next 10 threads. i'd lose a lot of time. maybe that's okay? let's see what the changes look like…&lt;/p&gt;

&lt;p&gt;the changes aren't great. it's okay so far. wait a minute… i should double check this change…&lt;/p&gt;

&lt;p&gt;oh. okay.&lt;/p&gt;

&lt;p&gt;next thread.&lt;/p&gt;

&lt;p&gt;this is how we get to where we are today…&lt;/p&gt;

&lt;p&gt;i don't expect to switch back from thinking in prompts… i'm optimistic about LLMs! they are helping a lot of people do incredible things! i'm embarrassed that i'm still not doing some things that i can easily delegate with prompts.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>discuss</category>
      <category>career</category>
    </item>
    <item>
      <title>OpenAI Just Solved a Problem Open Since 1999. It Still Can't Ask Its Own Question.</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 05 Aug 2026 11:39:46 +0000</pubDate>
      <link>https://dev.to/dannwaneri/openai-just-solved-a-problem-open-since-1999-it-still-cant-ask-its-own-question-48j0</link>
      <guid>https://dev.to/dannwaneri/openai-just-solved-a-problem-open-since-1999-it-still-cant-ask-its-own-question-48j0</guid>
      <description>&lt;p&gt;Four days after I published a piece arguing LLMs can't make the jump, &lt;a href="https://openai.com/index/ten-advances-in-mathematics/" rel="noopener noreferrer"&gt;OpenAI announced&lt;/a&gt; that an internal model called Astra had solved ten open problems in mathematics and theoretical computer science. One of them had been open since 1999.&lt;/p&gt;

&lt;p&gt;I'm not going to pretend that's a comfortable coincidence to sit with. So let's sit with it properly instead of pretending it didn't happen.&lt;/p&gt;




&lt;p&gt;The headline result is a non-sofic group. Mikhail Gromov introduced the concept of soficity in 1999 and asked whether every countable group has to be sofic. Twenty-seven years, no mathematician managed to prove or disprove it. Astra built the counterexample. The certificate ships on GitHub in Lean 4, formally verified, "sorry" count zero — meaning no step in the proof was left unproven, no trust in OpenAI required. Total inference cost for all ten results combined: about $2,000.&lt;/p&gt;

&lt;p&gt;Thomas Bloom, who curates the Erdős problems catalogue at Manchester, called it big news. Worth knowing: Bloom is the same mathematician who publicly dismantled an earlier false OpenAI math claim last October. His endorsement here isn't a company's own press release getting nodded along. It's the field's most skeptical reader saying this one holds.&lt;/p&gt;

&lt;p&gt;So: extraordinary, verified, real. Now the question that actually matters for the piece I wrote.&lt;/p&gt;




&lt;p&gt;&lt;a href="https://x.com/ValerioCapraro" rel="noopener noreferrer"&gt;Valerio Capraro&lt;/a&gt;, a mathematician who did his PhD on a problem adjacent to Gromov's conjecture, posted the sharpest version of the distinction I was reaching for and didn't quite land. Astra solved difficult problems inside existing conceptual worlds. Calculus, topology, scheme theory did something different — they didn't answer questions sitting inside a framework, they built frameworks new questions could be asked in.&lt;/p&gt;

&lt;p&gt;Non-sofic groups existing or not was always a well-posed question inside group theory as it already stood. Astra found the object. It didn't invent group theory. That's the line: solving hard problems inside a conceptual world is not the same act as inventing the world.&lt;/p&gt;

&lt;p&gt;Worth naming plainly, because it cuts the other way against overclaiming too: even a Lean certificate that type-checks doesn't confirm the formal statement actually captures the open problem the way mathematicians understood it. Someone still has to judge whether the formalization is asking the right question. That judgment is exactly the kind of move nobody automated here.&lt;/p&gt;




&lt;p&gt;A commenter, &lt;a href="https://dev.to/seo_4d8e85d23c06d94326f27"&gt;Seo&lt;/a&gt;, pushed on something I'd been sloppy about. Is "the jump" one mechanism, or several? Einstein's move was importing an outside framework — he read Hume and Mach until he had the nerve to throw out absolute simultaneity. Dirac's move was different. He wasn't handed a wrong answer. Bohr told him Klein had already solved the relativistic electron problem. Dirac went and found a different one anyway, because Klein's didn't fit what he called his darling theory.&lt;/p&gt;

&lt;p&gt;Importing something from outside the problem, and rejecting a correct-but-unsatisfying answer on the strength of your own priors, are not obviously the same action. I don't have a clean answer for whether they reduce to one mechanism. I'd rather leave that open than force it, because forcing it is exactly the kind of premature tidiness the whole piece is arguing against.&lt;/p&gt;




&lt;p&gt;Here's where it stops being abstract. &lt;a href="https://www.seangoedecke.com/llms-reward-expertise/" rel="noopener noreferrer"&gt;Sean Goedecke wrote about&lt;/a&gt; Terence Tao's public conversation with ChatGPT on a counterexample to the Jacobian Conjecture. Tao's messages are short. The model's outputs, talking to him, are unusually concise — expertise shunts it out of explaining-to-amateurs mode. He pushes back without contradicting directly: "this looks more complex than I was hoping for." And the detail that matters most: Tao makes the leaps himself. He almost never takes the model's suggested next move.&lt;/p&gt;

&lt;p&gt;Goedecke's conclusion: the human is the bottleneck, not the model, because the hard part is communicating exactly what kind of solution you want. The information is already in the model. It takes a very smart human to pull it out.&lt;/p&gt;

&lt;p&gt;That's my bookmark-time argument, relocated. I've been deciding what's worth saving since 2016, one bookmark at a time, and calling that curation. Tao is doing the same thing in real time, inside a chat window, calling it prompting. Different timescale, same move: supply the frame, let the model fill it.&lt;/p&gt;




&lt;p&gt;So the thesis needs updating, not abandoning. Not "LLMs can't jump." Something narrower and, I think, more true. Inside closed, formally verifiable worlds — math, code, games, anything with a Lean checker or a compiler or a scoreboard — the jump is getting crackable by scale and search, and Astra just proved it faster than I expected. Outside those worlds, in anything ambiguous, causally tangled, unverifiable in advance, nobody's shown it yet. Not the actual Einstein case. Not the actual geophysics case. Not the actual "is this bookmark worth keeping" case. And the people getting the most out of these models, Tao included, are the ones still doing that part themselves.&lt;/p&gt;

&lt;p&gt;I got four days. Most theses don't get tested this fast, or this publicly. I'd rather be corrected in the open than be right by accident.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>openai</category>
      <category>rag</category>
    </item>
    <item>
      <title>dev.to's Dashboard Can't Count Its Own Posts</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Mon, 03 Aug 2026 07:38:53 +0000</pubDate>
      <link>https://dev.to/dannwaneri/devtos-dashboard-cant-count-its-own-posts-3fci</link>
      <guid>https://dev.to/dannwaneri/devtos-dashboard-cant-count-its-own-posts-3fci</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;forem is the open source platform behind dev.to itself. I've had it starred and cloned for months and never opened the codebase — it's Rails, and I don't write Ruby. Jess's post was the reason that finally changed.&lt;/p&gt;

&lt;p&gt;github.com/forem/forem&lt;/p&gt;

&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/forem/forem/issues/23687" rel="noopener noreferrer"&gt;#23687&lt;/a&gt; is a one-line report: a user published exactly one post, and the dashboard's "Posts" counter said 2.&lt;/p&gt;

&lt;p&gt;Not writing Ruby meant I couldn't guess my way to the fix from vibes. I had to actually trace it — reading &lt;code&gt;DashboardsController&lt;/code&gt;, the sidebar partials, and the &lt;code&gt;Article&lt;/code&gt; model until the shape of the bug was undeniable, not assumed.&lt;/p&gt;

&lt;p&gt;The "Posts" badge in the dashboard sidebar renders &lt;code&gt;@user.articles_count&lt;/code&gt; — a &lt;code&gt;counter_culture&lt;/code&gt; cache on &lt;code&gt;User&lt;/code&gt; that increments for every &lt;code&gt;Article&lt;/code&gt; row belonging to that user, full stop. No filter on type, no filter on state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ruby"&gt;&lt;code&gt;&lt;span class="c1"&gt;# app/models/article.rb&lt;/span&gt;
&lt;span class="n"&gt;counter_culture&lt;/span&gt; &lt;span class="ss"&gt;:user&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;But the link that badge sits on always opens the same default view: &lt;code&gt;DashboardsController#show&lt;/code&gt; with no params. That view only lists &lt;strong&gt;non-archived, full-post-type&lt;/strong&gt; articles:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ruby"&gt;&lt;code&gt;&lt;span class="c1"&gt;# app/controllers/dashboards_controller.rb&lt;/span&gt;
&lt;span class="vi"&gt;@articles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;articles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_subforem&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="ss"&gt;:organization&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="vi"&gt;@articles&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="ss"&gt;:state&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="s2"&gt;"status"&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="vi"&gt;@articles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;statuses&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="vi"&gt;@articles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;full_posts&lt;/span&gt;
&lt;span class="vi"&gt;@show_archived&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="ss"&gt;:filter&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;to_s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;casecmp&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;"archived"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;zero?&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Forem has three article types — &lt;code&gt;full_post&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt; (a short "Boost" update), and &lt;code&gt;fullscreen_embed&lt;/code&gt; — and the counter doesn't distinguish between them, or between archived and active. The badge counts everything. The list under it shows a strict subset. Anyone who's ever posted a status update, or archived a post, sees a number that doesn't match what they can actually click into and see — exactly what got reported in #23687.&lt;/p&gt;

&lt;p&gt;I couldn't verify that in Ruby, but I recognized the shape of it instantly once it was laid out: a cached count drifting from what a filtered view actually renders. I've shipped that exact bug in JavaScript. Same failure, different syntax.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;github.com/forem/forem/pull/23690&lt;/p&gt;

&lt;p&gt;The fix doesn't touch the shared &lt;code&gt;articles_count&lt;/code&gt; counter — that cache is read elsewhere for badges and spam heuristics, where "every article this user has ever made" is the correct meaning. Instead, &lt;code&gt;DashboardsController&lt;/code&gt; gets a helper scoped to match what the Posts tab actually renders, and both the full-page and AJAX sidebar actions use it instead of the raw cache:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight ruby"&gt;&lt;code&gt;&lt;span class="c1"&gt;# The "Posts" nav item always links to the default (non-archived, full posts&lt;/span&gt;
&lt;span class="c1"&gt;# only) view of the user's own dashboard, so its indicator should reflect&lt;/span&gt;
&lt;span class="c1"&gt;# that same scope rather than the user's raw articles_count, which also&lt;/span&gt;
&lt;span class="c1"&gt;# includes statuses and archived posts that never show up in that list.&lt;/span&gt;
&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;posts_count_for&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="n"&gt;user&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;articles&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;from_subforem&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;full_posts&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;where&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="ss"&gt;archived: &lt;/span&gt;&lt;span class="kp"&gt;false&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;count&lt;/span&gt;
&lt;span class="k"&gt;end&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;There was no way for me to eyeball this and trust it — I can't read Ruby well enough for that, and there's no Ruby or Postgres on the machine I was working from, so I couldn't run the spec suite locally either. Verification had to happen somewhere else: I wrote regression specs asserting a user with one full post, one status, and one archived post should see a count of exactly 1, pushed the branch, and let Forem's own CI be the judge instead of my own confidence.&lt;/p&gt;

&lt;p&gt;CI caught something real on the first run — not in the fix, in my test. &lt;code&gt;create(:article, type_of: "status")&lt;/code&gt; failed its own model validation, because status-type articles in Forem aren't allowed to have body markdown, and the factory's default does. I found the pattern already used elsewhere in the suite (&lt;code&gt;body_markdown: "", main_image: nil&lt;/code&gt;), fixed the two specs, and pushed again.&lt;/p&gt;

&lt;p&gt;That failure is the actual proof this wasn't guesswork dressed up as a fix. If I'd been able to run specs locally I might have caught it before pushing; instead the project's own CI did the job a local run would have.&lt;/p&gt;

&lt;p&gt;Same lesson my other two entries kept landing on: &lt;a href="https://dev.to/dannwaneri/the-cloudflare-worker-that-ran-perfectly-and-still-failed-twice-17l2"&gt;The Cloudflare Worker That Ran Perfectly and Still Failed Twice&lt;/a&gt; and &lt;a href="https://dev.to/dannwaneri/i-was-filming-a-demo-of-my-monitoring-tool-the-monitor-wasnt-monitoring-1p7d"&gt;I Was Filming a Demo of My Monitoring Tool. The Monitor Wasn't Monitoring.&lt;/a&gt; — "it compiled" and "it's correct" are different claims, and only one of them is worth trusting.&lt;/p&gt;

&lt;p&gt;Everything's green now — 19 successful checks, 1 skipped, 0 failures, including the shard that runs &lt;code&gt;dashboard_spec.rb&lt;/code&gt;. The PR is open against forem/forem and waiting on a maintainer review, since third-party fork PRs need one before merge. Not merged yet as of writing this — I'd rather say that plainly than imply otherwise.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Different from my other two entries in one way: I don't write Ruby. Claude found the bug and wrote the fix. I picked the issue and gated everything that left my machine — the fork, the push, the PR, the CLA. Full delegation on the code, not on whether it shipped.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>bugsmash</category>
      <category>devchallenge</category>
      <category>ai</category>
    </item>
    <item>
      <title>How BrowserAct Fixed the Stale-Selector Failures Breaking My Browser Tasks</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Fri, 31 Jul 2026 11:06:54 +0000</pubDate>
      <link>https://dev.to/dannwaneri/how-browseract-fixed-the-stale-selector-failures-breaking-my-browser-tasks-52b5</link>
      <guid>https://dev.to/dannwaneri/how-browseract-fixed-the-stale-selector-failures-breaking-my-browser-tasks-52b5</guid>
      <description>&lt;p&gt;&lt;em&gt;Disclosure: BrowserAct sponsored this piece. The BrowserAct links below are affiliate-tracked — I get credit if you sign up through them.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;I keep hitting the same failure building agent workflows in Claude Code: the agent captures a selector once, the page re-renders with a new build hash, and the next run breaks on an element that's still right there on the screen — just under a different id. BrowserAct is a browser automation platform built for AI agents — real browser control, persistent browser identity, task sessions, verification handling, and human handoff, all behind one CLI. This test is about one narrow slice of that: how it handles a page whose ids and classes change under you, using Claude Code to drive it.&lt;/p&gt;

&lt;p&gt;I tested that against a page I built myself, compared directly against raw Playwright.&lt;/p&gt;

&lt;h2&gt;
  
  
  The problem, concretely
&lt;/h2&gt;

&lt;p&gt;Here's the shape of it. A frontend re-renders with fresh CSS-module or styled-components hashes on every deploy. You write:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;locator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;button&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;getAttribute&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;reload&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;page&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;locator&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`#&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;click&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3000&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That works the day you write it. It breaks the next time the build hash changes, and the failure you get is a bare timeout with no explanation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;locator.click: Timeout 3000ms exceeded.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Playwright's own role/text locators fix this specific case — &lt;code&gt;getByRole("button", { name: "Submit" })&lt;/code&gt; survives id/class churn fine, because it never depended on the hash in the first place. That's a real fix, not a workaround. What it doesn't give you is a representation of the page an agent can reason about turn by turn, or a signal that says "the page under you just changed, stop and re-check." That's the gap BrowserAct is actually filling.&lt;/p&gt;

&lt;h2&gt;
  
  
  The interaction model
&lt;/h2&gt;

&lt;p&gt;BrowserAct doesn't hand Claude Code a selector at all. The loop is: read the current state, choose an action from what's actually there, execute it, reassess.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;state  -&amp;gt; indexed list of interactive elements, as of right now
click &amp;lt;index&amp;gt;  -&amp;gt; act on one of them
state  -&amp;gt; read again, because the page may have changed
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The index is scoped to that one snapshot. It's not a selector, not a stable id — it's a pointer into "what state just returned," and it expires the moment the page does something a new &lt;code&gt;state&lt;/code&gt; call would notice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Setting up BrowserAct
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;uv tool &lt;span class="nb"&gt;install &lt;/span&gt;browser-act-cli &lt;span class="nt"&gt;--python&lt;/span&gt; 3.12
browser-act &lt;span class="nt"&gt;--version&lt;/span&gt;
browser-act browser create &lt;span class="nt"&gt;--name&lt;/span&gt; &lt;span class="s2"&gt;"dom-drift-test"&lt;/span&gt; &lt;span class="nt"&gt;--type&lt;/span&gt; chrome &lt;span class="nt"&gt;--desc&lt;/span&gt; &lt;span class="s2"&gt;"local churn test"&lt;/span&gt;
browser-act browser list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;browser list&lt;/code&gt; after creation is how you get the browser id you'll pass to every session command — it's not returned anywhere else. I used the &lt;code&gt;chrome&lt;/code&gt; browser type (a local, blank browser, immediate to create) rather than &lt;code&gt;stealth&lt;/code&gt;, since this test is about selector churn on a page I control, not anti-bot evasion — &lt;code&gt;stealth&lt;/code&gt; requires an API key and a billed purchase flow that has nothing to do with what I was testing. Creating a local &lt;code&gt;chrome&lt;/code&gt; browser completed immediately with no purchase page involved.&lt;/p&gt;

&lt;p&gt;Versions tested: &lt;code&gt;browser-act-cli&lt;/code&gt; v1.1.0, Playwright 1.61.1, Python 3.12.13. To reproduce this: the test page is a small Node server that reshuffles element ids, classes, and order on every reload — no third-party site involved — plus the Playwright scripts run against the same page.&lt;/p&gt;

&lt;p&gt;Skill source: &lt;a href="https://github.com/browser-act/skills/tree/main/browser-act" rel="noopener noreferrer"&gt;browser-act/skills&lt;/a&gt;. Installation details: &lt;a href="https://github.com/browser-act/skills/blob/main/docs/installation.md" rel="noopener noreferrer"&gt;docs/installation.md&lt;/a&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; domtest browser open &amp;lt;browser-id&amp;gt; http://localhost:8934
browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; domtest state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;state&lt;/code&gt; returns an indexed list, not a selector:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;[1]&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;item-r8k7gx&lt;/span&gt; &lt;span class="na"&gt;invalid=&lt;/span&gt;&lt;span class="s"&gt;false&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
        Bravo
[2]&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;class=&lt;/span&gt;&lt;span class="s"&gt;item-cr2ajb&lt;/span&gt; &lt;span class="na"&gt;invalid=&lt;/span&gt;&lt;span class="s"&gt;false&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
        Charlie
[5]&lt;span class="nt"&gt;&amp;lt;button&lt;/span&gt; &lt;span class="na"&gt;id=&lt;/span&gt;&lt;span class="s"&gt;submit-btn-ckeewz&lt;/span&gt; &lt;span class="na"&gt;invalid=&lt;/span&gt;&lt;span class="s"&gt;false&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
        Submit
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You act on the index (&lt;code&gt;click 5&lt;/code&gt;), not the class or id, so the build-hash problem doesn't apply to it — there's no hash in the reference to go stale.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual comparison
&lt;/h2&gt;

&lt;p&gt;I built a local test page that reshuffles element ids and item order on every reload, specifically to force this failure. Playwright's role locators, as noted above, survive that fine. The difference isn't that Playwright can't handle churn — it's what happens when you act on a reference that's gone stale anyway, whether from a captured selector or an old snapshot:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Playwright, selector captured once, reused after reload:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FAILED clicking #submit-btn-g9mpcm:
  locator.click: Timeout 3000ms exceeded.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;BrowserAct, index captured once, reused after reload:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error 210603: The snapshot belongs to a different page or tab than the
current one. Run 'browser-act state' again on the current page.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both refuse to silently click the wrong thing. But BrowserAct's error names the exact cause and the exact fix — re-run &lt;code&gt;state&lt;/code&gt; — where Playwright's is a generic timeout you already have to know how to interpret.&lt;/p&gt;

&lt;h2&gt;
  
  
  The recovery flow, actually run
&lt;/h2&gt;

&lt;p&gt;That error is only useful if what follows it actually works. Here's the full loop, one continuous session, real output, no steps skipped — the error and the successful recovery from it, back to back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 browser open &amp;lt;browser-id&amp;gt; http://localhost:8934
browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[5]&amp;lt;button id=submit-btn-5zhwmn invalid=false /&amp;gt;
        Submit
load 1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft60c9eb4eygllvrtn8c7.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ft60c9eb4eygllvrtn8c7.png" alt="state before the page changes, load 1, Submit at index 5" width="799" height="362"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Page changes, and the old reference is used anyway — the error from earlier, live:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 reload
browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 click 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Error 210603: The snapshot belongs to a different page or tab than the
current one. Run 'browser-act state' again on the current page.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Re-check, don't retry blind:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 state
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[5]&amp;lt;button id=submit-btn-c66evj invalid=false /&amp;gt;
        Submit
load 3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The id changed (&lt;code&gt;submit-btn-5zhwmn&lt;/code&gt; to &lt;code&gt;submit-btn-c66evj&lt;/code&gt;), and the fresh &lt;code&gt;state&lt;/code&gt; call caught it. Acting on the new index closes the loop:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;browser-act &lt;span class="nt"&gt;--session&lt;/span&gt; shot1 click 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;clicked=5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh0awk6iyzooe696a85h3.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fh0awk6iyzooe696a85h3.png" alt="reload, stale click producing Error 210603, fresh state at load 3, then the successful click" width="796" height="97"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;State, act, hit the error, reassess, act again — completed, in the same session the error happened in.&lt;/p&gt;

&lt;h2&gt;
  
  
  What BrowserAct Actually Solved
&lt;/h2&gt;

&lt;p&gt;It reduced my dependence on generated ids and CSS classes to zero for anything clickable, and gave Claude Code a predictable recovery path the moment the page changed underneath it: re-run &lt;code&gt;state&lt;/code&gt;, act on what's actually there now.&lt;/p&gt;

&lt;p&gt;Two real limits, not softened. &lt;code&gt;state&lt;/code&gt; only indexes interactive elements — in my test, the list items only got indices once I made them buttons. For plain text or document content, the right tool isn't &lt;code&gt;state&lt;/code&gt; at all, it's BrowserAct's content-extraction commands (&lt;code&gt;get markdown&lt;/code&gt;, &lt;code&gt;get text &amp;lt;index&amp;gt;&lt;/code&gt;) — treating &lt;code&gt;state&lt;/code&gt; as a universal page parser is the wrong mental model. And separately: the &lt;code&gt;title&lt;/code&gt; field in &lt;code&gt;state&lt;/code&gt;'s own output lagged behind the actual page content in one of my runs — the visible elements had already changed, &lt;code&gt;title&lt;/code&gt; hadn't caught up yet. Worth knowing if you're tempted to key any check off &lt;code&gt;title&lt;/code&gt; specifically.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Check I Still Kept Explicit
&lt;/h2&gt;

&lt;p&gt;BrowserAct solved the selector problem. It didn't remove the need to think about website authentication separately. Three different things are in play here, and it's worth naming them precisely: the browser identity (the &lt;code&gt;chrome&lt;/code&gt; browser instance itself), the BrowserAct task session (&lt;code&gt;--session domtest&lt;/code&gt;), and the target website's own authentication session.&lt;/p&gt;

&lt;p&gt;In an earlier run against the same test page, the website authentication state expired while the BrowserAct task session remained active — &lt;code&gt;state&lt;/code&gt; kept responding normally, it just started describing a "please sign in again" page instead of the dashboard. BrowserAct exposed that changed page state clearly enough for Claude Code to decide what to do next: continue, re-authenticate, or hand off to a human. It didn't decide that automatically, and I don't think it should — that's still a call I want made explicitly, not inferred.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where this leaves it
&lt;/h2&gt;

&lt;p&gt;Claude Code and BrowserAct solved the actual failure mode I set out to test: brittle selector maintenance replaced with a workflow built on current page state, indexed actions, and an explicit recovery step when the ground shifts. That's a narrower claim than "browser automation solved," and it's the one I can actually stand behind.&lt;/p&gt;




&lt;p&gt;Try BrowserAct: &lt;a href="https://www.browseract.ai/Daniel" rel="noopener noreferrer"&gt;browseract.ai/Daniel&lt;/a&gt;&lt;br&gt;
BrowserAct Skills: &lt;a href="https://www.browseract.com/?co-from=Daniel&amp;amp;redirect=https://github.com/browser-act/skills/tree/main" rel="noopener noreferrer"&gt;github.com/browser-act/skills&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;em&gt;(Both links above are affiliate-tracked to my account.)&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>automation</category>
      <category>discuss</category>
      <category>productivity</category>
    </item>
    <item>
      <title>Why Kimi K3 Still Can't Do What Einstein Did</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 29 Jul 2026 09:30:59 +0000</pubDate>
      <link>https://dev.to/dannwaneri/why-kimi-k3-still-cant-do-what-einstein-did-2l6d</link>
      <guid>https://dev.to/dannwaneri/why-kimi-k3-still-cant-do-what-einstein-did-2l6d</guid>
      <description>&lt;p&gt;In geophysics you almost never get to see the thing you're studying. You get a seismic trace, a gravity anomaly, a resistivity curve. You don't get the rock. You get the rock's echo, and you have to guess at a structure underground that would produce exactly that echo and no other. Nobody hands you the answer key. You infer the case from the result.&lt;/p&gt;

&lt;p&gt;I hadn't thought about that part of my degree in years, until I built Bookmark Brain.&lt;/p&gt;




&lt;p&gt;Bookmark Brain is a RAG pipeline trained on my own X bookmarks and likes, saved since 2016. It was around 50,000 the last time I wrote about this. It's 70,000 now. Cron jobs pull in new saves on their own; there's no manual re-curation involved. Ask it something and it retrieves the closest matching saved content, then composes an answer that sounds like me. It works well. Too well, honestly. When I asked it about API design opinions, it sounded more like me than most general-purpose models do when I prompt them to write in my voice.&lt;/p&gt;

&lt;p&gt;The reason isn't the model. It's the retrieval layer. My bookmarks are coherent because I spent a decade curating them into a specific worldview. The bot just finds the nearest neighbor and composes a fluent sentence around it. What it can't do is the thing I actually needed a few times while testing it: resolve a contradiction between two things I'd bookmarked years apart. It doesn't reconcile them. It picks whichever one is semantically closer to the question and hands that back.&lt;/p&gt;

&lt;p&gt;That's not a bug in my pipeline. That's the whole category of thing retrieval can't do.&lt;/p&gt;




&lt;p&gt;A 2014 blog post by Amni Rusli, "Irreplaceable Us," made roughly this same argument, minus the RAG pipeline. Machines work within given parameters, she wrote — feed them data and code and they'll optimize inside that space forever. What they can't do is leap into a different pond of parameters entirely and haul something back. She used Einstein and Dirac as her examples. Einstein reading Hume and Mach until he had the nerve to throw out absolute simultaneity. Dirac telling Bohr that Klein had already solved the relativistic electron problem and going looking for a different answer anyway, because the existing one didn't fit his "darling" theory.&lt;/p&gt;

&lt;p&gt;Twelve years later, in January, Google DeepMind published a paper making the same claim with the informality stripped out. Tom Zahavy's "LLMs Can't Jump" starts from a diagram Einstein actually drew, in a letter to Maurice Solovine: sense experience jumping to a system of axioms, then deduction working forward from there. Peirce had a name for the gap in that jump, and the paper borrows it. Deduction: rule plus case gives you a result, the only mode that guarantees truth. Induction: case plus result gives you a rule, which is close to what training an LLM on a trillion documents actually is. Abduction: rule plus a surprising result gives you a new case, or a new rule, to explain it. That's the geophysics move. That's resolving the contradiction in my own bookmarks. That's the one nobody has automated.&lt;/p&gt;

&lt;p&gt;The paper's argument for why scaling doesn't fix this is data scarcity. General relativity wasn't induced from a mountain of prior experimental results, because there wasn't one. The axioms can't have been deduced either, since deduction only runs forward from premises someone already has. Something else produced the premises. That something is the jump, and it's still missing.&lt;/p&gt;




&lt;p&gt;Then Kimi K3 shipped. 2.8 trillion parameters, the largest open-weight model ever released, benchmarking close behind Fable 5 and GPT-5.6 Sol on coding and agentic tasks. Moonshot's own claim is 2.5x the intelligence per unit of compute over their last generation. None of that changes the argument. A bigger training set makes induction better and deduction more reliable over longer chains. It doesn't add a third capability that wasn't there before. Feed a model more of the internet and you get a more convincing compositor, not a different kind of thing.&lt;/p&gt;

&lt;p&gt;I said this to a commenter under my original bot post, before I'd read the DeepMind paper: a model trained on Newtonian physics at sufficient scale would produce better Newtonian predictions, not special relativity. Turns out that's the whole thesis, just arrived at from a different direction — one from watching my own retrieval logs, one from a formal read of Peirce.&lt;/p&gt;




&lt;p&gt;Which brings me back to my own bookmarks.&lt;/p&gt;

&lt;p&gt;What actually happened at bookmark-time, every time I decided a tweet was worth saving, was a small version of the jump. This connects to that. This contradicts what I believed last year. Keep it. I've been doing that since 2016 — a decade of small jumps, one save at a time. Bookmark Brain inherited the residue of all of them. It never makes one itself.&lt;/p&gt;

&lt;p&gt;That's the part that should worry people more than the benchmark charts do. Not that the model can't out-think Einstein. That most of what gets paid for isn't the jump either.&lt;/p&gt;

&lt;p&gt;I write the essay, but I bookmark the argument first. That's where I'm putting the hours now — not in the composing, which the model will keep getting better at, but in the deciding what's worth saving. The jump doesn't scale. Mine, at least, still has to happen one bookmark at a time.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>llm</category>
      <category>rag</category>
      <category>discuss</category>
    </item>
    <item>
      <title>I Was Filming a Demo of My Monitoring Tool. The Monitor Wasn't Monitoring.</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Mon, 27 Jul 2026 08:40:19 +0000</pubDate>
      <link>https://dev.to/dannwaneri/i-was-filming-a-demo-of-my-monitoring-tool-the-monitor-wasnt-monitoring-1p7d</link>
      <guid>https://dev.to/dannwaneri/i-was-filming-a-demo-of-my-monitoring-tool-the-monitor-wasnt-monitoring-1p7d</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Smash Stories&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The project
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;workers-monitor&lt;/code&gt; is a Cloudflare Worker I run to watch a small fleet of&lt;br&gt;
my own Workers. Hourly cron, deterministic threshold gate, Claude Haiku&lt;br&gt;
judgement only if the gate trips, Telegram alert if Haiku confirms&lt;br&gt;
something's actually wrong. A quiet hour makes zero LLM calls.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/dannwaneri/workers-monitor" rel="noopener noreferrer"&gt;github.com/dannwaneri/workers-monitor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;It pages me on Telegram when something's wrong with the fleet it&lt;br&gt;
watches. What it didn't have, until recently, was an equivalent safety&lt;br&gt;
net for itself — if &lt;code&gt;workers-monitor&lt;/code&gt; broke, nobody got told. So I added&lt;br&gt;
Sentry: automatic error capture, and &lt;code&gt;Sentry.withMonitor()&lt;/code&gt; around the&lt;br&gt;
hourly run so a dead cron trigger would be caught immediately instead of&lt;br&gt;
silently waiting up to 24 hours for the next daily heartbeat to go quiet.&lt;/p&gt;

&lt;p&gt;I deployed it. A cron tick ran clean. I moved on, fairly pleased with&lt;br&gt;
myself.&lt;/p&gt;
&lt;h2&gt;
  
  
  The bug
&lt;/h2&gt;

&lt;p&gt;A few days later I sat down to record a demo video of the Sentry&lt;br&gt;
integration — screen capture, narration, the whole thing, for a separate&lt;br&gt;
part of this challenge. Simple plan: open the Sentry dashboard, show the&lt;br&gt;
captured errors, show the cron monitor's check-in, done.&lt;/p&gt;

&lt;p&gt;I opened the Monitors page to get the screenshot.&lt;/p&gt;

&lt;p&gt;There was nothing there. Just Sentry's own auto-created generic "Error&lt;br&gt;
Monitor" — no &lt;code&gt;workers-monitor-hourly-poll&lt;/code&gt;, no Cron-type entry&lt;br&gt;
whatsoever, despite the code having run successfully, every hour, for&lt;br&gt;
days. I'd been carrying around a completely false belief: that deploying&lt;br&gt;
&lt;code&gt;Sentry.withMonitor()&lt;/code&gt; and watching it execute without error meant it was&lt;br&gt;
monitoring. It wasn't. It never had been.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Sentry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withMonitor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;workers-monitor-hourly-poll&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Nothing about this throws. It compiles. It deploys. The wrapped function&lt;br&gt;
runs exactly as intended, every hour, on schedule. And it still wasn't&lt;br&gt;
doing the one thing it existed to do.&lt;/p&gt;
&lt;h2&gt;
  
  
  How I found it
&lt;/h2&gt;

&lt;p&gt;By accident, honestly. Not through review, not through testing — I found&lt;br&gt;
it because I was trying to film proof that something worked, and the&lt;br&gt;
proof wasn't there. If I hadn't been making a video, I might not have&lt;br&gt;
looked at that specific dashboard page for weeks.&lt;/p&gt;
&lt;h2&gt;
  
  
  The fix
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;withMonitor&lt;/code&gt;'s check-in has nowhere to attach without a schedule —&lt;br&gt;
Sentry needs to know what "on time" even means for this monitor before it&lt;br&gt;
will create a Cron Monitor entity to check in against. Without that&lt;br&gt;
config, the call silently has no monitor to report to.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Sentry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withMonitor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;workers-monitor-hourly-poll&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;crontab&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;0 * * * *&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;// matches the real cron trigger&lt;/span&gt;
    &lt;span class="na"&gt;checkinMargin&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;maxRuntime&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UTC&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;One config object. I deployed it, then made myself actually wait for a&lt;br&gt;
real hourly tick — not a local test, the genuine production cron —&lt;br&gt;
before I let myself believe it was fixed. &lt;code&gt;workers-monitor-hourly-poll&lt;/code&gt;&lt;br&gt;
showed up as a real Cron monitor afterward, "Every hour" schedule and&lt;br&gt;
all.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this one mattered to fix
&lt;/h2&gt;

&lt;p&gt;Because it's the same lie twice, if I'm honest about it. I built a&lt;br&gt;
monitoring tool because I don't trust systems to tell me the truth about&lt;br&gt;
their own state unprompted. Then I shipped a piece of that exact tool&lt;br&gt;
without applying the same skepticism to itself. "It ran without throwing"&lt;br&gt;
is not evidence of "it's doing its job" — I know that, I'd have said it&lt;br&gt;
confidently to anyone who asked — and I still fell for the gap between&lt;br&gt;
those two claims on my own code, in the one project whose entire purpose&lt;br&gt;
is not falling for that gap.&lt;/p&gt;

&lt;h2&gt;
  
  
  See it
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://youtu.be/dSvSOw0fy-E" rel="noopener noreferrer"&gt;https://youtu.be/dSvSOw0fy-E&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Real Sentry dashboard, the actual moment the cron monitor showed up after&lt;br&gt;
the fix — not a re-enactment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I'd do differently next time
&lt;/h2&gt;

&lt;p&gt;Treat "I deployed it and it didn't error" as a hypothesis, not a&lt;br&gt;
conclusion — for every feature, not just the ones I already suspect are&lt;br&gt;
fragile. The KV read bug I fixed earlier in this project (a separate&lt;br&gt;
entry, if you're comparing notes) came from a structured spec review.&lt;br&gt;
This one came from dumb luck — I happened to need a screenshot. I'd&lt;br&gt;
rather it come from the habit than the accident next time.&lt;/p&gt;

&lt;h2&gt;
  
  
  PR
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/dannwaneri/workers-monitor/pull/3" rel="noopener noreferrer"&gt;github.com/dannwaneri/workers-monitor/pull/3&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;One file changed, 14 insertions, 1 deletion — the isolated fix, nothing&lt;br&gt;
else riding along with it.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>sentry</category>
      <category>ai</category>
    </item>
    <item>
      <title>The Cloudflare Worker That Ran Perfectly and Still Failed Twice</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Mon, 27 Jul 2026 08:40:07 +0000</pubDate>
      <link>https://dev.to/dannwaneri/the-cloudflare-worker-that-ran-perfectly-and-still-failed-twice-17l2</link>
      <guid>https://dev.to/dannwaneri/the-cloudflare-worker-that-ran-perfectly-and-still-failed-twice-17l2</guid>
      <description>&lt;p&gt;&lt;em&gt;This is a submission for &lt;a href="https://dev.to/bugsmash"&gt;DEV's Summer Bug Smash: Clear the Lineup&lt;/a&gt; powered by &lt;a href="https://sentry.io/" rel="noopener noreferrer"&gt;Sentry&lt;/a&gt;.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Project Overview
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;workers-monitor&lt;/code&gt; is a Cloudflare Worker that watches a small fleet of my&lt;br&gt;
own Workers. Hourly cron, pulls fleet metrics from Cloudflare's GraphQL&lt;br&gt;
Analytics API, runs a deterministic threshold gate, and only calls Claude&lt;br&gt;
Haiku to judge signal-vs-noise if the gate trips. If Haiku confirms&lt;br&gt;
something's actually wrong, it sends a Telegram alert. A quiet hour makes&lt;br&gt;
zero LLM calls.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/dannwaneri/workers-monitor" rel="noopener noreferrer"&gt;github.com/dannwaneri/workers-monitor&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://youtu.be/APot7n2wYB8" rel="noopener noreferrer"&gt;https://youtu.be/APot7n2wYB8&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Real alert history from the deployed Worker — the daily heartbeat, and&lt;br&gt;
scrolled back, the genuine incidents this system has actually caught.&lt;/p&gt;

&lt;p&gt;This submission covers two bugs, found days apart, that turn out to be&lt;br&gt;
the same failure mode wearing different clothes: code that compiles,&lt;br&gt;
deploys cleanly, and runs without throwing — and still doesn't do the one&lt;br&gt;
thing it was written for. "No error" is not the same claim as "working&lt;br&gt;
correctly," and both bugs below only exist because that distinction got&lt;br&gt;
missed once.&lt;/p&gt;
&lt;h2&gt;
  
  
  Bug Fix or Performance Improvement
&lt;/h2&gt;
&lt;h3&gt;
  
  
  Bug 1: the maintenance window's fail-open gap
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;workers-monitor&lt;/code&gt; has a maintenance-window feature — &lt;code&gt;POST&lt;/code&gt; a start/end&lt;br&gt;
time and it suppresses Telegram alerts during a planned deploy, without&lt;br&gt;
stopping the gate or the logging. The function that reads that window had&lt;br&gt;
one non-negotiable rule, written as a comment directly above it:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Malformed data fails OPEN (returns null → no suppression) — a broken&lt;br&gt;
window read must never accidentally silence a real incident.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;The code didn't actually satisfy that comment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;readMaintenanceWindow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Env&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MaintenanceWindow&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;STATE&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;MAINTENANCE_KEY&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;MaintenanceWindow&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isNaN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isNaN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unparseable start/end timestamps&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="cm"&gt;/* ... */&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only &lt;code&gt;JSON.parse&lt;/code&gt; and the timestamp validation sit inside the &lt;code&gt;try&lt;/code&gt; block&lt;br&gt;
— &lt;code&gt;await env.STATE.get(...)&lt;/code&gt; on the line above does not. A genuine KV&lt;br&gt;
read failure (a real outage, not bad data) throws an exception nothing in&lt;br&gt;
this function catches. It propagates out of &lt;code&gt;readMaintenanceWindow&lt;/code&gt;, out&lt;br&gt;
of the &lt;code&gt;run()&lt;/code&gt; function that calls it, and aborts the entire hourly run —&lt;br&gt;
not just alert suppression, but the deterministic gate, the Haiku&lt;br&gt;
judgement call, and the structured logging for that hour, all of it,&lt;br&gt;
silently skipped. Every KV blip during that window meant a full hour of&lt;br&gt;
zero fleet visibility, not just a missed alert.&lt;/p&gt;

&lt;p&gt;I found this via a structured review of the diff against the original&lt;br&gt;
spec's acceptance criteria, not a live incident. One criterion was almost&lt;br&gt;
word-for-word the comment above the function — the review caught that the&lt;br&gt;
code only satisfied half of that sentence.&lt;/p&gt;
&lt;h3&gt;
  
  
  Between the two bugs: closing the class of gap, not just the instance
&lt;/h3&gt;

&lt;p&gt;Fixing bug 1 raised an obvious follow-up question: what happens when&lt;br&gt;
&lt;code&gt;workers-monitor&lt;/code&gt; itself breaks, not the fleet it watches? Nothing —&lt;br&gt;
console.error into &lt;code&gt;wrangler tail&lt;/code&gt;, unread unless I happened to be&lt;br&gt;
watching. So I wired in Sentry: &lt;code&gt;Sentry.withSentry()&lt;/code&gt; around the handler&lt;br&gt;
for automatic error capture, an explicit &lt;code&gt;Sentry.captureException()&lt;/code&gt; on&lt;br&gt;
the deliberately-swallowed exception in the scheduled handler's own&lt;br&gt;
catch, and &lt;code&gt;Sentry.withMonitor()&lt;/code&gt; around the hourly &lt;code&gt;run()&lt;/code&gt; call so a&lt;br&gt;
dead cron trigger gets caught immediately instead of waiting up to 24&lt;br&gt;
hours for the next Telegram heartbeat to go silent.&lt;/p&gt;

&lt;p&gt;I deployed it, watched a cron tick complete without error, and considered&lt;br&gt;
cron monitoring done.&lt;/p&gt;
&lt;h3&gt;
  
  
  Bug 2: the cron monitor that was never actually monitoring
&lt;/h3&gt;

&lt;p&gt;It wasn't done. While recording a demo video of the Sentry integration —&lt;br&gt;
for a separate piece of this challenge — I opened the Monitors dashboard&lt;br&gt;
to screenshot the cron check-in, and there wasn't one. Only Sentry's own&lt;br&gt;
auto-created generic "Error Monitor," nothing for &lt;code&gt;workers-monitor-hourly-poll&lt;/code&gt;&lt;br&gt;
at all, despite the code executing successfully every hour for days.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Sentry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withMonitor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;workers-monitor-hourly-poll&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This isn't wrong syntax. It compiles, deploys, executes the wrapped&lt;br&gt;
function fine. But the check-in has nowhere to attach without a schedule&lt;br&gt;
— Sentry needs to know what "on time" means for this monitor before it&lt;br&gt;
will create a Cron Monitor entity to check in against. Without that&lt;br&gt;
config, the call silently has no monitor to report to. No error, no&lt;br&gt;
warning, just nothing showing up. Exactly the same shape as bug 1: code&lt;br&gt;
that runs clean and still isn't doing its job.&lt;/p&gt;
&lt;h2&gt;
  
  
  Code
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://github.com/dannwaneri/workers-monitor/pull/1" rel="noopener noreferrer"&gt;github.com/dannwaneri/workers-monitor/pull/1&lt;/a&gt; — the fail-open fix, isolated, 12 insertions / 1 deletion.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://github.com/dannwaneri/workers-monitor/pull/3" rel="noopener noreferrer"&gt;github.com/dannwaneri/workers-monitor/pull/3&lt;/a&gt; — the cron monitor fix, isolated, 14 insertions / 1 deletion.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Bug 1 fix — env.STATE.get() gets its own try/catch, separate from the&lt;/span&gt;
&lt;span class="c1"&gt;// existing one around JSON.parse:&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;readMaintenanceWindow&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Env&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MaintenanceWindow&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="na"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;STATE&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;MAINTENANCE_KEY&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;event&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;maintenance_window_read_error&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt; &lt;span class="k"&gt;instanceof&lt;/span&gt; &lt;span class="nb"&gt;Error&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nc"&gt;String&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;}));&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="nx"&gt;MaintenanceWindow&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isNaN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;start&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isNaN&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;end&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unparseable start/end timestamps&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;w&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="cm"&gt;/* ... */&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Bug 2 fix — the missing monitorConfig, matching the real deployed schedule:&lt;/span&gt;
&lt;span class="nx"&gt;Sentry&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;withMonitor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;workers-monitor-hourly-poll&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;run&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;schedule&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;crontab&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;0 * * * *&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="c1"&gt;// matches wrangler.jsonc's trigger&lt;/span&gt;
    &lt;span class="na"&gt;checkinMargin&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// minutes late before considered missed&lt;/span&gt;
    &lt;span class="na"&gt;maxRuntime&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;    &lt;span class="c1"&gt;// minutes before considered timed out&lt;/span&gt;
    &lt;span class="na"&gt;timezone&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;UTC&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;  &lt;span class="c1"&gt;// Cloudflare cron triggers always run in UTC&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  My Improvements
&lt;/h2&gt;

&lt;p&gt;Neither fix is large. What connects them is how each was actually&lt;br&gt;
verified, not assumed:&lt;/p&gt;

&lt;p&gt;For bug 1, I ran a two-part review before trusting the diff: does it&lt;br&gt;
satisfy every Given/When/Then acceptance criterion from the original&lt;br&gt;
spec, and separately, does it resolve every assumption flagged during&lt;br&gt;
design, including ones nobody had explicitly revisited. That second check&lt;br&gt;
is what surfaced it — the acceptance criterion for KV-failure behavior&lt;br&gt;
was already written down before the code existed, and the implementation&lt;br&gt;
simply didn't fully match its own spec.&lt;/p&gt;

&lt;p&gt;For bug 2, the same discipline applied to my own claim: I'd told myself&lt;br&gt;
"deployed, ran without error, cron monitoring is working." I checked the&lt;br&gt;
Monitors dashboard before the fix (empty — confirmed the bug was real,&lt;br&gt;
not a hunch), deployed the fix, waited for a genuine production cron&lt;br&gt;
tick — not a local test, not a manual trigger — and only considered it&lt;br&gt;
done once &lt;code&gt;workers-monitor-hourly-poll&lt;/code&gt; actually appeared as a registered&lt;br&gt;
Cron monitor with an "Every hour" schedule.&lt;/p&gt;

&lt;p&gt;The broader lesson, true of both bugs: "the code deployed and nothing&lt;br&gt;
crashed" is a dramatically weaker claim than "the thing I built does what&lt;br&gt;
I said it does." Those two can look identical from a terminal and be&lt;br&gt;
completely different in reality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Best Use of Sentry
&lt;/h2&gt;

&lt;p&gt;Both Error Monitoring and Cron Monitoring are used here, and bug 2 is&lt;br&gt;
specifically about making the Cron Monitoring half actually work as&lt;br&gt;
intended, not just installed:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Automatic error capture, verified against real output&lt;/strong&gt; —
&lt;code&gt;Sentry.withSentry()&lt;/code&gt; wraps the handler, instrumenting &lt;code&gt;fetch&lt;/code&gt; and
&lt;code&gt;scheduled&lt;/code&gt;. Deploying a temporary route that threw an intentional
error confirmed the capture path end to end — and surfaced something I
didn't expect: alongside the intentional error, Sentry also caught a
second, unrelated &lt;code&gt;TypeError: Cannot read properties of undefined
(reading 'scriptName')&lt;/code&gt;, thrown from inside the SDK's own fetch
wrapper on the same request. Not my bug — Seer's root-cause analysis
traced it to the SDK reading &lt;code&gt;scriptName&lt;/code&gt; off a Cloudflare
execution-context field that wasn't populated for a &lt;code&gt;curl&lt;/code&gt;-originated
request — but it's the kind of thing you only catch by generating a
real event and reading what actually came back, not by trusting that
&lt;code&gt;withSentry()&lt;/code&gt; compiled and calling it done.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Explicit capture on the swallowed exception&lt;/strong&gt; — the scheduled
handler already had a deliberate top-level &lt;code&gt;try/catch&lt;/code&gt; so one bad run
doesn't crash the Worker. That's exactly the case &lt;code&gt;withSentry&lt;/code&gt;'s
automatic uncaught-exception capture can't see, since the exception is
caught before it ever becomes "uncaught" — so the catch block also
calls &lt;code&gt;Sentry.captureException(err)&lt;/code&gt; directly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cron monitoring, now actually registering&lt;/strong&gt; — after the bug 2 fix,
&lt;code&gt;workers-monitor-hourly-poll&lt;/code&gt; shows up as a real Cron monitor with an
"Every hour" schedule. If the trigger stops firing entirely, Sentry
now catches that immediately, instead of the previous 24-hour blind
spot (the daily Telegram heartbeat was the only prior proof-of-life).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Verified, not assumed&lt;/strong&gt;, on both fronts: for error capture, I deployed&lt;br&gt;
a temporary route that threw an intentional error, confirmed it was&lt;br&gt;
captured in the Sentry dashboard, then removed the route before opening&lt;br&gt;
the PR. For cron monitoring, the Monitors dashboard was checked empty&lt;br&gt;
before the fix and populated after a real hourly tick — not a claim, a&lt;br&gt;
before/after I watched happen.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://youtu.be/dSvSOw0fy-E" rel="noopener noreferrer"&gt;https://youtu.be/dSvSOw0fy-E&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The resolved issues, the real stack trace, and the moment the cron&lt;br&gt;
monitor actually registered — the same evidence, on screen.&lt;/p&gt;

</description>
      <category>devchallenge</category>
      <category>bugsmash</category>
      <category>ai</category>
      <category>cloudflare</category>
    </item>
    <item>
      <title>The Only AI Tell That Doesn't Need a Detector</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 22 Jul 2026 13:36:22 +0000</pubDate>
      <link>https://dev.to/dannwaneri/the-only-ai-tell-that-doesnt-need-a-detector-3gkd</link>
      <guid>https://dev.to/dannwaneri/the-only-ai-tell-that-doesnt-need-a-detector-3gkd</guid>
      <description>&lt;p&gt;&lt;a href="https://x.com/paulg/status/2079877240699940955" rel="noopener noreferrer"&gt;Paul Graham posted a test&lt;/a&gt; this week that has nothing to do with sentence structure. Slop gives itself away, he said, when the diction doesn't match the idea, when something completely ordinary gets delivered with the excitement of someone announcing a discovery.&lt;/p&gt;

&lt;p&gt;That's a different axis than everything else I've been reading about detection this month. Pangram scores a pattern: token by token, sentence by sentence, does the shape of this text statistically resemble a machine's output. Sloan, the human version I ran into on DEV.to, did the same thing by ear instead of by classifier, GPTZero as a second opinion. Both are measuring the same thing: surface. PG's test measures a relationship. What's actually being said, against how much weight the delivery is putting behind saying it.&lt;/p&gt;




&lt;p&gt;Run my own flagged pieces through it and they'd pass clean. What got them flagged wasn't excitement outrunning substance, it was named data points and short paragraphs doing real argumentative work, plainly. Pangram's classifier and PG's ear would disagree with each other on the same writing. That's worth sitting with. The tool built to formalize the intuition doesn't actually agree with the intuition once you test them against the same text.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/pascal_cescato_692b7a8a20"&gt;Pascal's&lt;/a&gt; comment on the last piece fits here too. Rephrasing his own English for flow, after 6 to 20 hours of writing the argument himself, isn't a case of ordinary ideas dressed up as brilliant ones. It's someone's real thinking, translated. Nothing in that process produces the register mismatch PG is describing. A classifier flagged it anyway, at 97% confidence, on a post from 2017.&lt;/p&gt;




&lt;p&gt;The other thing PG's test doesn't need is a platform. Chris Best's whole pitch for &lt;a href="https://post.substack.com/p/against-claudefishing" rel="noopener noreferrer"&gt;shipping Pangram into Substack&lt;/a&gt; is that reader intuition doesn't scale, you need a tool doing this at the volume a feed operates at. PG is quietly arguing the opposite: the tell was always available to anyone reading carefully, you just have to know what you're listening for. That's closer to Josh Puckett's complaint in &lt;a href="https://x.com/joshpuckett" rel="noopener noreferrer"&gt;"In Defense of Writing,"&lt;/a&gt; that readers already discount slop on sight, than it is to anything Substack shipped this week.&lt;/p&gt;

&lt;p&gt;I don't think that makes the tool pointless. Most readers aren't reading carefully, most of the time, and a feed moving fast enough rewards that. But it does mean the two approaches are solving different versions of the problem. One scales an ear. The other replaces it.&lt;/p&gt;




&lt;p&gt;I know which one caught real writing and called it fake, and I know which one would have let it through. Diction that matches the idea isn't something a classifier trained on hard negatives is built to notice. It's something you have to actually read for.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>writing</category>
      <category>discuss</category>
    </item>
    <item>
      <title>Substack's New AI Detector Has the Same Blind Spot DEV.to's Did</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 22 Jul 2026 08:28:46 +0000</pubDate>
      <link>https://dev.to/dannwaneri/substacks-new-ai-detector-has-the-same-blind-spot-devtos-did-103j</link>
      <guid>https://dev.to/dannwaneri/substacks-new-ai-detector-has-the-same-blind-spot-devtos-did-103j</guid>
      <description>&lt;p&gt;Substack shipped an AI detector this week. Every post, note, and comment over 100 words can now be scanned through Pangram to see how much of it reads as human or AI. &lt;a href="https://post.substack.com/p/against-claudefishing" rel="noopener noreferrer"&gt;Chris Best's launch post&lt;/a&gt; frames it as giving readers a choice: the platform isn't banning AI use, just surfacing it. Worth knowing going in: Pangram's own data already ranks Substack as the cleanest of the platforms it scans, a fraction of LinkedIn's AI-content rate. They're launching transparency tooling from the platform with the least to hide.&lt;/p&gt;

&lt;p&gt;I read that and thought: I already know exactly how this goes. I lived it on a smaller scale, with a person instead of a platform integration.&lt;/p&gt;




&lt;p&gt;A few months ago I got flagged twice in one day by &lt;a href="https://dev.to/dannwaneri/i-got-flagged-by-sloan-sloan-is-a-guy-i-know-3d0e"&gt;"Sloan,"&lt;/a&gt; DEV.to's moderation-warning system. Not a bot quietly scoring posts in the background. A specific community member, reading articles and running them through GPTZero, then sending the same message a blunt classifier would have sent.&lt;/p&gt;

&lt;p&gt;The two pieces that got flagged were the ones that generated the most technical discussion I'd published all year. Short paragraphs. Named data points. Rhetorical questions doing real work. The features that make an argument land are the same features that read as "AI-shaped" to anyone calibrated to notice them, human or model.&lt;/p&gt;

&lt;p&gt;Write worse, look more human. Write well, get flagged.&lt;/p&gt;

&lt;p&gt;That thread also surfaced the part nobody had a clean answer for: the policy creates a dishonesty incentive. Two equally AI-assisted pieces, equally good — the one with a disclosure gets flagged, because now there's something to catch. The one without doesn't. The system was catching transparency, not AI use.&lt;/p&gt;

&lt;p&gt;And then Marco showed up in the comments. Forty years in tech, writing in his second language, using AI to make sure his Italian didn't flatten into something stiffer than he meant. Same Sloan message. Same classifier verdict. Nothing to do with what the policy was built for.&lt;/p&gt;




&lt;p&gt;Pangram is a real classifier with real engineering behind it: &lt;a href="https://www.pangram.com/research/how-it-works" rel="noopener noreferrer"&gt;hard negative mining against its own false positives, training data deliberately mirrored&lt;/a&gt; so it can't just learn "formal writing = AI." That's more rigor than one guy running GPTZero between article reads. I'll give it that.&lt;/p&gt;

&lt;p&gt;But it inherits the same structural problem Sloan had, because it's answering the same narrow question: does this text look AI-shaped. Not: did a human do the thinking. Chris Best's own post admits as much. Pangram can't tell you whether care went into something, only whether the sentences pattern-match to a machine's output.&lt;/p&gt;

&lt;p&gt;That gap is where Marco lives. Detectors trained without deliberately mirrored data have a documented habit of flagging non-native English writing, since careful, formal phrasing correlates with both AI output and someone translating in their head before they type — enough of a problem that several major universities have stopped letting instructors use AI detectors at all. Pangram claims their mirror-prompt method fixes it. Maybe. Most of the numbers backing that claim trace back to Pangram or a study Pangram commissioned.&lt;/p&gt;

&lt;p&gt;Someone with no stake in the answer already looked. &lt;a href="https://www.theatlantic.com/technology/2026/05/pangram-ai-detection-accuracy/687381/" rel="noopener noreferrer"&gt;The Atlantic's Matteo Wong traced a recent wave of AI-writing accusations back to Pangram itself&lt;/a&gt;, including a horror novel pulled from a major publisher days before its release. His argument wasn't that the tool is broken. It's that a detector that's mostly reliable can be more dangerous than one that's obviously unreliable, because people stop checking. A 99.98% accuracy rate sounds like certainty. Applied across millions of posts, the failures are still real people, still real reputations, just quieter about it.&lt;/p&gt;

&lt;p&gt;That's Marco's risk, and mine, in one sentence: the false positive doesn't feel like a statistic when it's your byline.&lt;/p&gt;




&lt;p&gt;I write from Port Harcourt, in English, the language I was taught in and think in, using AI as part of an actual workflow: not to generate opinions I don't have, but to get from a rough draft to a clean one without losing the argument along the way. Sloan already showed me what a false positive costs, close enough that I don't need to imagine it. Marco is the version of that risk I can't unsee.&lt;/p&gt;

&lt;p&gt;It's also the whole reason I stopped using a generic humanizer and built &lt;a href="https://github.com/dannwaneri/voice-humanizer" rel="noopener noreferrer"&gt;one calibrated to my own published corpus&lt;/a&gt; instead. A tool trained to strip "AI-shaped" patterns from anyone's writing will also strip the parts of your writing that are just yours, an em dash you use structurally, a habit of compressing three examples into two. Voice-humanizer checks against what I actually sound like, not against a mirrored dataset of nobody in particular.&lt;/p&gt;

&lt;p&gt;Sloan and Pangram are both answering "does this look like AI." I don't think that's the question that matters. The question is whether someone can be asked "did you know what you were writing about, and do you stand behind it," and answer yes.&lt;/p&gt;

&lt;p&gt;I do.&lt;/p&gt;

</description>
      <category>devto</category>
      <category>ai</category>
      <category>meta</category>
      <category>discuss</category>
    </item>
    <item>
      <title>A bug in Qwen3-TTS taught me voice is biometric</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Tue, 21 Jul 2026 09:01:28 +0000</pubDate>
      <link>https://dev.to/dannwaneri/a-bug-in-qwen3-tts-taught-me-voice-is-biometric-568o</link>
      <guid>https://dev.to/dannwaneri/a-bug-in-qwen3-tts-taught-me-voice-is-biometric-568o</guid>
      <description>&lt;p&gt;The trained voice cloning model for my project is 50 megabytes. Anyone with those 50 megabytes can convincingly be me on a phone call.&lt;/p&gt;

&lt;p&gt;I did not build the project to prove this point. I built it because every AI voice cloning tool I tried erased my Nigerian accent. ElevenLabs, XTTS, F5-TTS — Western English training data, generic African-accented output. The pipeline that finally worked chains Qwen3-TTS (Alibaba's 1.7B voice cloning model, released January 2026) with a reference clip of my own voice. Runs on free Kaggle GPU. Full write-up at &lt;a href="https://github.com/dannwaneri/naija-voice" rel="noopener noreferrer"&gt;github.com/dannwaneri/naija-voice&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The technical hurdle was small and specific. Qwen3-TTS ships with a hardcoded &lt;code&gt;min_new_tokens=2&lt;/code&gt; at line 2046 of &lt;code&gt;modeling_qwen3_tts.py&lt;/code&gt;. The README documents that you can pass Hugging Face &lt;code&gt;generate()&lt;/code&gt; kwargs, but this line silently overrides &lt;code&gt;min_new_tokens&lt;/code&gt; no matter what you pass. Half the seeds I tried produced mid-sentence truncation because the model was free to emit EOS within a few tokens. Once I patched that line, the pipeline held. Voice consistent across a full minute. Nigerian accent intact.&lt;/p&gt;

&lt;p&gt;Issue #55 on the repo has been open since January, filed by a Korean user reporting the same truncation symptom in a different language. Six months, no maintainer response, no root cause identified. I commented with the diagnosis and opened a PR.&lt;/p&gt;

&lt;p&gt;That is the surface story.&lt;/p&gt;

&lt;h2&gt;
  
  
  The .pth file is me
&lt;/h2&gt;

&lt;p&gt;The trained model weights are 50 megabytes. That file, run through the pipeline, produces audio indistinguishable from my voice. It carries my accent. It handles the way I compress vowels. Anyone with those bytes can generate audio of "me" saying anything. A phone call to my bank. A video message to my mother. A confession to a crime.&lt;/p&gt;

&lt;p&gt;I did not commit the weights to GitHub. The repo has the code and the notebooks. It does not have the file that IS me. This is the first line of the README:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;The trained voice model (&lt;code&gt;models/&lt;/code&gt;) is intentionally not published — a voice fingerprint is biometric data and should not be committed to public repositories. If you fork this, generate your own from your own voice.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;I wrote that line before I understood what it meant. I understand it now.&lt;/p&gt;

&lt;h2&gt;
  
  
  Voice is not like a password
&lt;/h2&gt;

&lt;p&gt;You can rotate a password. If your Gmail password leaks, you change it and the leak is contained. The new password does not carry any history of the old one.&lt;/p&gt;

&lt;p&gt;You cannot rotate a voice. If someone has a working clone of my voice today, that clone works for the rest of my life. Every phone call I make afterward has to compete with a generated version indistinguishable from the real one. I cannot upload a new voice next Tuesday.&lt;/p&gt;

&lt;p&gt;Fingerprints have this same property, but fingerprint capture requires physical proximity. Voice capture requires 60 seconds of clean audio. I have posted 60 seconds of clean audio on the internet. Every podcast host has. Every Twitter Spaces speaker has. Every founder pitching on Loom has.&lt;/p&gt;

&lt;p&gt;The threat model is not exotic. It is a scammer with a phone, my mother's number, and 60 seconds of me talking about AI voice cloning on YouTube.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I do about it
&lt;/h2&gt;

&lt;p&gt;Publishing the weights to a public Hugging Face space would be the standard AI-project move. Star count, forkability, community reach. I am not doing it.&lt;/p&gt;

&lt;p&gt;I documented the pipeline and the training procedure. Anyone who wants their own voice can build it. Anyone who wants my voice specifically has to compromise my laptop.&lt;/p&gt;

&lt;p&gt;That is a small mitigation. It does not defend against the 60 seconds of audio I have already released. It does not defend against a service that offers cheap voice cloning to anyone with a URL. But it means the artifact I control is not the source of the leak.&lt;/p&gt;

&lt;p&gt;If you are building anything voice-related: do not publish the weights of a specific person's voice unless that person has explicitly consented in writing. If you clone your own voice for your own content, keep the weights on hardware you control. If you offer voice cloning as a service, refuse the "clone this pastor" and "clone this politician" requests. One viral misuse case poisons the entire category.&lt;/p&gt;

&lt;p&gt;The tooling to clone voices well is now open source, small enough to run on a free GPU, and documented well enough that a solo developer can figure it out in a weekend. The tooling to defend against voice cloning misuse does not exist yet.&lt;/p&gt;

&lt;p&gt;That is not a call to arms. It is where we are.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>security</category>
      <category>opensource</category>
      <category>python</category>
    </item>
    <item>
      <title>the part of this year I don't put in the commit messages</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Mon, 20 Jul 2026 10:22:22 +0000</pubDate>
      <link>https://dev.to/dannwaneri/the-part-of-this-year-i-dont-put-in-the-commit-messages-l6m</link>
      <guid>https://dev.to/dannwaneri/the-part-of-this-year-i-dont-put-in-the-commit-messages-l6m</guid>
      <description>&lt;p&gt;24 hours. That's how long a contract lasted before someone pulled it because I was running Windows, not Mac. Not my work. My OS.&lt;/p&gt;

&lt;p&gt;It almost broke me. Not in the dramatic way. In the quiet way — the kind where you keep shipping code during the day and don't tell your family anything at night.&lt;/p&gt;

&lt;p&gt;Then came FR8. A fully-funded, three-month residency in Helsinki for young builders. I found it in &lt;a href="https://dev.to/hemapriya_kanagala"&gt;Hemapriya's&lt;/a&gt; Dev Opportunity Radar #2, applied the same day, got a confirmation email. A few days later: "didn't quite align with what we're currently looking for — deep technical, high-conviction bets." I took it well in the replies. Not so well internally.&lt;/p&gt;

&lt;p&gt;The Windows contract was the lowest. Finland was close behind it.&lt;/p&gt;

&lt;p&gt;I kept showing up for the DEV Challenges anyway. Almost every one that's run since I got back. Won OpenClaw. Lost more than I won — a lot more. Those losses did something the wins didn't: they prepared me. I'm in the Global AI Hackathon this month. And the Africa Deep Tech Challenge. I want to win one. Maybe both.&lt;/p&gt;

&lt;p&gt;There was a third one too, closer to home than the other two: the Community Program Manager role at DEV itself. MLH had just acquired DEV and needed someone to run it. Not a writing role. Not a building role. My profile fit because of the platform depth, not despite it.&lt;/p&gt;

&lt;p&gt;Rejected on location. Small team, no infrastructure to hire in Nigeria right now. Not my resume. Not my answers. A payroll line I couldn't fix from my side of the application.&lt;/p&gt;

&lt;p&gt;This isn't even my original account. I had one back in 2019. Lost it. Found my way back here in November 2025 — eight months, not a year, but it's felt longer.&lt;/p&gt;

&lt;p&gt;My family doesn't know any of this. Not the Windows call, not Finland, not the account I lost in 2019. They know I code. They don't know what coding cost this year.&lt;/p&gt;

&lt;p&gt;DEV.to is where the people who do know are. Not because I told the whole story to anyone here. Because writing honestly and showing up did what explaining never could.&lt;/p&gt;

&lt;p&gt;Somewhere in the middle of all of it, one of those pieces — &lt;a href="https://dev.to/dannwaneri/someone-else-pays-for-your-ai-access-5149"&gt;"Someone Else Pays for Your AI Access"&lt;/a&gt; — got picked up in AI Engineer World's Fair Daily Context, &lt;a href="https://dev.to/swyx"&gt;swyx's&lt;/a&gt; curated daily covering the World's Fair conference in San Francisco. I wrote that piece from my phone in Port Harcourt. My mom doesn't know what Daily Context is or who swyx is, or what any of this means. I still want to find a way to show her.&lt;/p&gt;

&lt;p&gt;If you've had a stretch like this — a low point, another one right behind it, and a community that held you without needing the backstory — I see you. I feel you like we're in the same room.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://dev.to/jess"&gt;Jess&lt;/a&gt;, &lt;a href="https://dev.to/ben"&gt;Ben&lt;/a&gt;, the whole DEV team — thank you. You're touching lives in ways I don't think you fully know.&lt;/p&gt;

&lt;p&gt;Here's to the next stretch being kinder. And if it isn't, here's to showing up anyway.&lt;/p&gt;

</description>
      <category>career</category>
      <category>devto</category>
      <category>community</category>
      <category>mentalhealth</category>
    </item>
    <item>
      <title>Building an AI Agent That Knows When Not to Guess (Qwen + MCP)</title>
      <dc:creator>Daniel Nwaneri</dc:creator>
      <pubDate>Wed, 15 Jul 2026 16:23:34 +0000</pubDate>
      <link>https://dev.to/dannwaneri/building-an-ai-agent-that-knows-when-not-to-guess-qwen-mcp-19kl</link>
      <guid>https://dev.to/dannwaneri/building-an-ai-agent-that-knows-when-not-to-guess-qwen-mcp-19kl</guid>
      <description>&lt;p&gt;A payment landed for exactly half an invoice's value. The payer's email matched the customer on file. The reference generated by Paystack — the Stripe-equivalent payment processor across Africa — didn't match anything at all.&lt;/p&gt;

&lt;p&gt;Qwen looked at it and came back with 30% confidence and no invoice named.&lt;/p&gt;

&lt;p&gt;I built &lt;a href="https://github.com/dannwaneri/recona" rel="noopener noreferrer"&gt;Recona&lt;/a&gt; for the Global AI Hackathon Series with Qwen Cloud, deadline July 20, 2026 — an agent that reconciles Paystack payments against open invoices and chases the overdue ones, with no human involved on the easy cases. That transaction wasn't supposed to be the interesting part of the demo. It became the whole point.&lt;/p&gt;




&lt;h2&gt;
  
  
  What Recona does
&lt;/h2&gt;

&lt;p&gt;If you freelance or run a small business taking payments in Nigeria, money lands with a reference like &lt;code&gt;PMT final tunde&lt;/code&gt;, and you spend the evening figuring out which invoice it settles — and which client you forgot to chase. Recona automates both halves. It matches incoming payments against open invoices using Qwen, and it runs a daily collections sweep that drafts and sends increasingly firm reminders as invoices age.&lt;/p&gt;

&lt;p&gt;Cloudflare Workers and D1 handle ingestion and orchestration — signature-verified Paystack webhooks, idempotent against duplicate delivery. Alibaba Cloud SAS runs a Dockerized Node service that holds all the Qwen reasoning, deployed separately from the ingestion layer. The reconciler exposes its matching engine both as REST and as MCP tools — &lt;code&gt;match_transaction_to_invoice&lt;/code&gt;, &lt;code&gt;draft_payment_reminder&lt;/code&gt; — over streamable HTTP. Telegram is the human-in-the-loop surface, because the actual job here is a workflow closing itself, not another dashboard to log into.&lt;/p&gt;

&lt;p&gt;The rule I designed around: the model proposes, deterministic code disposes. Auto-closing an invoice requires exact amount, matching currency, and confidence above a threshold — checked in code after Qwen responds, never trusted from the prompt. The LLM reads the messy handwriting. The calculator authorizes the deposit.&lt;/p&gt;




&lt;h2&gt;
  
  
  What I expected to demo
&lt;/h2&gt;

&lt;p&gt;I had a clean story planned. A client pays half an invoice. Qwen correctly identifies which one it is. My deterministic guard blocks the auto-close anyway, because the amount is wrong. Model is right, code overrules it for safety. Good demo beat.&lt;/p&gt;

&lt;p&gt;That's not what happened.&lt;/p&gt;

&lt;p&gt;I ran the real transaction through the real system — the actual Cloudflare Worker at &lt;code&gt;recon-ingest.fpl-test.workers.dev&lt;/code&gt;, the actual deployed reconciler, the actual Qwen API. I ran it twice: once against the original invoice, once after re-seeding a fresh one at exactly double the payment amount, to rule out a fluke.&lt;/p&gt;

&lt;p&gt;Both times, given a payment that matched an invoice's customer email but was exactly half the amount, with a reference that had zero connection to any invoice number, Qwen returned 30% confidence and no committed invoice ID — even though its own reasoning text named the right invoice by ID. It wasn't wrong. It just wouldn't commit to an answer it didn't have enough signal to support.&lt;/p&gt;

&lt;p&gt;I had a choice: force the demo video to match the script I'd already written, or let it show what the model actually did. I rewrote the narration to match reality.&lt;/p&gt;




&lt;h2&gt;
  
  
  Why the honest version is the better demo
&lt;/h2&gt;

&lt;p&gt;I designed against the failure mode I was worried about — a confident wrong answer sliding past my guards. I didn't design as carefully against the opposite one: a system so wrapped in caution that the model's own certainty never becomes a usable signal, and a human ends up reviewing everything regardless of whether the model actually knew the answer.&lt;/p&gt;

&lt;p&gt;What I saw sits in between. Qwen reasoned out loud about the correct invoice, declined to assert it, and handed a legible number to the orchestration layer — 30%, here's why. That's exactly the kind of thing you can build policy around. My auto-close gate doesn't have to grade whether the model's guess is right. It just has to trust the confidence number Qwen already computed about itself, and default to a human whenever that number is low.&lt;/p&gt;

&lt;p&gt;Don't build your safety layer to catch the model when it's wrong. Build it to treat the model's own uncertainty as a first output, and put your guardrails on that. The alternative requires you to be smarter than the model at judging its own answers. This one just requires the model to be honest about what it doesn't know — and Qwen, in my testing, was.&lt;/p&gt;

&lt;p&gt;A junior hire who's always certain is expensive to trust. One who says "I'm 30% sure, and here's why" is the one you can actually build a process around.&lt;/p&gt;




&lt;p&gt;Repo: &lt;a href="https://github.com/dannwaneri/recona" rel="noopener noreferrer"&gt;github.com/dannwaneri/recona&lt;/a&gt; — MIT licensed. Built for the &lt;a href="https://qwencloud-hackathon.devpost.com/" rel="noopener noreferrer"&gt;Global AI Hackathon Series with Qwen Cloud&lt;/a&gt;, Track 4: Autopilot Agent.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>hackathon</category>
      <category>llm</category>
    </item>
  </channel>
</rss>
