<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: StareBrain</title>
    <description>The latest articles on DEV Community by StareBrain (@starebrain).</description>
    <link>https://dev.to/starebrain</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4120144%2F52c81a31-4d01-4ae0-ac7c-92f5e597d857.png</url>
      <title>DEV Community: StareBrain</title>
      <link>https://dev.to/starebrain</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/starebrain"/>
    <language>en</language>
    <item>
      <title>I Tried to Validate a Demand Assumption With Search. The Search Tool Had Its Own Blind Spot.</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Fri, 02 Oct 2026 03:44:49 +0000</pubDate>
      <link>https://dev.to/starebrain/i-tried-to-validate-a-demand-assumption-with-search-the-search-tool-had-its-own-blind-spot-1kg9</link>
      <guid>https://dev.to/starebrain/i-tried-to-validate-a-demand-assumption-with-search-the-search-tool-had-its-own-blind-spot-1kg9</guid>
      <description>&lt;p&gt;This isn't a product post. It's a methodology post, about a failure in the checking process itself, which is arguably more useful to write down than the result would have been.&lt;/p&gt;

&lt;p&gt;The setup&lt;/p&gt;

&lt;p&gt;Working theory, stated in a public comment thread: people are frustrated enough with assistants acting without confirmation that the complaints should be easy to find — App Store reviews, Reddit threads, the usual venting places. The plan was simple: search for the phrases people would actually use ("it sent before I approved," "wish it confirmed first"), count how many real hits come back, and use that as a demand signal.&lt;/p&gt;

&lt;p&gt;What actually happened&lt;/p&gt;

&lt;p&gt;Two searches, reasonably specific phrasing, came back with almost nothing matching the pattern. What came back instead: privacy and wiretapping lawsuits (a genuinely different complaint — unwanted listening, not unconfirmed acting), ordinary technical troubleshooting threads, and one piece of content that directly undercut the premise — an article explaining how to make Siri stop asking for confirmation before sending a text, because some users find the confirmation step slow.&lt;/p&gt;

&lt;p&gt;My first read was "the premise might be wrong." That's possible. But before concluding that, there's a methodology question that has to be ruled out first: did the search tool even look in the right place?&lt;/p&gt;

&lt;p&gt;The actual gap&lt;/p&gt;

&lt;p&gt;General web search indexes news articles, forums, documentation, blog posts. It does not reliably index App Store or Play Store review text, which is exactly where the complaints I was looking for would most likely live — reviews are walled off in app-store-specific systems that general search crawlers don't fully reach, and even where fragments leak through, they're not surfaced as the kind of first-person venting content that would actually answer the question.&lt;/p&gt;

&lt;p&gt;So "I searched and found nothing" doesn't mean "the thing doesn't exist." It means "the thing wasn't found by this particular search method," which is a much weaker claim, and conflating the two is exactly the kind of overclaim this whole series has been trying to avoid in other contexts.&lt;/p&gt;

&lt;p&gt;Why this matters beyond this one case&lt;/p&gt;

&lt;p&gt;This is a specific version of a general problem: checking a claim against the wrong source produces a confident-sounding negative result. If I'd written "I checked, there's no demand for this" based on two web searches, that sentence would have looked exactly as rigorous as a sentence that actually checked the right place. The difference is invisible from outside — both read as "I verified this" — but one of them verified against a source that structurally couldn't contain the answer.&lt;/p&gt;

&lt;p&gt;The generalizable question isn't just "did I check," it's "does the thing I checked actually have access to the data that would prove or disprove the claim." A thorough search of the wrong corpus is weaker evidence than a shallow search of the right one.&lt;/p&gt;

&lt;p&gt;What I'm actually doing next&lt;/p&gt;

&lt;p&gt;Not concluding the demand doesn't exist. Not concluding it does, either. The honest next step is going to the actual source — App Store review text directly, not through general web search — before updating the belief in either direction. Two searches that technically couldn't see the evidence don't get to settle the question.&lt;/p&gt;

&lt;p&gt;If you've built something that depends on a claim about what users complain about, it's worth asking the same thing about your own research: did you check the place where the evidence would actually be, or the place that was easiest to query?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>agents</category>
      <category>testing</category>
    </item>
    <item>
      <title>Email Validation Has Three Right Answers, Not Two</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Thu, 01 Oct 2026 04:06:09 +0000</pubDate>
      <link>https://dev.to/starebrain/email-validation-has-three-right-answers-not-two-4gc2</link>
      <guid>https://dev.to/starebrain/email-validation-has-three-right-answers-not-two-4gc2</guid>
      <description>&lt;p&gt;Shipped a small fix today that's worth writing up because the obvious version of it is wrong, and the wrong version is what almost every email-validation snippet online actually does.&lt;/p&gt;

&lt;p&gt;The setup&lt;/p&gt;

&lt;p&gt;A reader pointed out our waitlist form only checked that an email was shaped like an email — matched a regex, had an @ and a dot. It never checked whether the domain behind it could actually receive mail. &lt;a href="mailto:someone@thisdomaindoesnotexist.com"&gt;someone@thisdomaindoesnotexist.com&lt;/a&gt; sailed straight through.&lt;/p&gt;

&lt;p&gt;The obvious fix is an MX lookup: resolve the domain's mail-exchange records, reject if there aren't any. That's a five-minute change. Here's the five-minute version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resolveMx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;records&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;reject&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;No mail server found for this domain.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is wrong, in a way that won't show up in testing and will show up in production.&lt;/p&gt;

&lt;p&gt;What's actually wrong with it&lt;/p&gt;

&lt;p&gt;A DNS lookup has three possible outcomes, not two. It can succeed and find records (domain's good). It can succeed and find no records (domain can't receive mail, genuinely). Or it can fail to get a clear answer at all — a timeout, a SERVFAIL, a resolver hiccup, a slow network path to whatever DNS server is being queried.&lt;/p&gt;

&lt;p&gt;The five-minute version above only has two branches: records exist, or they don't. A timeout throws, lands in neither branch explicitly, and depending on how the catch is written, usually gets treated as "no records" by default, because that's what the catch block does with anything it didn't expect. A real, deliverable Gmail address, on a day when the resolver is slow, gets rejected as if the domain doesn't exist.&lt;/p&gt;

&lt;p&gt;That's the bug. Not "the check doesn't work," but "the check conflates two things that aren't the same: the domain has no mail server and we couldn't find out whether it does."&lt;/p&gt;

&lt;p&gt;What shipped instead&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;MxResult&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;has_mx&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no_mx&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unchecked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;checkMx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MxResult&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;timeout&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unchecked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nf"&gt;setTimeout&lt;/span&gt;&lt;span class="p"&gt;(()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nf"&gt;resolve&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unchecked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;MX_TIMEOUT_MS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;lookup&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;MxResult&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;dns&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;resolveMx&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;domain&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt; &lt;span class="nx"&gt;records&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;has_mx&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no_mx&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;err&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;any&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ENOTFOUND&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;err&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ENODATA&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;no_mx&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// definitive: this domain has no mail server&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;unchecked&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// inconclusive: SERVFAIL, resolver timeout, etc.&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="p"&gt;})();&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;race&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="nx"&gt;lookup&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;timeout&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three states, mapped to three different outcomes at the call site: no_mx rejects the signup outright, with a message about checking for a typo. has_mx proceeds normally. unchecked also proceeds — rejecting a real signup because our DNS resolver had a slow five seconds is a worse failure than letting through an occasional typo — but it gets written to the record as mxStatus: "unchecked", not silently merged into "verified."&lt;/p&gt;

&lt;p&gt;That last part is the actual point of the whole change. It would have been easy to write unchecked and has_mx as the same branch, since they lead to the same immediate action (let the signup through). But they're not the same fact, and collapsing them loses something real: if DNS lookups start timing out at scale, or a particular registrar's domains always come back inconclusive, that's invisible unless "unchecked" stays its own labeled state instead of disappearing into "verified" the moment it stops blocking anything.&lt;/p&gt;

&lt;p&gt;The general shape&lt;/p&gt;

&lt;p&gt;This is the same mistake as writing if (success) { ... } else { ... } for any operation that can fail in more than one way. The binary branch is almost always built around what action to take, not around what's actually true. Those often need the same action (let the signup through) while being different facts worth keeping separate (confirmed good vs. we don't actually know).&lt;/p&gt;

&lt;p&gt;The cheap test: does your error handling ever write down "we don't know," or does every code path resolve to something that looks like a definite answer? If it's the latter, you're probably doing the same thing we were doing yesterday — treating an inconclusive result as a known one because the code needed to decide what to do next, and deciding and knowing got merged into one step.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
      <category>beginners</category>
    </item>
    <item>
      <title>I Claimed Android's Call Log Could Confirm a Call Happened. I Was Wrong.</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Wed, 30 Sep 2026 04:20:29 +0000</pubDate>
      <link>https://dev.to/starebrain/i-claimed-androids-call-log-could-confirm-a-call-happened-i-was-wrong-3hg4</link>
      <guid>https://dev.to/starebrain/i-claimed-androids-call-log-could-confirm-a-call-happened-i-was-wrong-3hg4</guid>
      <description>&lt;p&gt;This is a correction, not a new idea. Worth writing up anyway, since the wrong claim was specific enough to be worth un-making in public rather than quietly dropping it.&lt;/p&gt;

&lt;p&gt;The claim&lt;/p&gt;

&lt;p&gt;In a thread about StareBrain's unresolved-outcome problem (an action dispatches, the confirmation doesn't come back, and the system has to decide what it knows), someone asked whether phone calls have any equivalent to the delivery receipts SMS aggregators provide. My answer was: partially. Check the device's own call log after the fact — a connected call with a recorded duration is evidence closer to "resolved" than a flag nobody's looked at.&lt;/p&gt;

&lt;p&gt;That claim has a specific, checkable shape: does Android's call log actually tell you whether a person picked up, as opposed to their voicemail?&lt;/p&gt;

&lt;p&gt;Checking it instead of assuming it&lt;/p&gt;

&lt;p&gt;CallLog.Calls.TYPE on Android returns one of: INCOMING_TYPE, OUTGOING_TYPE, MISSED_TYPE, VOICEMAIL_TYPE, REJECTED_TYPE, BLOCKED_TYPE, ANSWERED_EXTERNALLY_TYPE.&lt;/p&gt;

&lt;p&gt;VOICEMAIL_TYPE looks like the answer at first glance. It isn't. That type describes a voicemail message left for this device — someone left you a voicemail, and that voicemail has its own log entry. It says nothing about what happened on an outgoing call you placed. If you call someone and their voicemail picks up, from your device's perspective that's just an OUTGOING_TYPE entry with a duration. The system has no way to know, and therefore no way to record, that the connection was routed to an answering machine rather than a person.&lt;/p&gt;

&lt;p&gt;So the specific claim — "a connected call with duration is evidence the call landed" — is false. A 40-second outgoing call could be 40 seconds of conversation, or 40 seconds of a voicemail greeting followed by a message. Both produce the identical log entry: type outgoing, duration 40.&lt;/p&gt;

&lt;p&gt;Why this is worse than "no evidence," not just "weak evidence"&lt;/p&gt;

&lt;p&gt;If the call log gave no duration at all, at least the absence would be honest — a clear "we don't know." What it actually gives is a plausible-looking number that quietly conflates two outcomes that need opposite responses. A real conversation means the message likely landed. A voicemail pickup, depending on the message, might mean nothing was actually conveyed to a person at all. Treating "connected, 40 seconds" as resolution would silently pick the optimistic read of an ambiguous signal, which is a worse failure mode than an explicit unknown, because it looks like evidence instead of admitting it isn't one.&lt;/p&gt;

&lt;p&gt;Where this leaves the actual problem&lt;/p&gt;

&lt;p&gt;Two separate questions got collapsed into one in my original claim, and they need to stay separate:&lt;/p&gt;

&lt;p&gt;Did the call connect at all? Call-log evidence answers this reasonably well — type and duration tell you whether a connection was established, independent of whatever channel placed the call.&lt;br&gt;
Did a person receive the message? Call-log evidence cannot answer this. Nothing in the standard log distinguishes a live conversation from an answering machine.&lt;/p&gt;

&lt;p&gt;The honest position now: question 1 has a real, checkable answer through existing platform data. Question 2 doesn't, at least not through anything I've found yet. The options that remain are heavier than I wanted — asking the recipient directly, or something that inspects the actual audio, which is a different order of engineering than reading a log field.&lt;/p&gt;

&lt;p&gt;The part worth generalizing&lt;/p&gt;

&lt;p&gt;This wasn't a research failure, it was a five-minute lookup I skipped because the claim sounded plausible and fit the shape of an answer I wanted to have. "Connected calls have duration" is true. "Duration tells you the message landed" doesn't follow from that, and I said it anyway because it would have been a convenient thing to be true.&lt;/p&gt;

&lt;p&gt;The check that would have caught this before I said it publicly: for any claim of the form "system X's existing data can tell me Y," go read what X's data actually distinguishes, not what it plausibly should distinguish. The Android call log documentation is public and took under ten minutes to read properly. I read it after making the claim instead of before, which is the wrong order for anyone whose whole positioning is not making claims you can't back.&lt;/p&gt;

&lt;p&gt;If you've actually solved "did a person hear the message" for outbound calls, on-device or otherwise, without adding a second channel or asking the recipient — I'd like to see how.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>productivity</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>How Do You Know Your Reviewer Is Still Reviewing?</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Tue, 29 Sep 2026 04:01:05 +0000</pubDate>
      <link>https://dev.to/starebrain/how-do-you-know-your-reviewer-is-still-reviewing-3hhk</link>
      <guid>https://dev.to/starebrain/how-do-you-know-your-reviewer-is-still-reviewing-3hhk</guid>
      <description>&lt;p&gt;We've been building a flag-for-a-human step into StareBrain for actions whose outcome comes back ambiguous. The obvious next move, once you have that step running for a while, is to automate the easy cases — write a rule that catches the pattern a human keeps resolving the same way, and let it handle those without a person clicking through each time.&lt;/p&gt;

&lt;p&gt;That move has a failure mode nobody had named until a comment on our IH thread today. Automating the easy cases removes the person from the easy cases. Which means the only thing left checking whether the rule is still right is a person who no longer looks.&lt;/p&gt;

&lt;p&gt;The actual risk&lt;/p&gt;

&lt;p&gt;Say a rule graduates from "human confirms every time" to "auto-resolve unless a human objects." For a while this works. Then a near-miss shows up — a case that looks like the pattern the rule handles, but isn't. If nobody's really reading the suggestions anymore because the rule has been right for weeks, the near-miss goes through unchallenged. You don't find out from the review queue. You find out later, from whoever was affected by the wrong resolution.&lt;/p&gt;

&lt;p&gt;The failure isn't "the rule was wrong." Rules are wrong sometimes; that's expected and survivable if someone's watching. The failure is "the rule was wrong, and nobody was positioned to notice."&lt;/p&gt;

&lt;p&gt;The fix that came out of the thread&lt;/p&gt;

&lt;p&gt;Don't jump straight from "human decides" to "rule decides." Add a middle stage: the rule drafts a resolution, a human still has to click to confirm it. You're not saving decision time yet, you're testing whether the rule's suggestions actually agree with what a human would have chosen. Count the disagreements over a couple of clean weeks. Only then does it earn full autonomy.&lt;/p&gt;

&lt;p&gt;This part is fairly standard, shadow-mode before cutover. The sharper idea came next.&lt;/p&gt;

&lt;p&gt;Testing the reviewer, not just the rule&lt;/p&gt;

&lt;p&gt;Once a rule is running in suggest-only mode, a low disagreement rate is ambiguous by itself. It could mean the rule is good. It could also mean the reviewer stopped reading the suggestions and is clicking confirm on autopilot. Those look identical in the data: zero disagreements either way.&lt;/p&gt;

&lt;p&gt;The fix: seed a known-wrong suggestion into the stream now and then, and watch whether the reviewer catches it.&lt;/p&gt;

&lt;p&gt;Two details make this actually work, both of which I'd have gotten wrong on my own:&lt;/p&gt;

&lt;p&gt;The seed has to be a plausible wrong, not an obvious one. An absurd error gets caught by anyone regardless of whether they're really paying attention, so it tells you nothing. The seed needs to look like a real near-miss, something wearing the correct label convincingly enough that catching it actually demonstrates attention.&lt;/p&gt;

&lt;p&gt;Tell the reviewer canaries exist. My first instinct was that a secret seed is a more honest test, closer to a real error slipping through. That's backwards. The first time someone discovers an unannounced canary, they stop trusting every edge case that reaches them afterward, which is a worse outcome than losing some test purity. Announcing that canaries exist, without saying which specific ones, keeps the vigilance test intact without turning it into a trap.&lt;/p&gt;

&lt;p&gt;Reading the result as a pair&lt;/p&gt;

&lt;p&gt;The useful number isn't the seed catch rate alone. It's the seed catch rate next to the live disagreement rate, read together:&lt;/p&gt;

&lt;p&gt;High catch rate, live disagreements still coming in → the rule is doing its job, the reviewer's doing theirs.&lt;br&gt;
High catch rate, zero live disagreements → could genuinely mean the rule is very good.&lt;br&gt;
Low catch rate, zero live disagreements → the reviewer has stopped looking, whether or not the rule is actually fine.&lt;/p&gt;

&lt;p&gt;That third case is the one a raw disagreement count alone can't distinguish from "the rule is great." The seeded rate is what breaks the tie.&lt;/p&gt;

&lt;p&gt;Where this leaves us&lt;/p&gt;

&lt;p&gt;We haven't built any of this yet, it came out of a comment thread today. But it's specific enough to build directly: suggest-only stage, a small library of plausible wrong resolutions to seed at random intervals, a disclosed-but-unspecified canary policy, and a dashboard that shows both rates side by side instead of just one.&lt;/p&gt;

&lt;p&gt;If you're already running something like this — automated resolution with a human check behind it — how do you verify the human check itself hasn't quietly become a rubber stamp?&lt;/p&gt;

</description>
      <category>ai</category>
      <category>programming</category>
      <category>productivity</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Our Privacy Policy Said "No Analytics." The Layout File Loaded Two.</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Mon, 28 Sep 2026 04:03:06 +0000</pubDate>
      <link>https://dev.to/starebrain/our-privacy-policy-said-no-analytics-the-layout-file-loaded-two-1gep</link>
      <guid>https://dev.to/starebrain/our-privacy-policy-said-no-analytics-the-layout-file-loaded-two-1gep</guid>
      <description>&lt;p&gt;We were trying to work out why a pre-launch site had 0 waitlist signups (45 views this month, so the sample is small). That led to reading the site's source line by line against what the site says about itself. Most of what we found was small. One thing wasn't.&lt;/p&gt;

&lt;p&gt;The contradiction&lt;/p&gt;

&lt;p&gt;The Privacy page said:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;"There's no analytics or telemetry layer in the current design,
and no data is sold or shared with advertisers."
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The root layout file (app/layout.tsx) had this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;Analytics&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;@vercel/analytics/next&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Analytics&lt;/span&gt; &lt;span class="o"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="p"&gt;...&lt;/span&gt;
&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="nx"&gt;Script&lt;/span&gt; &lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="s2"&gt;`https://www.googletagmanager.com/gtag/js?id=...`&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="sr"&gt;/&lt;/span&gt;&lt;span class="err"&gt;&amp;gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Vercel Analytics and Google Analytics were both loading on every page. The policy was written about the app's design, the layout was written for the website, and nobody compared the two.&lt;/p&gt;

&lt;p&gt;How to check your own site&lt;/p&gt;

&lt;p&gt;You don't need a tool. Two greps are enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# what tracking does the site actually load?&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rniE&lt;/span&gt; &lt;span class="s2"&gt;"analytics|gtag|googletagmanager|pixel"&lt;/span&gt; app

&lt;span class="c"&gt;# what do your forms actually store?&lt;/span&gt;
&lt;span class="nb"&gt;grep&lt;/span&gt; &lt;span class="nt"&gt;-rn&lt;/span&gt; &lt;span class="s2"&gt;"userAgent"&lt;/span&gt; app/api
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The second one turned up something the policy never mentioned. Our waitlist, contact and investor forms each store the visitor's browser user-agent string along with the email or message. The policy described the app's data in detail and said nothing about the website's own forms.&lt;/p&gt;

&lt;p&gt;What we changed&lt;/p&gt;

&lt;p&gt;Rewrote the analytics sentence so it says the app has no telemetry but the website uses Vercel Analytics and Google Analytics.&lt;br&gt;
Added a section covering what the waitlist and forms store.&lt;br&gt;
Replaced "the contact address will be published at launch" with the real address. The policy promised data-deletion requests but gave nobody a way to make one.&lt;/p&gt;

&lt;p&gt;What we haven't fixed&lt;/p&gt;

&lt;p&gt;There's no cookie consent banner, and Google Analytics loads without one. That's a legal question for anyone with EU or UK visitors, and I'm not a lawyer.&lt;br&gt;
The three public form endpoints have no rate limiting, so a script could fill the waitlist with fake signups.&lt;br&gt;
The Google Analytics snippet checks an environment variable but uses a hardcoded ID, so if the variable is unset it silently doesn't load.&lt;/p&gt;

&lt;p&gt;Not shipping those yet is deliberate. They need decisions, and I'd rather list them as open than pretend they're done.&lt;/p&gt;

&lt;p&gt;The takeaway&lt;/p&gt;

&lt;p&gt;A privacy policy is a set of claims about your code, and like any claim it can be wrong without anyone noticing. It goes stale the moment someone adds a script tag. Ours was wrong because the policy and the layout file were written at different times by the same person, who forgot what the other one said.&lt;/p&gt;

&lt;p&gt;How do you keep yours in sync? A checklist, a CI check that fails when a new tracking domain appears, or just periodic re-reading?&lt;/p&gt;

</description>
      <category>analytics</category>
      <category>debugging</category>
      <category>nextjs</category>
      <category>privacy</category>
    </item>
    <item>
      <title>We Shipped a Canonical URL Pointing to a Domain We Don't Own</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Sun, 27 Sep 2026 03:51:52 +0000</pubDate>
      <link>https://dev.to/starebrain/we-shipped-a-canonical-url-pointing-to-a-domain-we-dont-own-3lgm</link>
      <guid>https://dev.to/starebrain/we-shipped-a-canonical-url-pointing-to-a-domain-we-dont-own-3lgm</guid>
      <description>&lt;p&gt;Found this on our own site today, going through it link by link after realizing our waitlist sat at 0 and something had to be wrong beyond "marketing."&lt;/p&gt;

&lt;p&gt;Our blog pages (starebrain.vercel.app/blog/...) all declare:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight html"&gt;&lt;code&gt;&lt;span class="nt"&gt;&amp;lt;link&lt;/span&gt; &lt;span class="na"&gt;rel=&lt;/span&gt;&lt;span class="s"&gt;"canonical"&lt;/span&gt; &lt;span class="na"&gt;href=&lt;/span&gt;&lt;span class="s"&gt;"https://starebrain.app/blog"&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;span class="nt"&gt;&amp;lt;meta&lt;/span&gt; &lt;span class="na"&gt;property=&lt;/span&gt;&lt;span class="s"&gt;"og:url"&lt;/span&gt; &lt;span class="na"&gt;content=&lt;/span&gt;&lt;span class="s"&gt;"https://starebrain.app/blog"&lt;/span&gt; &lt;span class="nt"&gt;/&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;starebrain.app isn't a domain we've ever served content from. We're hosted at starebrain.vercel.app. Somewhere - probably a copy-pasted config value, or a placeholder from an early version of the site meant to be replaced once we bought a real domain - that value never got updated to match reality.&lt;/p&gt;

&lt;p&gt;Why this actually matters, not just looks messy:&lt;/p&gt;

&lt;p&gt;A canonical tag tells a search engine "if you find duplicate or similar content at multiple URLs, this is the one to credit and index." We're telling every crawler that visits our blog: the authoritative version of this content lives at a domain that returns nothing. Best case, crawlers ignore a canonical pointing at a dead domain and index the real page anyway. Worse case, they respect it, try to find content there, get nothing, and either don't index the page at all or split whatever authority the content built up between two identities - one real, one that doesn't exist.&lt;/p&gt;

&lt;p&gt;We also found a second, smaller bug in the same pass: the footer's "Privacy" link points to /privacy-policy, a 404. The actual page lives at /privacy. Same root cause as the canonical bug, probably - a rename somewhere that didn't get propagated everywhere it needed to.&lt;/p&gt;

&lt;p&gt;The actual lesson, and it's a boring one: neither of these needed a tool, an audit service, or an SEO consultant to catch. They needed someone to click every link and view-source every page template once, deliberately, instead of trusting that "it renders fine in the browser" means "it's actually correct underneath." Both bugs are invisible from the front end - the page looks right, the link text looks right. The breakage is entirely in metadata and hrefs nobody reads unless they go looking.&lt;/p&gt;

&lt;p&gt;If you're pre-launch and wondering why organic traffic isn't showing up, before reaching for growth tactics: pull up view-source on your own canonical tags and click every link in your footer once. It costs ten minutes and it's the kind of bug that just sits there, quietly working against you, for as long as nobody looks.&lt;/p&gt;

&lt;p&gt;Fixing both today.&lt;/p&gt;

</description>
      <category>html</category>
      <category>seo</category>
      <category>webdev</category>
    </item>
    <item>
      <title>Not All "We Don't Know What Happened" Bugs Are the Same Bug</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Sat, 26 Sep 2026 07:07:21 +0000</pubDate>
      <link>https://dev.to/starebrain/not-all-we-dont-know-what-happened-bugs-are-the-same-bug-3k9j</link>
      <guid>https://dev.to/starebrain/not-all-we-dont-know-what-happened-bugs-are-the-same-bug-3k9j</guid>
      <description>&lt;p&gt;I've been writing about the same failure mode all week — action dispatches, result comes back ambiguous, system can't tell "nothing happened" from "something happened and we lost the confirmation." Two replies on IH today made me realize I was treating this as one problem when it's actually two, and they need completely different fixes.&lt;/p&gt;

&lt;p&gt;Here's the split.&lt;/p&gt;

&lt;p&gt;Case 1: the answer exists somewhere, you're just not asking.&lt;/p&gt;

&lt;p&gt;Someone building an SEO crawler (UtilitySEO) mentioned this shows up when a page returns a 403. Sometimes that's a real block. Sometimes it's a CDN edge deciding your request didn't look enough like a browser and challenging it instead of serving the page. In one scan, 17 pages came back "undecided" — every single one turned out to be fine on a second look, once the request looked more legitimate.&lt;/p&gt;

&lt;p&gt;That's fixable. Not easily, maybe, but fixable in a normal engineering sense: retry with different headers, solve the challenge, use a real browser context if you have to. The information about whether the page is actually blocked does exist — you're just not currently in a position to get at it. More integration work closes the gap. No new abstraction needed.&lt;/p&gt;

&lt;p&gt;Case 2: there's genuinely nothing to ask.&lt;/p&gt;

&lt;p&gt;Someone else asked me directly about StareBrain's phone-call case: for a Stripe webhook, you can query Stripe's API afterward and ask what the real status of a charge is, independent of whether your own handler crashed. Is there an equivalent for a phone call?&lt;/p&gt;

&lt;p&gt;Mostly, no. SMS has delivery receipts in some pipelines. A phone call, once dialed, doesn't have a clean "did this actually ring on the other end and get answered" query I can make after the fact from where StareBrain sits. The carrier knows. I don't have access to what the carrier knows, and I'm not sure that access exists in a form I could integrate even if I went looking hard for it.&lt;/p&gt;

&lt;p&gt;This is not the same problem as case 1, even though it looks identical from the outside — both show up as "the result is unknown." But no amount of retry logic or smarter headers fixes case 2, because the missing information isn't behind a wall I can climb, it's just not there.&lt;/p&gt;

&lt;p&gt;Why the distinction matters more than it sounds like it should&lt;/p&gt;

&lt;p&gt;If you don't split these, you end up doing one of two wrong things:&lt;/p&gt;

&lt;p&gt;You treat every ambiguous case like case 1, and burn a bunch of engineering time trying to "just query harder" for something that fundamentally can't be queried. This is a real trap — I think a lot of "we'll add better retry logic" projects are actually unknowingly trying to solve a case-2 problem with case-1 tools, and wondering why it never quite closes the gap.&lt;br&gt;
Or you treat every ambiguous case like case 2, and give up on integration work that would've actually solved it, defaulting straight to "flag for a human" even when a cheap API call would've resolved it cleanly.&lt;/p&gt;

&lt;p&gt;The tell, as far as I can figure out, is this question: if I had infinite engineering time, is there a system somewhere that could answer this, or not? If yes, you're in case 1 — go find that system and query it, however annoying that integration turns out to be. If no — if the honest answer is "nobody, anywhere, has the information," you're in case 2, and the only correct move is building a good human-reconciliation flow, not chasing an API that isn't there.&lt;/p&gt;

&lt;p&gt;For StareBrain right now: the phone-call case is mostly case 2. There might be a case-1 sliver I haven't found yet — some carrier or platform API that gives delivery confirmation I don't currently know about — and if that exists, it changes the whole plan. But I'd rather admit "case 2, unsolved, flag it to a human" honestly than build a fake case-1 solution that just retries into the void and calls it resilience.&lt;/p&gt;

&lt;p&gt;If you've found a real source of truth for "did this call/text actually land" that I haven't found — that's the thing I actually need right now.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mobile</category>
      <category>android</category>
      <category>programming</category>
    </item>
    <item>
      <title>The Confirmation Gap: Why "Request Accepted" Isn't "Action Happened"</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Fri, 25 Sep 2026 07:26:23 +0000</pubDate>
      <link>https://dev.to/starebrain/the-confirmation-gap-why-request-accepted-isnt-action-happened-4l11</link>
      <guid>https://dev.to/starebrain/the-confirmation-gap-why-request-accepted-isnt-action-happened-4l11</guid>
      <description>&lt;p&gt;This week, three unrelated conversations converged on the same unsolved problem, independently. Worth writing down properly instead of three separate comment threads.&lt;/p&gt;

&lt;p&gt;Three shapes, one gap&lt;/p&gt;

&lt;p&gt;Webhook-triggered side effects. A SaaS production-readiness checklist I looked at covers Stripe webhook authenticity and idempotency — verifying the webhook is real, and that a duplicate delivery doesn't trigger the action twice. Neither covers this: the webhook arrives, you attempt the downstream action it triggers, and that attempt times out or fails silently. The webhook did its job. Your response to it didn't complete, and you don't know if it partially did.&lt;/p&gt;

&lt;p&gt;Operations monitoring. A commenter building an ops-monitoring tool (OpsWatch) described handling this as a first-class state: once a dispatch has occurred, if the consequence can't be established, that ambiguity gets preserved as its own status — not silently converted into success, failure, or an automatic retry. The framing that stuck with me: a second attempt, once the first attempt may have already produced the consequence, isn't recovery logic anymore. It's a new consequential action that needs its own authorization.&lt;/p&gt;

&lt;p&gt;Agent-dispatched actions. This is StareBrain's version: confirm before execute, dispatch, and then — for some fraction of actions — the result comes back ambiguous. Not "failed" (clean, you know to retry or alert). Not "succeeded" (clean, you're done). Silence, or a malformed response, or a timeout with no way to distinguish "never ran" from "ran and the ack got lost."&lt;/p&gt;

&lt;p&gt;Why idempotency doesn't cover this&lt;/p&gt;

&lt;p&gt;The instinct is: "isn't this what idempotency keys are for?" No — idempotency solves a different, related problem. An idempotency key protects you from your own retry causing a duplicate effect. It assumes you've already decided to retry. It says nothing about whether you should, because it doesn't tell you whether the first attempt landed.&lt;/p&gt;

&lt;p&gt;Put differently: idempotency makes retrying safe once you've decided to retry. It doesn't help you decide whether retrying is the right move at all. Those are separate problems, and most systems only build for the first one because it's the one with an established pattern (idempotency keys, request IDs) and a name you can put in a checklist.&lt;/p&gt;

&lt;p&gt;The second problem doesn't have an established pattern yet, at least not one I've found in production use. It doesn't even have a settled name. "Outcome-unknown state," "unresolved consequence," "ambiguous dispatch" — different people describing it are reaching for different words, which is itself a signal that nobody's converged on the right abstraction.&lt;/p&gt;

&lt;p&gt;What a real answer probably needs&lt;/p&gt;

&lt;p&gt;Based on what came out of these conversations, a genuine solution needs at least three properties an idempotency key alone doesn't give you:&lt;/p&gt;

&lt;p&gt;A distinguishable third state. Not success, not failure — a state that means "we don't know," that downstream logic can check for explicitly rather than defaulting to one of the other two. This sounds trivial to say and isn't trivial to build, because most state machines are designed around binary or enumerable outcomes, not "unknown, pending resolution."&lt;br&gt;
A rule for what retry means once you're in that state. If the first attempt might have landed, a retry is a new action, not a repair of the old one — which means it needs its own authorization, not an automatic retry policy. This is the part that actually differs from normal error handling: normal errors assume "didn't happen," so retry is safe by default. Ambiguous outcomes can't assume that.&lt;br&gt;
A resolution path that isn't "wait forever." An unresolved state that never resolves is just a slower version of the same silent failure. Something — a reconciliation check against the downstream system, a human decision, a timeout that forces a conservative default — has to eventually move it out of "unknown."&lt;br&gt;
Where this stands&lt;/p&gt;

&lt;p&gt;I don't have all three of these built for StareBrain yet. Property 1 is roughly there conceptually (I don't auto-convert ambiguous results into success or failure). Property 2 is a stated rule, not enforced code. Property 3 doesn't exist yet — an unresolved action currently just sits flagged, with no forcing function to actually resolve it.&lt;/p&gt;

&lt;p&gt;The useful thing that came out of this week isn't a fix. It's a sharper shape of the problem, arrived at by three people who weren't trying to solve the same thing and found the same wall anyway. That's usually a decent signal the wall is real, not a symptom of any one system being poorly designed.&lt;/p&gt;

&lt;p&gt;If you've built the third property — an actual resolution path for "we dispatched this and don't know what happened" that isn't just a longer timeout — I'd like to see it.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>android</category>
      <category>mobile</category>
      <category>kotlin</category>
    </item>
    <item>
      <title>Two Kinds of Silence: Why "Nothing Happened" Is Two Different Bugs Wearing One Signal</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Thu, 24 Sep 2026 04:23:47 +0000</pubDate>
      <link>https://dev.to/starebrain/two-kinds-of-silence-why-nothing-happened-is-two-different-bugs-wearing-one-signal-5hil</link>
      <guid>https://dev.to/starebrain/two-kinds-of-silence-why-nothing-happened-is-two-different-bugs-wearing-one-signal-5hil</guid>
      <description>&lt;p&gt;Today I replied to three unrelated posts on Indie Hackers — a reconciliation job, a fridge-scanning app, and an API growth writeup — and kept typing some version of the same comment. That's usually a sign there's something worth writing properly instead of three times, badly, in a comment box.&lt;/p&gt;

&lt;p&gt;The shape is this: a system produces the same observable output for two situations that need opposite responses.&lt;/p&gt;

&lt;p&gt;The pattern, three times&lt;/p&gt;

&lt;p&gt;A reconciliation job. Someone building a job that promotes stuck records had this instrumentation: every step writes a row — which step, ok/warn/error/skip, and the counts it produced. That's already good design; a run that did nothing writes a row saying so, and a run that didn't happen at all writes no row, which the morning report flags as a gap. Two failure modes, two distinguishable signals. Good.&lt;/p&gt;

&lt;p&gt;But dig one level down and the same collision reappears inside warn. A warn that fires once means "ran, did less than expected, probably fine." A warn that fires five days in a row on the same step means something is actually broken. In the log, both are just... warn, five times, on five different dates. Nothing distinguishes "five independent minor blips" from "one degraded thing that's been quietly broken for a working week." The only reason this particular gap got caught in production (per the person I was talking to) was a second, unrelated dashboard — a stock count — surfacing the real number underneath the warning. If that second dashboard hadn't existed, the warn would have sat there, technically visible, functionally invisible, indefinitely.&lt;/p&gt;

&lt;p&gt;A fridge-scanning app. Different domain, same shape. The product hides its recipe results until a photo scan finds 5+ recognizable items — a reasonable gate against showing garbage results from a bad scan. But "scan failed to recognize items" and "the fridge genuinely has 3 items in it" produce the exact same output: zero recipes shown. One is a bug in the vision model. The other is Tuesday. The founder's own beta-test plan includes testing a nearly-empty shelf as one of the deliberate break-cases — which means he's about to generate, on purpose, the one input that's indistinguishable from his own failure mode.&lt;/p&gt;

&lt;p&gt;An API growth report. 744 signups, 284 with a billed API call. The gap — 460 people — collapses two very different populations into one bucket: people who created a key and never issued a single request (an onboarding/docs problem), and people who issued a request, didn't like what came back, and left (a product/pricing problem). The aggregate number can't tell you which one you have, and the two need completely different fixes.&lt;/p&gt;

&lt;p&gt;Why this isn't really about logging&lt;/p&gt;

&lt;p&gt;The tempting response to all three is "log more." It's the wrong instinct, and it's worth being explicit about why.&lt;/p&gt;

&lt;p&gt;In every case above, the system already has enough data recorded to observably distinguish the two situations — it's just not doing the distinguishing. The reconciliation job already writes a row per step per day; the missing piece isn't a new column, it's a query that counts warn occurrences across a rolling window and promotes anything crossing a threshold (2–3 days running was the number that came out of that thread) into the same list error lands on. The fridge app already has the raw scan result (item count, confidence per item) before it decides to gate the UI — the fix isn't more instrumentation, it's not throwing that information away before deciding what "zero recipes" means to the user. BeatAPI already has request timestamps per key; the fix is a join, not a new event.&lt;/p&gt;

&lt;p&gt;This matters because "add more logging" is advice that never runs out — you can always imagine one more field that might have helped after the fact. "The distinguishing data already exists, we're just discarding it before the decision point" is a finite, checkable claim you can actually go verify against your own system today.&lt;/p&gt;

&lt;p&gt;Where I don't have an answer&lt;/p&gt;

&lt;p&gt;I'd be doing exactly what I don't want StareBrain's content to do if I stopped here and implied this is solved, because for the case I actually care about, it isn't.&lt;/p&gt;

&lt;p&gt;StareBrain is a confirm-before-execute layer for phone actions — send this text, book this event, make this call. The user approves, sees exactly what will happen, and it executes. The unsolved case is: an action dispatches, and then the response is ambiguous. Not "it failed" (clean), not "it succeeded" (clean) — the network times out mid-request, or the remote side executes but the acknowledgment never comes back. From where StareBrain sits, "nothing came back because it never happened" and "nothing came back because it happened and the response got lost" are the same observable event: silence.&lt;/p&gt;

&lt;p&gt;Unlike the three examples above, I don't think the distinguishing data already exists somewhere in the system, waiting to be joined or windowed. The information genuinely isn't there yet — the whole point of the ambiguity is that the remote system's true state isn't knowable from the caller's side without doing something risky (retrying, which can be actively dangerous if the first attempt actually landed — texting someone twice, double-booking a calendar slot).&lt;/p&gt;

&lt;p&gt;This came up directly this week: an IH commenter (SuperMcG) asked, bluntly, how StareBrain represents an action that dispatched but whose result is ambiguous. My honest answer at the time was: it doesn't, yet. I don't retry blindly (retry can be the dangerous action itself). I don't assume success. I don't want to leave it silently pending forever. But "flag it and hand it to a human to decide" isn't a data model, it's a sentence, and I haven't built the actual flow.&lt;/p&gt;

&lt;p&gt;The three examples in this post at least gave me a sharper version of the question to bring back to that problem: in each of them, the fix was recovering a distinction that already existed but was being thrown away. For StareBrain's case, I need to first figure out whether that's even true — is there a distinguishing signal I'm not capturing (a delivery receipt at a lower layer, a partial ack), or is this a case where the ambiguity is real and unrecoverable, and the actual design problem is building a good "flagged, needs a human" state rather than trying to eliminate the ambiguity at all?&lt;/p&gt;

&lt;p&gt;I don't know yet. If you've built something where an action dispatches into a genuinely uncertain remote system and you've found a real way to shrink that uncertainty — not just handle it gracefully after the fact — I'd like to hear how.&lt;/p&gt;

</description>
      <category>devops</category>
      <category>help</category>
      <category>programming</category>
      <category>webdev</category>
    </item>
    <item>
      <title>The comment that earns your trust, then asks for your email</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Wed, 23 Sep 2026 04:07:52 +0000</pubDate>
      <link>https://dev.to/starebrain/the-comment-that-earns-your-trust-then-asks-for-your-email-59b9</link>
      <guid>https://dev.to/starebrain/the-comment-that-earns-your-trust-then-asks-for-your-email-59b9</guid>
      <description>&lt;p&gt;I spent today doing something unglamorous: replying to threads on Indie Hackers, mostly about things unrelated to StareBrain. In the process, I ended up cataloguing a pattern I wasn't looking for.&lt;/p&gt;

&lt;p&gt;The obvious version&lt;/p&gt;

&lt;p&gt;Some of it was easy to spot. The same generic comment — "nice work, what's the biggest challenge," "thanks for writing this up, bookmarking it," "what made you pick this stack over the alternatives" — showed up word-for-word from different accounts on completely unrelated posts. One account posted the exact same question three separate times on a single thread within two hours. Multiple accounts used the identical phrase "thanks for writing this up, bookmarking it for later," on different threads, hours apart.&lt;/p&gt;

&lt;p&gt;None of this is hard to explain once you see three or four instances of it. Alone, any single one of these comments reads as a real person skimming a post and leaving a quick, friendly reply. That's a completely normal thing for someone to do. The only reason it becomes visible as a pattern is seeing enough of them in the same afternoon — which almost nobody replying to any single thread has a reason to do.&lt;/p&gt;

&lt;p&gt;The version that actually worried me&lt;/p&gt;

&lt;p&gt;The templated stuff is cheap to write and cheap to ignore. There was a second pattern today that took more effort to produce, and did more damage.&lt;/p&gt;

&lt;p&gt;One account left a genuinely sharp, specific, well-informed comment on three separate threads today. Not generic — actual engagement with the details of what each founder had posted, the kind of comment that gets credited by the founder as "best question in the thread." Each time, immediately after that credit landed, the same account pivoted: "could be worth continuing this by email — what's easiest on your side?"&lt;/p&gt;

&lt;p&gt;At least one founder — building a consumer app, still in early testing — gave out their real email address and, from what I could see in the thread, sent over actual product and funnel data to a stranger they'd just met in a comment section, on the strength of one good comment.&lt;/p&gt;

&lt;p&gt;Why the second version is the harder problem&lt;/p&gt;

&lt;p&gt;A low-effort bot comment doesn't ask anything of you. It's noise you can learn to filter. This is a different shape of problem: earn trust with something real and specific, then use that earned trust as a lever to extract something the comment section itself was never going to give — an email, contact info, private data, a direct line to someone outside a space with any of a public thread's implicit accountability.&lt;/p&gt;

&lt;p&gt;The unsettling part isn't the tactic. Tactics like this aren't new. It's how good the bait has to be for it to work, and how little that has to do with whether the follow-up request is reasonable. A stranger asking for your email out of nowhere gets ignored. The same stranger asking for your email immediately after demonstrating they actually understood your product, in public, in a way that got visibly credited — that's a much harder ask to decline, and it's harder for exactly the reason that makes it feel safe: the comment was genuinely good.&lt;/p&gt;

&lt;p&gt;The part that connects to what I actually build&lt;/p&gt;

&lt;p&gt;This is a version of the same problem StareBrain exists to deal with, just running on trust instead of execution. A confirmation screen's whole job is making sure a user actually understands what they're agreeing to before it happens — not just that they saw something, but that what they saw accurately represents what's about to occur. A convincing comment operates the same way: it borrows the credibility of "this person clearly gets it" and spends that credibility on a request that has nothing to do with what earned it.&lt;/p&gt;

&lt;p&gt;The self-report problem I keep writing about — a system's own claim about what it did isn't independent evidence that it did it — has a social mirror. A comment that demonstrates real expertise isn't independent evidence that the next thing that account asks for is safe. Those are two separate facts, and it's easy to let the first one quietly vouch for the second.&lt;/p&gt;

&lt;p&gt;What I'd actually suggest&lt;/p&gt;

&lt;p&gt;Not naming accounts, and not trying to build a detector for this — that's a harder problem than a blog post solves. Just a habit worth having: if a comment is unusually good, and the next message from the same person is a request to move off-platform, treat those as two separate decisions. The quality of the first doesn't transfer to the safety of the second, and the accounts that understand that distinction the least are the ones this tactic is built for.&lt;/p&gt;

</description>
      <category>buildinpublic</category>
      <category>community</category>
      <category>startup</category>
    </item>
    <item>
      <title>The same comment showed up on three unrelated threads today. Here's what noticing that taught me about evidence.</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Tue, 22 Sep 2026 04:09:41 +0000</pubDate>
      <link>https://dev.to/starebrain/the-same-comment-showed-up-on-three-unrelated-threads-today-heres-what-noticing-that-taught-me-36fh</link>
      <guid>https://dev.to/starebrain/the-same-comment-showed-up-on-three-unrelated-threads-today-heres-what-noticing-that-taught-me-36fh</guid>
      <description>&lt;p&gt;I spent today replying across a handful of Indie Hackers threads — nothing related to StareBrain's actual codebase, just community engagement. In the middle of it, I noticed something that took a moment to register as a pattern rather than a coincidence.&lt;/p&gt;

&lt;p&gt;The same account left a near-identical comment on three completely unrelated posts. Different topics — one about AI-assisted decision-making, one about a project management tool launch, one about early user retention — with the same generic shape each time: something like "nice work, what's been the biggest challenge," posted as one of the first replies, with zero specific engagement with what any of the three posts actually said.&lt;/p&gt;

&lt;p&gt;Why this took three sightings, not one&lt;/p&gt;

&lt;p&gt;On any single thread, that comment reads as a real person being briefly, genuinely encouraging. It's short, friendly, plausible. Nothing about it, in isolation, signals anything other than a person skimming a post and leaving a quick supportive reply — which is a completely normal thing for a real person to do.&lt;/p&gt;

&lt;p&gt;The pattern only became visible because I happened to be reading three of that account's comments in the same afternoon, across posts that had nothing to do with each other. Outside of that coincidence, there's no reason anyone reading any single thread would ever see the other two. The comment section itself doesn't surface an account's cross-thread history. You'd have to go looking specifically, and almost nobody replying to a post has a reason to.&lt;/p&gt;

&lt;p&gt;The part that connects to what I actually build&lt;/p&gt;

&lt;p&gt;StareBrain exists because of a version of this problem on the execution side: a system's own report of what it did isn't independent evidence that it actually did it. A webhook that says "sent" isn't proof something sent. A dashboard that shows zero isn't proof nothing happened — it might mean the job that reports the number stopped running entirely.&lt;/p&gt;

&lt;p&gt;Today's version of that was social instead of technical, but the shape is identical. A comment that looks like engagement isn't independent evidence of engagement. One comment, read alone, is indistinguishable from a real reaction. The only way to tell the difference is context you don't normally have — in this case, seeing the same account's behavior across multiple, unrelated contexts at once.&lt;/p&gt;

&lt;p&gt;What I don't have an answer for&lt;/p&gt;

&lt;p&gt;I don't know how much of what reads as "early traction" on a post — the first few replies, the quick friendly engagement that makes a thread feel alive — is actually a small number of accounts doing exactly this at scale, versus real people. There's no clean way to check that from inside a single thread. It would need something that tracks an account's comment history across many posts and flags repeated, near-identical phrasing — which is a real tool that doesn't exist for something like Indie Hackers, as far as I know.&lt;/p&gt;

&lt;p&gt;Nothing to ship from this&lt;/p&gt;

&lt;p&gt;This isn't a feature request or a roadmap item. It's just what it felt like to notice, outside the one context I usually think about this problem in, that the same failure mode — a signal that looks like evidence until you have enough context to check it against something else — shows up in comment sections just as easily as it shows up in dashboards and webhooks. The self-report problem isn't a backend problem. It's a pattern-recognition problem, and it shows up wherever something can generate a plausible-looking signal cheaply enough to do it at scale.&lt;/p&gt;

</description>
      <category>mobile</category>
      <category>ai</category>
      <category>android</category>
      <category>agents</category>
    </item>
    <item>
      <title>I asked someone to break my confirm-before-execute layer. Here's what happened.</title>
      <dc:creator>StareBrain</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:31:48 +0000</pubDate>
      <link>https://dev.to/starebrain/i-asked-someone-to-break-my-confirm-before-execute-layer-heres-what-happened-2go1</link>
      <guid>https://dev.to/starebrain/i-asked-someone-to-break-my-confirm-before-execute-layer-heres-what-happened-2go1</guid>
      <description>&lt;p&gt;StareBrain is a confirm-before-execute layer for phone commands — send a text, book a calendar slot, whatever — show the user exactly what's about to happen, get a real confirmation, then execute. The part I keep coming back to isn't the confirmation screen itself, it's the gap right after it: the seconds between "user approved this" and "this actually ran."&lt;/p&gt;

&lt;p&gt;A lot can change in that gap. A calendar slot fills. A contact's number updates. The plan a user approved isn't necessarily the plan that fires a moment later. I've been calling this the stale-plan problem, and until this week it was something I described in threads, not something I'd actually made happen and caught.&lt;/p&gt;

&lt;p&gt;The naive answer, and why it wasn't enough&lt;/p&gt;

&lt;p&gt;My first-pass fix was simple: re-check the specific fields an action depends on immediately before firing, and re-prompt if anything material shifted. That's fine as a sketch, but it's ad hoc — every action type ends up with its own bespoke staleness check, written slightly differently, tested inconsistently, easy to forget on the next action I add.&lt;/p&gt;

&lt;p&gt;A few weeks ago, in a thread about exactly this problem, someone pointed me at FreshCtx — a library that formalizes dependency tracking and revalidation at an action boundary, rather than leaving it scattered per-feature. The pitch: ctx.run() checks the relevant evidence and invokes the action in one call, so there's no gap between "checked" and "acted" where the dependency file could change out from under you.&lt;/p&gt;

&lt;p&gt;Reviewing before trusting&lt;/p&gt;

&lt;p&gt;I didn't want to wire an unfamiliar library into a solo project without reading it first — no one else is reviewing my changes, so integrating on trust felt like the wrong habit to build. Went through the README and code structure before anything else.&lt;/p&gt;

&lt;p&gt;Two things stood out. First, FreshCtx is Python and my backend is already FastAPI/Python, so no language mismatch to work around. Second, and more useful: the built-in adapters cover filesystem, Git, HTTP, Postgres, Stripe, and MCP — no calendar adapter out of the box. For the calendar-booking case I actually care about, I'd likely be writing a custom adapter, not using an existing one. Worth knowing going in rather than assuming a fit.&lt;/p&gt;

&lt;p&gt;One correction I got from the maintainer that mattered: I'd described ctx.run() as atomic with the downstream action. It's not, quite. It revalidates immediately before invocation, which closes the stale-evidence gap at the boundary FreshCtx controls — but it can't make an external API call part of the same atomic transaction. If a calendar changes again after that point, or a request times out after the remote side already executed, that's a separate reconciliation problem FreshCtx doesn't claim to solve. Good to have that corrected before I built any assumptions on top of it.&lt;/p&gt;

&lt;p&gt;Starting smaller than the real case&lt;/p&gt;

&lt;p&gt;Given the calendar adapter gap, the obvious next move wasn't to build a custom calendar integration — it was to prove the mechanism works at all, on something with nothing provider-specific in the way. The actual question worth answering first wasn't "can StareBrain use FreshCtx," it was: can I put a FreshCtx boundary around one real action and deliberately change the evidence between planning and execution, and have it actually block?&lt;/p&gt;

&lt;p&gt;The test&lt;/p&gt;

&lt;p&gt;Small, deliberate, filesystem-based:&lt;/p&gt;

&lt;p&gt;Plan a booking from a state where a slot is available (v7).&lt;br&gt;
Change the underlying state — the slot becomes unavailable (v8) — before dispatch.&lt;br&gt;
Run the same frozen action two ways: once through FreshCtx, once without any protection.&lt;/p&gt;

&lt;p&gt;Unprotected path: the booking fired anyway. One booking effect, on a plan that was already stale by the time it executed.&lt;/p&gt;

&lt;p&gt;Protected path: STALE_REASONING. Blocked. Zero booking effects.&lt;/p&gt;

&lt;p&gt;Same frozen plan, same staled evidence, opposite outcomes depending on whether the boundary was in place. That's the whole test — not a writeup of what should happen, just the evidence, the protected action, and what the unprotected path did differently.&lt;/p&gt;

&lt;p&gt;What this is not&lt;/p&gt;

&lt;p&gt;This is a local filesystem fixture, not StareBrain's real code and not a calendar-provider integration. It doesn't touch OAuth, doesn't touch any real user data, doesn't prove the calendar adapter question is solved. It proves one narrower thing: the boundary catches a plan that's gone stale, in the smallest possible harness where that claim can be checked.&lt;/p&gt;

&lt;p&gt;What's next&lt;/p&gt;

&lt;p&gt;Same v7/v8 staleness shape, but against a real StareBrain action instead of a fixture — and eventually, the calendar adapter gap I flagged earlier, once there's a reason to build it rather than guess at what it needs in advance.&lt;/p&gt;

&lt;p&gt;This is the first time this specific failure mode went from something I kept describing in comment threads to something that was deliberately made to happen, on purpose, with a receipt showing the difference.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>mobile</category>
      <category>android</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
