<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: FetchSandbox</title>
    <description>The latest articles on DEV Community by FetchSandbox (@fetchsandbox).</description>
    <link>https://dev.to/fetchsandbox</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3830177%2Fe63c5a3c-3f48-4ae2-ad6b-ed510712a396.png</url>
      <title>DEV Community: FetchSandbox</title>
      <link>https://dev.to/fetchsandbox</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/fetchsandbox"/>
    <language>en</language>
    <item>
      <title>Lovable Builders Break One Action On Purpose. Nobody Named It.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Thu, 01 Oct 2026 16:46:15 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/lovable-builders-break-one-action-on-purpose-nobody-named-it-2bgh</link>
      <guid>https://dev.to/fetchsandbox/lovable-builders-break-one-action-on-purpose-nobody-named-it-2bgh</guid>
      <description>&lt;p&gt;&lt;em&gt;In 47 posts, three builders describe failure testing before launch.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Lovable builders are already testing failures by hand. I read 47 posts from r/lovable and found three that point at the same habit: break an important action on purpose before launch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What breaks after the happy path
&lt;/h2&gt;

&lt;p&gt;The posts describe a familiar gap. A webhook route looks fine, then takes days to debug when it fails. A new app's error logs make sense to the person who wrote the backend, but not to the person trying to help a customer. From the outside, a failed payment or missing email can look like a confusing screen, not a broken integration.&lt;/p&gt;

&lt;p&gt;Generated apps usually exercise the success case while they are being built. Production adds a second webhook delivery, a payment that succeeds while the app's membership row stays free, or an email request that returns 200 but later bounces. None of those is answered by asking whether the page loaded.&lt;/p&gt;

&lt;h2&gt;
  
  
  Break the action that can fail quietly
&lt;/h2&gt;

&lt;p&gt;Choose the action that would cost you money or trust if it failed without an obvious error—often checkout or its confirmation email. Force a failure and inspect what the customer sees and what the database stores.&lt;/p&gt;

&lt;p&gt;Two useful first checks:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Deliver the same webhook twice. Does the app provision twice?&lt;/li&gt;
&lt;li&gt;Let payment succeed, then make the follow-up write fail. Does the customer remain marked free? &lt;a href="https://fetchsandbox.com/blog/lovable-stripe-paid-still-free" rel="noopener noreferrer"&gt;That mismatch has its own failure pattern.&lt;/a&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;People rebuild these checks by hand after an agent rewrites the backend. The hosted &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt; can start an app-level check: configure the returned twins, arm a returned probe with the session ID, exercise the reachable app, and submit the session ID, run ID, and probe. A green provider quickrun alone does not test the Lovable app.&lt;/p&gt;

&lt;p&gt;When code moves into a repository, carry applicable deterministic checks into CI and attach the receipt to the change. A reviewer can then inspect what ran after the builder chat is gone. Anything the run did not measure should remain unmeasured.&lt;/p&gt;

&lt;p&gt;The full &lt;a href="https://fetchsandbox.com/blog/lovable-failure-testing-47-reddit-posts" rel="noopener noreferrer"&gt;47-post write-up&lt;/a&gt; lists the titles and limits. This is a small sample biased toward people describing problems, not a measure of how often failures happen. One record was an unrelated game ad; I did not interview the posters. The useful signal is narrower: builders are already asking how to break the important path before users find it.&lt;/p&gt;

</description>
      <category>lovable</category>
      <category>testing</category>
      <category>webhooks</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Replit's Green Twin Run Never Hit Its App</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Thu, 01 Oct 2026 16:38:36 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/replits-green-twin-run-never-hit-its-app-2kl2</link>
      <guid>https://dev.to/fetchsandbox/replits-green-twin-run-never-hit-its-app-2kl2</guid>
      <description>&lt;p&gt;&lt;em&gt;A green quickrun can test the twin while leaving your app untested.&lt;/em&gt;&lt;/p&gt;

&lt;p&gt;Replit says the Paddle checkout works. The confirmation email appears. A green quickrun can still leave the app untested.&lt;/p&gt;

&lt;p&gt;That is because the quickrun exercises a FetchSandbox workflow against a provider twin. It does not call the Replit checkout, webhook handler, database, or email path. You can have a healthy twin and an app that still grants access after a declined payment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the app part of the test
&lt;/h2&gt;

&lt;p&gt;Connect FetchSandbox through Replit's hosted MCP URL, &lt;a href="https://fetchsandbox.com/mcp/v1" rel="noopener noreferrer"&gt;https://fetchsandbox.com/mcp/v1&lt;/a&gt;, with a Bearer key. Then ask the agent to call validate_integration for Paddle and Resend with the app's reachable URL. That starts an app-level session and returns the twin configuration and available probes for that run.&lt;/p&gt;

&lt;p&gt;The Replit app must actually use those temporary twin hosts from its development or disposable test configuration. Keep live provider keys out of this run. If the test needs to deliver a webhook, the app's webhook route must be reachable from FetchSandbox; a private preview that redirects to a login page is not reachable evidence.&lt;/p&gt;

&lt;p&gt;Next, arm one probe the session returned, exercise the matching action in the app, and submit that session_id, the run_id returned by arming, and the probe back to FetchSandbox. That links the provider request to what the running app did. A green quickrun is useful for checking the twin workflow, but it is not a substitute for this path.&lt;/p&gt;

&lt;p&gt;Inspect the run's delivery attempts and checkpoints. A redirect, signature rejection, or missing app observation means the relevant check did not run; repeating the same setup does not turn it into a pass. Fix reachability or configuration first, then start a fresh run when its temporary session has expired.&lt;/p&gt;

&lt;p&gt;Try this instruction in Replit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Use FetchSandbox MCP validate_integration for Paddle and Resend
against this app's reachable test URL. Configure only the returned
temporary twins in the test environment. Show the session_id and probes.
Arm one with session_id, exercise the app flow, then submit session_id,
run_id, and probe. Do not call quickrun an app test, and report every unmeasured
check plainly.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the app unlocks a membership after a failed payment, inspect the membership state—not just the HTTP 200. If a receipt rule is unmeasured, the run did not establish that behavior. A useful &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;MCP workflow&lt;/a&gt; makes that difference visible; it cannot turn missing evidence into a pass.&lt;/p&gt;

&lt;p&gt;When the code moves into a repository, carry the relevant deterministic checks into CI and attach the run receipt to the change. The &lt;a href="https://fetchsandbox.com/agent-api-workflow-testing" rel="noopener noreferrer"&gt;agent-to-pipeline workflow&lt;/a&gt; makes the result reviewable after the Replit chat is gone.&lt;/p&gt;

&lt;p&gt;The rule is simple: a green twin workflow proves the twin path; only observed app traffic and measured app behavior prove the app checks that ran.&lt;/p&gt;

</description>
      <category>mcp</category>
      <category>replit</category>
      <category>paddle</category>
      <category>resend</category>
    </item>
    <item>
      <title>Lovable's MCP Form Has a None Option. Don't Pick It.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Tue, 22 Sep 2026 15:58:09 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/lovables-mcp-form-has-a-none-option-dont-pick-it-2b2i</link>
      <guid>https://dev.to/fetchsandbox/lovables-mcp-form-has-a-none-option-dont-pick-it-2b2i</guid>
      <description>&lt;p&gt;Lovable will let you add a custom MCP server with &lt;strong&gt;No authentication&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That option is a trap for FetchSandbox. The hosted endpoint at &lt;code&gt;https://fetchsandbox.com/mcp/v1&lt;/code&gt; forwards &lt;code&gt;Authorization&lt;/code&gt; to the backend. Empty bearer is a 401. The 401's job is to send you to &lt;code&gt;/keys&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get the key first
&lt;/h2&gt;

&lt;p&gt;Sign in at &lt;a href="https://fetchsandbox.com/keys" rel="noopener noreferrer"&gt;fetchsandbox.com/keys&lt;/a&gt;. Hit &lt;strong&gt;Create another key&lt;/strong&gt;. Copy it once. The list later only shows &lt;code&gt;fsk_…&lt;/code&gt; prefixes.&lt;/p&gt;

&lt;p&gt;Do not paste it into Lovable chat.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then the form
&lt;/h2&gt;

&lt;p&gt;Connectors → &lt;strong&gt;+&lt;/strong&gt; → MCP server:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Server name: &lt;code&gt;FetchSandbox&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Server URL: &lt;code&gt;https://fetchsandbox.com/mcp/v1&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Authentication: &lt;strong&gt;Bearer token or API key&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Token: the secret you just copied&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Leave OAuth alone. FetchSandbox does not complete that dance from this form.&lt;/p&gt;

&lt;p&gt;Chat connectors stay on your account. They are not in the published app.&lt;/p&gt;

&lt;h2&gt;
  
  
  Then make it misbehave
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Spin up a Stripe test environment and run a payment through it.
Arm webhook_retries and run checkout again.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 200 on the first charge is the easy part. The question is whether the same event id credits twice.&lt;/p&gt;

&lt;p&gt;The screenshots, the JWT 401 path, and the "paid but still free" read-back are in the &lt;a href="https://fetchsandbox.com/blog/test-lovable-integration-custom-mcp-connector" rel="noopener noreferrer"&gt;canonical guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>lovable</category>
      <category>mcp</category>
      <category>stripe</category>
      <category>webhooks</category>
    </item>
    <item>
      <title>Deterministic Evals for Coding Agents: Prove the Fix</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Tue, 15 Sep 2026 12:22:28 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/deterministic-evals-for-coding-agents-prove-the-fix-2m01</link>
      <guid>https://dev.to/fetchsandbox/deterministic-evals-for-coding-agents-prove-the-fix-2m01</guid>
      <description>&lt;p&gt;Your coding agent says it fixed the retry bug. The diff looks reasonable, its new test passes, and an LLM reviewer gives the patch a high score. None of those results proves that a duplicate webhook can no longer charge the customer twice.&lt;/p&gt;

&lt;p&gt;The missing piece is a deterministic eval that makes the failure happen on the old code, runs the same failure against the proposed fix, and lets measured behavior decide the verdict.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is a deterministic eval engine for coding agents?
&lt;/h2&gt;

&lt;p&gt;A deterministic eval engine for coding agents executes a defined invariant against real code and calculates the result with fixed logic. Given the same code, scenario, and observations, it returns the same verdict. A model can help discover a bug or propose a test, but it does not get to turn the final run green.&lt;/p&gt;

&lt;p&gt;For a bug fix, the useful evaluation is differential:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start with the current, still-broken working tree.&lt;/li&gt;
&lt;li&gt;Apply the proposed patch to a separate copy.&lt;/li&gt;
&lt;li&gt;Inject the same provider failure into both copies.&lt;/li&gt;
&lt;li&gt;Run the same executable probe against both.&lt;/li&gt;
&lt;li&gt;Require the old code to violate the invariant.&lt;/li&gt;
&lt;li&gt;Require the patched code to hold the invariant.&lt;/li&gt;
&lt;li&gt;Record the observations in a receipt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;FetchSandbox implements this pattern for third-party API integrations. It is a verification layer between the coding agent's claim and the pull request.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is an LLM judge not enough to verify a code fix?
&lt;/h2&gt;

&lt;p&gt;An LLM judge is useful when the requirement is subjective: Is the explanation clear? Is the design unnecessarily complex? Does the patch appear maintainable?&lt;/p&gt;

&lt;p&gt;An integration invariant is usually not subjective.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One payment event should create one entitlement.&lt;/li&gt;
&lt;li&gt;A replayed webhook should not repeat a charge.&lt;/li&gt;
&lt;li&gt;A provider &lt;code&gt;409&lt;/code&gt; should not be swallowed and reported as success.&lt;/li&gt;
&lt;li&gt;A lost response should not make a retry create the resource twice.&lt;/li&gt;
&lt;li&gt;A stale event should not overwrite newer local state.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For those properties, a probability or prose assessment is weaker than an observation. The relevant question is not whether another model likes the patch. It is whether the required state held after the failure actually ran.&lt;/p&gt;

&lt;p&gt;Anthropic's guidance on &lt;a href="https://www.anthropic.com/engineering/demystifying-evals-for-ai-agents" rel="noopener noreferrer"&gt;evaluating AI agents&lt;/a&gt; makes the same grader distinction: use deterministic graders where correctness can be checked by code, and reserve model-based graders for criteria that require judgment.&lt;/p&gt;

&lt;p&gt;That separation matters in production. If generation and final grading are both model decisions, the verification layer inherits model variance. If execution produces a stable signal and pure logic maps that signal to a verdict, CI can rely on it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does the exit-code flip decide whether a fix is proven?
&lt;/h2&gt;

&lt;p&gt;FetchSandbox probes use a small exit-code contract:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;0&lt;/code&gt; means the invariant &lt;strong&gt;HELD&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;1&lt;/code&gt; means the invariant was &lt;strong&gt;VIOLATED&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;2&lt;/code&gt; or another harness error means &lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;a probe that did not run is also &lt;strong&gt;INCONCLUSIVE&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Only one measured transition can produce a green verdict:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;buggy tree   1  VIOLATED
fixed tree   0  HELD
result          PROVEN
green_allowed   true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every other result declines or rejects the proof:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;0 → 0&lt;/code&gt;: the supposed bug did not reproduce, so the patch is not proven&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;1 → 1&lt;/code&gt;: the bug reproduced, but the fix is incomplete&lt;/li&gt;
&lt;li&gt;either side inconclusive: the engine could not establish the claim&lt;/li&gt;
&lt;li&gt;self-reported before-and-after results: useful context, but never independently proven&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This gate is deliberately boring. It is a pure function of the two measured runs. Reference simulations, agent confidence, a persuasive explanation, or a passing unrelated test cannot change &lt;code&gt;green_allowed&lt;/code&gt; to &lt;code&gt;true&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why must the same failure run on old and fixed code?
&lt;/h2&gt;

&lt;p&gt;Running only the patched code proves less than it appears to.&lt;/p&gt;

&lt;p&gt;Suppose a coding agent adds a test for a duplicate webhook and the test passes. Several explanations are still possible:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the test never triggered a duplicate delivery&lt;/li&gt;
&lt;li&gt;the handler never booted&lt;/li&gt;
&lt;li&gt;the mock skipped the provider behavior&lt;/li&gt;
&lt;li&gt;the assertion checked only the HTTP response&lt;/li&gt;
&lt;li&gt;the patch suppressed every side effect, including the one the customer paid for&lt;/li&gt;
&lt;li&gt;the bug was never present in the tested path&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The old-code run is the negative control. It demonstrates that the scenario can detect the claimed bug in this repository. The fixed-code run then demonstrates that the same detector no longer finds it after the proposed change.&lt;/p&gt;

&lt;p&gt;Without the first leg, “pass” can mean “the check is blind.” Without the second, there is no fix. The pair is what turns a test result into differential evidence.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does FetchSandbox inject integration failures?
&lt;/h2&gt;

&lt;p&gt;FetchSandbox runs integration code against a stateful sandbox of the services it talks to. A scenario can deliberately introduce behavior that happy-path mocks often omit, including duplicate webhook delivery, &lt;code&gt;409&lt;/code&gt; conflict responses, retries, and observable repeated side effects.&lt;/p&gt;

&lt;p&gt;Consider a payment webhook handler that deduplicates on a delivery ID instead of the provider's stable event ID. A retry arrives with a new delivery ID, so the application treats one event as two and grants or charges twice.&lt;/p&gt;

&lt;p&gt;The executable invariant is not:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;The handler returned 200.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;One provider event creates one business side effect,
even when delivery is attempted more than once.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The probe drives the duplicate delivery, observes the application state, and exits &lt;code&gt;1&lt;/code&gt; if the side effect accumulates. The proposed patch is applied to a copy, the same delivery sequence runs again, and the probe exits &lt;code&gt;0&lt;/code&gt; only when the invariant holds.&lt;/p&gt;

&lt;p&gt;FetchSandbox also has a zero-effect backstop for this class of proof. A patch that stops the duplicate by granting or charging nothing does not get credit for correctness.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is there an LLM anywhere in the FetchSandbox eval path?
&lt;/h2&gt;

&lt;p&gt;The final verdict is not an LLM opinion. Once both probe runs complete, fixed logic classifies their exit codes and decides whether the transition qualifies for green.&lt;/p&gt;

&lt;p&gt;The scenario tier is a separate concern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;curated scenarios use reviewed, hand-authored probes&lt;/li&gt;
&lt;li&gt;declared scenarios use an integration's explicit invariant definition&lt;/li&gt;
&lt;li&gt;a novel bug class can use an LLM to propose a candidate probe&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A generated probe is not trusted because a model wrote it. It must first reproduce the failure on the buggy code, execute against both trees, state the invariant it checked, and survive qualification before it can be reused. If FetchSandbox cannot produce a probe that runs and reproduces the bug, the result stays unproven.&lt;/p&gt;

&lt;p&gt;This is the important boundary: models may help propose what to execute; models do not decide whether the measured execution passed.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a coding agent send to &lt;code&gt;prove_fix&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;The &lt;code&gt;prove_fix&lt;/code&gt; MCP tool needs the still-broken project and the proposed unified diff. Call it before applying the patch to the working tree.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;find_bugs
  → fix_bug returns a proposed diff
  → prove_fix receives broken tree + diff
  → FetchSandbox applies diff to a copy
  → same scenario runs against both trees
  → deterministic gate returns proven, rejected, or unproven
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A successful response includes machine-readable fields such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"proven"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"green_allowed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reproduced"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verified"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"receipt_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://fetchsandbox.com/runs/fix-...?flow=..."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the agent has already applied the patch, the original failing state is gone from the workspace. Reverting mentally or writing a new fixed-only test is not equivalent. The tool needs the real broken tree to establish the first half of the proof.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is in the proof receipt?
&lt;/h2&gt;

&lt;p&gt;The receipt is the review artifact for the run. It records the scenario, the before-and-after outcome, the invariant, and the measured verdict at a shareable URL.&lt;/p&gt;

&lt;p&gt;Here is a &lt;a href="https://fetchsandbox.com/runs/fix-417c74290218?flow=9c61192f9179" rel="noopener noreferrer"&gt;measured FetchSandbox proof receipt&lt;/a&gt; from a Paddle entitlement bug. The first proposed patch was rejected as incomplete. A later patch reproduced the bug on the old tree and held the invariant on the fixed tree, producing the required &lt;code&gt;1 → 0&lt;/code&gt; transition.&lt;/p&gt;

&lt;p&gt;A reviewer can paste the receipt into the pull request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Integration proof&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Failure: duplicate webhook repeats a paid side effect
&lt;span class="p"&gt;-&lt;/span&gt; Invariant: one provider event creates one side effect
&lt;span class="p"&gt;-&lt;/span&gt; Before patch: VIOLATED
&lt;span class="p"&gt;-&lt;/span&gt; After patch: HELD
&lt;span class="p"&gt;-&lt;/span&gt; Verdict: proven
&lt;span class="p"&gt;-&lt;/span&gt; Receipt: https://fetchsandbox.com/runs/fix-...?flow=...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The receipt does not replace review. It removes one factual question from review: whether the named integration failure reproduced before the patch and stopped after it.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does the proof fit into PR review and CI/CD?
&lt;/h2&gt;

&lt;p&gt;Use the proof receipt during pull-request review, then keep the corresponding provider workflow in CI.&lt;/p&gt;

&lt;p&gt;The two checks serve different moments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prove_fix&lt;/code&gt; evaluates a proposed patch against a reproduced bug and returns the before-and-after receipt&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;fetchsandbox run --all --json&lt;/code&gt; runs the configured integration workflows in a pipeline and exits non-zero when a workflow fails&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The current shipped workflow is to paste the &lt;code&gt;prove_fix&lt;/code&gt; receipt into the PR manually. FetchSandbox does not yet post GitHub comments automatically. In CI, store the JSON workflow output as an artifact and make the job a required check when its coverage matches the branch's risk.&lt;/p&gt;

&lt;p&gt;This creates a useful chain of evidence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agent proposes code
  → sandbox reproduces the failure
  → patch flips VIOLATED to HELD
  → receipt travels with the PR
  → workflow remains gated in CI/CD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent remains fast. The reviewer gets a falsifiable result instead of a longer assurance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Does FetchSandbox evaluate every kind of coding task?
&lt;/h2&gt;

&lt;p&gt;No. FetchSandbox is focused on code that integrates with third-party APIs and on failure modes its scenarios and probes can observe. The measured real-code runner currently boots Node.js and Python projects.&lt;/p&gt;

&lt;p&gt;It does not prove subjective code quality, migration safety, authorization design, or every requirement in a pull request. Those still need static checks, ordinary test suites, security review, and human judgment.&lt;/p&gt;

&lt;p&gt;It also does not turn missing coverage into success. If the app cannot boot, the scenario cannot reproduce the bug, or the probe cannot judge the outcome, the verdict remains unproven.&lt;/p&gt;

&lt;p&gt;That refusal is part of the product. A deterministic eval engine is valuable not because it always says yes, but because nothing except the required measured transition can make it say yes.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I use FetchSandbox from a coding agent?
&lt;/h2&gt;

&lt;p&gt;Connect &lt;a href="https://dev.to/mcp"&gt;FetchSandbox MCP&lt;/a&gt; to Claude Code, Cursor, Claude Desktop, Codex, or another MCP-compatible client. The agent can discover integrations, run stateful workflows, inject failures, investigate a bug, propose a patch, and call &lt;code&gt;prove_fix&lt;/code&gt; before editing the working tree.&lt;/p&gt;

&lt;p&gt;For the broader setup, see &lt;a href="https://dev.to/agent-api-workflow-testing"&gt;agent API workflow testing&lt;/a&gt; and &lt;a href="https://dev.to/api-integration-testing-in-ci"&gt;API integration testing in CI&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions about deterministic coding-agent evals
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the difference between a coding-agent benchmark and a fix proof?
&lt;/h3&gt;

&lt;p&gt;A benchmark compares agent performance across a task set. A fix proof evaluates one concrete claim in one repository: the named failure occurred before this patch and stopped after it under the same executable scenario.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can the same coding agent generate the fix and request the proof?
&lt;/h3&gt;

&lt;p&gt;Yes. Independence comes from the external scenario, real execution, and deterministic gate. The agent can request the run, but it cannot set &lt;code&gt;green_allowed&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why not trust a test that the coding agent wrote?
&lt;/h3&gt;

&lt;p&gt;Agent-written tests can be useful, but they can repeat the same mistaken assumption as the patch. The stronger pattern includes a negative control: the probe must catch the bug on the old code before its pass on the new code counts.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens if the old code passes?
&lt;/h3&gt;

&lt;p&gt;The result is &lt;code&gt;not_reproduced&lt;/code&gt;, not proven. A &lt;code&gt;0 → 0&lt;/code&gt; result shows that the probe did not establish the claimed before-and-after change.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens if the runner or probe crashes?
&lt;/h3&gt;

&lt;p&gt;The run is inconclusive and green is blocked. A harness error is not evidence that the application violated or held the business invariant.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can an LLM-generated probe produce a green result?
&lt;/h3&gt;

&lt;p&gt;It can only contribute a candidate executable check. The candidate must run against real code, reproduce the old failure, and observe the fixed invariant. The deterministic gate, not the model, decides the result.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>mcp</category>
      <category>ci</category>
    </item>
    <item>
      <title>RevenueCat Returned 200. The Wallet Still Went Backward.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Tue, 15 Sep 2026 12:19:29 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/revenuecat-returned-200-the-wallet-still-went-backward-4klf</link>
      <guid>https://dev.to/fetchsandbox/revenuecat-returned-200-the-wallet-still-went-backward-4klf</guid>
      <description>&lt;p&gt;A paying customer cancelled, then recovered. RevenueCat sent both events. The handler verified the signatures and returned &lt;code&gt;200&lt;/code&gt; twice.&lt;/p&gt;

&lt;p&gt;The wallet ended on the cancelled state.&lt;/p&gt;

&lt;p&gt;The older &lt;code&gt;EXPIRATION&lt;/code&gt; had been retried after the newer &lt;code&gt;RENEWAL&lt;/code&gt;. Arrival order was not event order, and the last write won.&lt;/p&gt;

&lt;p&gt;Someone in an &lt;a href="https://www.reddit.com/r/indiehackers/comments/1vrbe6m/" rel="noopener noreferrer"&gt;Indie Hackers thread&lt;/a&gt; put the same bug in one sentence:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;My RevenueCat webhooks write to a wallet balance, and the real failures were never bad code, just two events landing out of order.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  What the handler assumed
&lt;/h2&gt;

&lt;p&gt;The code looked ordinary:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/webhooks/revenuecat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;verifyRevenueCat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;RENEWAL&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;credits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;100&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;entitled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;EXPIRATION&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;credits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;entitled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendStatus&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It treated each POST as current truth. That only works if HTTP arrival order matches the subscription lifecycle.&lt;/p&gt;

&lt;p&gt;RevenueCat does not promise that. Failed deliveries retry after about 5, 10, 20, 40, and 80 minutes. A retry reuses &lt;code&gt;event.id&lt;/code&gt; and &lt;code&gt;event_timestamp_ms&lt;/code&gt;. It recomputes the signature timestamp &lt;code&gt;t&lt;/code&gt; for that HTTP attempt.&lt;/p&gt;

&lt;p&gt;So this is a legal sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10:00  EXPIRATION happens
10:01  RENEWAL recovery happens
10:02  RENEWAL webhook arrives     → credits = 100
11:22  EXPIRATION retry arrives    → credits = 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both signatures can be valid. Both handlers can return 200. The customer is locked out because the stale write landed last.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three clocks that disagree
&lt;/h2&gt;

&lt;p&gt;Do not sort RevenueCat events by whichever timestamp is closest to the request.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Field&lt;/th&gt;
&lt;th&gt;What it actually is&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;purchased_at_ms&lt;/code&gt; / expiration dates&lt;/td&gt;
&lt;td&gt;When the store action happened&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;event_timestamp_ms&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;When RevenueCat generated the event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;signature &lt;code&gt;t&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;When this delivery attempt was signed&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;t&lt;/code&gt; is the worst field to use as business order. An 80-minute retry of an old event gets a fresh signature time, so it looks newer than events that genuinely happened after it.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;event.id&lt;/code&gt; is the idempotency key. Use it to refuse a duplicate delivery. It does not tell you whether this event is still current.&lt;/p&gt;

&lt;h2&gt;
  
  
  The doorbell pattern
&lt;/h2&gt;

&lt;p&gt;The Indie Hackers replies kept circling the same fix: stop treating the payload as a command.&lt;/p&gt;

&lt;p&gt;Store the event. Return 200. Then ask RevenueCat what is true now.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;processedEvents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;insert&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="k"&gt;catch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;duplicate&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sendStatus&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;entitlements&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;revenuecat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;customers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;activeEntitlements&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;users&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;app_user_id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;entitled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;entitlements&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;premium&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A late retry of an old event then triggers the same current-state read. It cannot roll the wallet backward just because it arrived last.&lt;/p&gt;

&lt;p&gt;Wallet credits still need a ledger. If one event means one additive grant, insert &lt;code&gt;event.id&lt;/code&gt; under a unique constraint in the same transaction as the credit row. Idempotency answers “did I already process this id?” Ordering answers “is this still the latest truth?” Those are different questions.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test happy-path fixtures hide
&lt;/h2&gt;

&lt;p&gt;Most RevenueCat webhook tests send one event, assert 200, and stop. A &lt;a href="https://fetchsandbox.com/webhook-sandbox" rel="noopener noreferrer"&gt;webhook sandbox&lt;/a&gt; is only useful here if it can replay the same signed events in more than one order.&lt;/p&gt;

&lt;p&gt;The useful fixture sends the same event set in more than one order:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;RENEWAL → EXPIRATION
EXPIRATION → RENEWAL
RENEWAL, then the same RENEWAL again
old EXPIRATION retried after a newer RENEWAL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For every permutation, read the application state, not the HTTP status:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;final entitlement == RevenueCat active entitlements
each event.id creates at most one side effect
a stale delivery cannot overwrite newer state
the handler did not “fix” order by ignoring every event
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line is the trap. A patch that writes nothing is perfectly order-independent and completely wrong.&lt;/p&gt;

&lt;p&gt;Another comment in that thread asked for the same honesty in the receipt: if order dependence between two events was never tested, say so instead of reporting green.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prove it from the coding agent, then keep it in CI
&lt;/h2&gt;

&lt;p&gt;Connect &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt; in Cursor or Claude Code. The RevenueCat workflows do not stop at “we received the webhook.” &lt;code&gt;webhook_event_verified&lt;/code&gt; and &lt;code&gt;entitlement_verified_after_purchase&lt;/code&gt; both read &lt;code&gt;/active_entitlements&lt;/code&gt; after the write.&lt;/p&gt;

&lt;p&gt;Ask the agent:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Audit my RevenueCat webhook for out-of-order wallet updates.
Keep the still-broken tree. Propose a diff. Call prove_fix before
writing it to disk. The probe must fail on the old code when an
older EXPIRATION arrives after a newer RENEWAL, and hold on the
patched code. Paste the receipt.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Green is only allowed on a measured flip:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;buggy tree  → terminal state diverged  → exit 1
fixed tree  → all tested orders agree  → exit 0
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the probe cannot reproduce the order bug, the result stays unproven. A model saying “this looks idempotent now” is not the gate.&lt;/p&gt;

&lt;p&gt;Then keep the same fixtures in CI so a later change does not bring the last-write-wins handler back:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run API integration workflows&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;npx fetchsandbox run "$FETCHSANDBOX_ID" --all --json \&lt;/span&gt;
      &lt;span class="s"&gt;&amp;gt; fetchsandbox-workflows.json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Paste the &lt;code&gt;prove_fix&lt;/code&gt; receipt into the pull request by hand. Automatic GitHub comments are not a shipped FetchSandbox feature.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## RevenueCat order proof&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; event set: RENEWAL, EXPIRATION
&lt;span class="p"&gt;-&lt;/span&gt; duplicate event.id: exercised
&lt;span class="p"&gt;-&lt;/span&gt; before: wallet followed arrival order
&lt;span class="p"&gt;-&lt;/span&gt; after: wallet matched active entitlements
&lt;span class="p"&gt;-&lt;/span&gt; not verified: two concurrent deliveries for the same customer
&lt;span class="p"&gt;-&lt;/span&gt; receipt: https://fetchsandbox.com/runs/...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The full timestamp map, retry schedule, and checklist live in the &lt;a href="https://fetchsandbox.com/blog/revenuecat-webhooks-out-of-order" rel="noopener noreferrer"&gt;canonical guide&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>revenuecat</category>
      <category>webhooks</category>
      <category>testing</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Your Paddle Webhook Is Signed. The Seat Count Is Not.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Sat, 12 Sep 2026 12:49:17 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/your-paddle-webhook-is-signed-the-seat-count-is-not-2m67</link>
      <guid>https://dev.to/fetchsandbox/your-paddle-webhook-is-signed-the-seat-count-is-not-2m67</guid>
      <description>&lt;p&gt;The webhook passed signature verification. Paddle said the transaction was complete. The customer paid for 5 seats.&lt;/p&gt;

&lt;p&gt;The database granted 500.&lt;/p&gt;

&lt;p&gt;Nothing forged the webhook. The dangerous value had entered earlier, through &lt;code&gt;customData&lt;/code&gt; in browser checkout code. Paddle stored it, included it in &lt;code&gt;transaction.completed&lt;/code&gt;, and correctly signed the entire payload.&lt;/p&gt;

&lt;p&gt;The handler confused &lt;strong&gt;authentic delivery&lt;/strong&gt; with &lt;strong&gt;authoritative data&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The two quantities inside one valid event
&lt;/h2&gt;

&lt;p&gt;Paddle Checkout can receive custom metadata from frontend code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="nx"&gt;Paddle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Checkout&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;items&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;priceId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pri_seat&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;customData&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;workspaceId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ws_123&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;seats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When the transaction completes, the webhook may carry both values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"event_type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"transaction.completed"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"custom_data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"workspaceId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"ws_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"seats"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"items"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"price"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"pri_seat"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"quantity"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Both are inside the body protected by &lt;code&gt;Paddle-Signature&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;But the fields answer different questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;custom_data.seats&lt;/code&gt; says what checkout code attached&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;items[0].quantity&lt;/code&gt; says what the completed transaction billed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The signature proves Paddle sent those bytes. It does not turn browser-originated metadata into a billing fact.&lt;/p&gt;

&lt;h2&gt;
  
  
  The handler that looked secure
&lt;/h2&gt;

&lt;p&gt;The vulnerable code was not missing signature verification:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;paddle&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;webhooks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;unmarshal&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;rawBody&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PADDLE_WEBHOOK_SECRET&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;eventType&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;transaction.completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;customData&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;seats&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;grantSeats&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;customData&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;workspaceId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nx"&gt;seats&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It checked the signature, waited for the completed event, and wrote the entitlement on the server.&lt;/p&gt;

&lt;p&gt;Then it trusted the wrong field.&lt;/p&gt;

&lt;p&gt;The same problem exists with &lt;code&gt;workspaceId&lt;/code&gt;. If browser code supplies the tenant identifier and the webhook handler treats it as authorization, payment for one checkout can be routed to an account the server never approved.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the fixed trust boundary looks like
&lt;/h2&gt;

&lt;p&gt;Use custom data for correlation, not authority.&lt;/p&gt;

&lt;p&gt;The safer flow is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a purchase intent on your server.&lt;/li&gt;
&lt;li&gt;Store the intended workspace against an opaque checkout reference.&lt;/li&gt;
&lt;li&gt;Put only that reference in Paddle &lt;code&gt;customData&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;On &lt;code&gt;transaction.completed&lt;/code&gt;, resolve the reference on your server.&lt;/li&gt;
&lt;li&gt;Find the expected, server-known &lt;code&gt;price.id&lt;/code&gt; inside &lt;code&gt;data.items[]&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Grant the quantity on that paid line item.&lt;/li&gt;
&lt;li&gt;Deduplicate on Paddle's &lt;code&gt;event_id&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The quantity extraction should be explicit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seatItem&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;items&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;find&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt;
    &lt;span class="nx"&gt;item&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;price&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;PADDLE_SEAT_PRICE_ID&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seats&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;seatItem&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;quantity&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isSafeInteger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;seats&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
  &lt;span class="nx"&gt;seats&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt;
  &lt;span class="nx"&gt;seats&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="nx"&gt;_000&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Invalid billed seat quantity&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That upper bound is an application decision. The important part is that the entitlement comes from the completed, paid line item rather than an arbitrary metadata key.&lt;/p&gt;

&lt;p&gt;Then resolve ownership separately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;checkoutRef&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;customData&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;checkoutRef&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;purchase&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;db&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;purchaseIntent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;findUnique&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;where&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;checkoutRef&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;purchase&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;throw&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Unknown checkout reference&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;applyEntitlement&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;workspaceId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;purchase&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;workspaceId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;eventId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;eventId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;transactionId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;seats&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Payment facts, correlation metadata, and tenant authorization are now three separate concerns.&lt;/p&gt;

&lt;h2&gt;
  
  
  The test that exposes the bug
&lt;/h2&gt;

&lt;p&gt;A happy-path webhook fixture usually gives both fields the same value:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;custom_data.seats      = 5
items[0].quantity      = 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The vulnerable and fixed handlers both pass.&lt;/p&gt;

&lt;p&gt;The useful test makes the trust boundary disagree:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Paddle-Signature       = valid
custom_data.seats      = 500
items[0].quantity      = 5
expected grant         = 5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not stop at &lt;code&gt;response.status === 200&lt;/code&gt;. Read the entitlement your handler wrote.&lt;/p&gt;

&lt;p&gt;The complete invariant is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;5 units billed
500 claimed in client metadata
5 seats granted
same event delivered again
still 5 seats granted
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That last line matters because Paddle uses at-least-once webhook delivery. Fixing the quantity source while leaving the grant non-idempotent still doubles access on a retry.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run it from the coding agent
&lt;/h2&gt;

&lt;p&gt;FetchSandbox has a deterministic &lt;code&gt;paddle_seat_quantity&lt;/code&gt; probe for this exact mismatch. It sends a freshly signed &lt;code&gt;transaction.completed&lt;/code&gt; event to the application's real handler and observes how many seats the application grants.&lt;/p&gt;

&lt;p&gt;With &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt; connected to Cursor or Claude Code:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;./fetchsandbox audit my Paddle webhook for client-controlled seat
quantity. Propose the fix, then call prove_fix before applying it.
Show whether the buggy tree grants 500 while the fixed tree grants
the 5 units actually billed. Return the proof receipt.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same probe runs against both trees:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;buggy tree  → granted=500, billed=5 → VIOLATED
fixed tree  → granted=5,   billed=5 → HELD
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the probe cannot load the handler or measure the grant, the result is inconclusive. It does not turn “could not run” into green.&lt;/p&gt;

&lt;p&gt;Here is a &lt;a href="https://fetchsandbox.com/runs/fix-417c74290218?flow=9c61192f9179" rel="noopener noreferrer"&gt;measured Paddle seat-quantity receipt&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Carry the proof into review and CI
&lt;/h2&gt;

&lt;p&gt;Attach the receipt to the pull request with the exact invariant:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Paddle entitlement proof&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; signature: valid
&lt;span class="p"&gt;-&lt;/span&gt; paid quantity: 5
&lt;span class="p"&gt;-&lt;/span&gt; client metadata claim: 500
&lt;span class="p"&gt;-&lt;/span&gt; before patch: granted 500
&lt;span class="p"&gt;-&lt;/span&gt; after patch: granted 5
&lt;span class="p"&gt;-&lt;/span&gt; receipt: https://fetchsandbox.com/runs/fix-...?flow=...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That receipt proves the proposed diff against the mismatched-data failure. Keep the broader provider workflows running in CI as a separate regression gate:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run API integration workflows&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;npx fetchsandbox run "$FETCHSANDBOX_ID" --all --json \&lt;/span&gt;
      &lt;span class="s"&gt;&amp;gt; fetchsandbox-workflows.json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The PR receipt and CI check have different jobs. The receipt shows the bug existed and the proposed fix changed the measured behavior. CI keeps the provider workflows from drifting as more code lands.&lt;/p&gt;

&lt;h2&gt;
  
  
  Signed does not mean trusted
&lt;/h2&gt;

&lt;p&gt;Verify every Paddle webhook signature. That is non-negotiable.&lt;/p&gt;

&lt;p&gt;Then make a second decision for every field: &lt;strong&gt;where did this value originate, and is it authoritative for the action I am about to take?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For seat provisioning:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;custom_data&lt;/code&gt; can correlate the transaction&lt;/li&gt;
&lt;li&gt;a server-side record decides the workspace&lt;/li&gt;
&lt;li&gt;paid &lt;code&gt;items[]&lt;/code&gt; decide the quantity&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;event_id&lt;/code&gt; makes processing idempotent&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A valid signature can carry an unsafe instruction. The verifier has to test both the cryptography and the business trust boundary.&lt;/p&gt;

&lt;p&gt;The full implementation guide is on &lt;a href="https://fetchsandbox.com/blog/paddle-webhook-custom-data-seat-quantity" rel="noopener noreferrer"&gt;FetchSandbox&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>paddle</category>
      <category>webhooks</category>
      <category>testing</category>
      <category>mcp</category>
    </item>
    <item>
      <title>AI Code Is Cheap. Your Pull Request Still Needs Proof.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Tue, 08 Sep 2026 20:59:59 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/ai-code-is-cheap-your-pull-request-still-needs-proof-5hb2</link>
      <guid>https://dev.to/fetchsandbox/ai-code-is-cheap-your-pull-request-still-needs-proof-5hb2</guid>
      <description>&lt;p&gt;Your coding agent opens a plausible pull request before lunch. The tests it wrote pass, the diff reads cleanly, and the reviewer is left with the expensive question: &lt;strong&gt;did this change survive the provider failure it claims to handle?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Writing the integration is now the cheap step. Establishing that the integration behaves correctly under retries, lost responses, stale events, webhooks, and partial failure is still the hard step.&lt;/p&gt;

&lt;h2&gt;
  
  
  How should you verify AI-generated code before merging it?
&lt;/h2&gt;

&lt;p&gt;Verify AI-generated integration code with an independent runtime check, not only the agent's explanation or the tests it wrote. Reproduce the provider failure against the still-broken code, apply the proposed diff to a copy, run the same failure again, and require a measured broken-to-fixed result before merge.&lt;/p&gt;

&lt;p&gt;The output should be a shareable receipt that records:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the failure scenario that ran&lt;/li&gt;
&lt;li&gt;whether the bug reproduced before the patch&lt;/li&gt;
&lt;li&gt;whether the invariant held after the patch&lt;/li&gt;
&lt;li&gt;the requests, responses, and observed side effects&lt;/li&gt;
&lt;li&gt;an explicit proven, rejected, or unproven verdict&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That receipt changes the pull-request conversation from “the agent says this works” to “here is what ran, and here is what changed.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is reviewing AI-written code becoming the bottleneck?
&lt;/h2&gt;

&lt;p&gt;AI coding agents can produce more diffs than a team can carefully inspect. Faster generation increases review pressure because plausible code still has to be checked against the real behavior of every external system it touches.&lt;/p&gt;

&lt;p&gt;Static review remains useful. It can catch an unsafe query, a missing authorization check, or an obviously incorrect SDK call. But many integration bugs do not live on one suspicious line:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the same webhook is delivered twice&lt;/li&gt;
&lt;li&gt;a successful response disappears after the provider commits the write&lt;/li&gt;
&lt;li&gt;a stale event arrives after a newer event&lt;/li&gt;
&lt;li&gt;the first request creates an ID that the second request never persists&lt;/li&gt;
&lt;li&gt;a handler returns &lt;code&gt;200&lt;/code&gt; but leaves local state unchanged&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The diff can look reasonable while the lifecycle is still broken.&lt;/p&gt;

&lt;p&gt;Most current writing about AI code verification says the same broad thing: define acceptance criteria, run tests, and keep humans in the loop. That advice is correct, but it usually stops before the reviewer gets an inspectable artifact showing the failure on the actual code.&lt;/p&gt;

&lt;p&gt;The missing layer is &lt;strong&gt;runtime evidence attached to the change&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is a green test written by the same agent not enough?
&lt;/h2&gt;

&lt;p&gt;A test written by the same agent can repeat the agent's mistaken assumption. It may mock away the provider behavior, assert only an HTTP status, or accept a patch that removes the desired side effect along with the bug.&lt;/p&gt;

&lt;p&gt;Consider a webhook that grants seats twice when Paddle redelivers one &lt;code&gt;transaction.completed&lt;/code&gt; event. A weak verifier asks:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Did the seat count stop increasing?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;An agent can satisfy that check by granting no seats at all. The duplicate is gone, the customer receives nothing, and CI is green.&lt;/p&gt;

&lt;p&gt;The real invariant is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;One purchase grants the purchased seats exactly once,
however many times the same event is delivered.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That distinction matters because CI is an optimization target. If “make the reported symptom disappear” is enough to pass, a plausible but destructive patch can pass.&lt;/p&gt;

&lt;p&gt;An independent verifier must be able to reject the first fix.&lt;/p&gt;

&lt;p&gt;That gate behavior showed up in a separate measured FetchSandbox acceptance run against a Paddle billing demo. The bug in that run was a webhook trusting a client-controlled seat quantity. The first proposed patch came back &lt;code&gt;fix_incomplete&lt;/code&gt;; the agent iterated, and only the second patch produced:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;buggy tree   → invariant VIOLATED
fixed tree   → invariant HELD
verdict      → green_allowed: true
proof grade  → measured
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;a href="https://fetchsandbox.com/runs/fix-417c74290218?flow=9c61192f9179" rel="noopener noreferrer"&gt;Inspect the measured client-controlled-seat receipt&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does FetchSandbox verify on an AI-generated fix?
&lt;/h2&gt;

&lt;p&gt;FetchSandbox acts as the integration verification layer between the coding agent and the pull request. Its &lt;code&gt;prove_fix&lt;/code&gt; MCP tool runs the customer's real code before and after the proposed patch against the same provider failure scenario.&lt;/p&gt;

&lt;p&gt;The sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The agent identifies an integration bug with &lt;code&gt;find_bugs&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;It proposes a diff with &lt;code&gt;fix_bug&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Before writing that diff into the working tree, it calls &lt;code&gt;prove_fix&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;FetchSandbox materializes the still-broken project.&lt;/li&gt;
&lt;li&gt;It applies the diff to a separate copy.&lt;/li&gt;
&lt;li&gt;It runs the same provider scenario against both versions.&lt;/li&gt;
&lt;li&gt;It returns green only when the bug reproduces before and the required invariant holds after.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If FetchSandbox cannot reproduce the bug, it does not convert uncertainty into a pass. The result is unproven, and &lt;code&gt;green_allowed&lt;/code&gt; remains false.&lt;/p&gt;

&lt;p&gt;A successful result has a machine-readable shape:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"green_allowed"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"state"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"proven"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reproduced"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"verified"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"receipt_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://fetchsandbox.com/runs/fix-...?flow=..."&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The receipt URL is public to anyone who has the link. A reviewer can inspect it without access to the original agent conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should go into the pull request?
&lt;/h2&gt;

&lt;p&gt;An AI-authored pull request should include the claim, the invariant, and the evidence. A short verification block is enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## Integration verification&lt;/span&gt;
&lt;span class="p"&gt;
-&lt;/span&gt; Failure: duplicate Paddle transaction webhook
&lt;span class="p"&gt;-&lt;/span&gt; Invariant: one purchase grants seats exactly once
&lt;span class="p"&gt;-&lt;/span&gt; Before patch: reproduced
&lt;span class="p"&gt;-&lt;/span&gt; After patch: held
&lt;span class="p"&gt;-&lt;/span&gt; Receipt: https://fetchsandbox.com/runs/fix-...?flow=...
&lt;span class="p"&gt;-&lt;/span&gt; Remaining review: authorization, data migration, rollout safety
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The receipt does not replace code review. It narrows code review to what still needs human judgment.&lt;/p&gt;

&lt;p&gt;The reviewer can now check architecture, maintainability, security boundaries, and whether the invariant represents the intended business rule. They do not have to infer whether the agent ever ran the failure it claims to have fixed.&lt;/p&gt;

&lt;h2&gt;
  
  
  How does the proof continue into CI/CD?
&lt;/h2&gt;

&lt;p&gt;The pre-merge receipt and the CI workflow gate answer different questions. The receipt proves that one proposed diff changed a reproduced failure into the required behavior. CI keeps the provider workflow from regressing as the branch changes.&lt;/p&gt;

&lt;p&gt;FetchSandbox's CLI produces structured JSON and exits non-zero when a normal workflow run fails:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;Run API integration workflows&lt;/span&gt;
  &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;|&lt;/span&gt;
    &lt;span class="s"&gt;npx fetchsandbox run "$FETCHSANDBOX_ID" --all --json \&lt;/span&gt;
      &lt;span class="s"&gt;&amp;gt; fetchsandbox-workflows.json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Store &lt;code&gt;fetchsandbox-workflows.json&lt;/code&gt; as a build artifact and make the job a required status check when your team is comfortable with its coverage. The same workflow definitions can be explored locally through &lt;a href="https://dev.to/mcp"&gt;FetchSandbox MCP&lt;/a&gt; and run headlessly in the pipeline.&lt;/p&gt;

&lt;p&gt;Today, the &lt;code&gt;prove_fix&lt;/code&gt; receipt is attached to the pull request as a link. An automatic GitHub Action that posts the receipt and check result is not yet a shipped FetchSandbox feature, so do not describe that automation as live. The available path is explicit and auditable: prove through MCP, paste the receipt into the PR, and keep the provider workflows running in CI.&lt;/p&gt;

&lt;h2&gt;
  
  
  What does a proof receipt prove?
&lt;/h2&gt;

&lt;p&gt;A FetchSandbox proof receipt proves the behavior shown on the artifact: the named scenario ran, the bug reproduced on the pre-patch code, and the measured invariant held or did not hold after the patch.&lt;/p&gt;

&lt;p&gt;It does not prove that every requirement in the pull request is correct. It does not prove an untested security property, a safe database migration, or a production rollout plan. Good evidence has boundaries.&lt;/p&gt;

&lt;p&gt;That is also why an unproven verdict is useful. “We could not reproduce this failure” tells the reviewer where the evidence ends. A confident green with no execution would be worse.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changes when every pull request carries evidence?
&lt;/h2&gt;

&lt;p&gt;The unit of review changes from a diff plus a confident summary to a diff plus a falsifiable result.&lt;/p&gt;

&lt;p&gt;The coding agent can still generate quickly. The human still owns intent and risk. The verification layer owns one narrower job: make the provider failure happen, run the proposed fix against it, and preserve what happened in a receipt.&lt;/p&gt;

&lt;p&gt;Cheap code is useful only when teams can qualify it for production without creating an equally large review queue. The next improvement in agentic software development is not another faster generator. It is a better way to say &lt;strong&gt;no&lt;/strong&gt; to an unproven patch.&lt;/p&gt;

&lt;p&gt;Connect &lt;a href="https://dev.to/mcp"&gt;FetchSandbox MCP&lt;/a&gt; to Cursor or Claude Code, run &lt;code&gt;find_bugs → fix_bug → prove_fix&lt;/code&gt;, and attach the returned receipt to the next integration pull request. For broader workflow setup, see &lt;a href="https://dev.to/agent-api-workflow-testing"&gt;agent API workflow testing&lt;/a&gt; and &lt;a href="https://dev.to/api-integration-testing-in-ci"&gt;API integration testing in CI&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions developers ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does passing CI prove AI-generated code is correct?
&lt;/h3&gt;

&lt;p&gt;No. Passing CI proves only that the checks configured in CI passed. Verify that those checks exercise the relevant provider failure and assert the intended business state, not just an HTTP status or the absence of an exception.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should the coding agent write its own tests?
&lt;/h3&gt;

&lt;p&gt;The agent can write tests, but those tests should not be the only judge of its patch. Use independent scenarios, reviewed invariants, negative controls, or a separate verification system that can reject the agent's first answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does FetchSandbox replace human pull-request review?
&lt;/h3&gt;

&lt;p&gt;No. FetchSandbox supplies runtime evidence for API integration behavior. Humans still review intent, architecture, security, migrations, and rollout risk.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can FetchSandbox automatically comment on every GitHub pull request?
&lt;/h3&gt;

&lt;p&gt;Not currently. The shipped workflow returns a shareable proof receipt through MCP, which you can attach manually. The CLI can run workflows as a CI status check; automatic PR comments are planned rather than live.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when FetchSandbox cannot reproduce the bug?
&lt;/h3&gt;

&lt;p&gt;The proof declines to green. An unavailable or inconclusive scenario remains unproven instead of being reported as verified.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>devops</category>
      <category>ci</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Test ElevenLabs 429 Handling Without Burning API Quota</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Mon, 07 Sep 2026 14:21:39 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/test-elevenlabs-429-handling-without-burning-api-quota-o6m</link>
      <guid>https://dev.to/fetchsandbox/test-elevenlabs-429-handling-without-burning-api-quota-o6m</guid>
      <description>&lt;p&gt;Your app loads the ElevenLabs voice list successfully. The user picks a voice, clicks Generate, and the text-to-speech request returns &lt;code&gt;429&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;That is a more useful test than an API that always rate-limits. Authentication worked. Discovery worked. The failure happened on the expensive operation, after the UI had already promised progress.&lt;/p&gt;

&lt;h2&gt;
  
  
  The branch most mocks skip
&lt;/h2&gt;

&lt;p&gt;A generated ElevenLabs integration usually proves this sequence:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET  /v1/voices                    -&amp;gt; 200
POST /v1/text-to-speech/{voiceId}  -&amp;gt; 200
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The fixture is green forever. It does not answer:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What happens when voice discovery succeeds but generation is throttled?&lt;/li&gt;
&lt;li&gt;Does the client retry immediately?&lt;/li&gt;
&lt;li&gt;Does the SDK already retry before your wrapper retries?&lt;/li&gt;
&lt;li&gt;Does the UI create several concurrent generation jobs?&lt;/li&gt;
&lt;li&gt;Is there a maximum attempt count?&lt;/li&gt;
&lt;li&gt;What happens when no &lt;code&gt;Retry-After&lt;/code&gt; header is present?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are client-policy questions. Burning production quota until ElevenLabs returns 429 is an expensive way to discover the answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  A 429 is not permission to retry forever
&lt;/h2&gt;

&lt;p&gt;Rate-limit handling needs a decision function, not a recursive catch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;RetryDecision&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;wait&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;delayMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt;
  &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stop&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nl"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;decide429&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;Response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="nx"&gt;remainingMs&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;
&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nx"&gt;RetryDecision&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stop&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;retry budget exhausted&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;retryAfter&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;retry-after&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;fallbackMs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="nx"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;delayMs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;retryAfter&lt;/span&gt;
    &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="nf"&gt;parseRetryAfter&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;retryAfter&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;fallbackMs&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;delayMs&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;remainingMs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;stop&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;reason&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;deadline exceeded&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;action&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;wait&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;delayMs&lt;/span&gt; &lt;span class="p"&gt;};&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important assertions are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;generation_attempts &amp;lt;= 3
voice_list_calls == 1
no duplicate UI jobs
fallback backoff used when Retry-After is absent
controlled error returned when the retry budget ends
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not assert only that the client eventually threw. A tight loop can throw after hammering the provider twenty times.&lt;/p&gt;

&lt;h2&gt;
  
  
  Watch for stacked retries
&lt;/h2&gt;

&lt;p&gt;Many SDKs retry transient failures internally. If your application wrapper also retries, the budgets multiply:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SDK attempts:          3
application attempts:  3
actual HTTP requests:  9
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is why the test needs a request count. Pick one layer to own the policy, or coordinate the limits explicitly.&lt;/p&gt;

&lt;p&gt;The failure can get worse in a browser. A disabled button that re-enables too early allows the user to start another retry tree while the first is waiting.&lt;/p&gt;

&lt;h2&gt;
  
  
  Force the real workflow shape without an API key
&lt;/h2&gt;

&lt;p&gt;FetchSandbox ships an ElevenLabs twin with a curated &lt;code&gt;text_to_speech&lt;/code&gt; workflow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;GET /v1/voices&lt;/code&gt; to discover available voices&lt;/li&gt;
&lt;li&gt;Select the returned &lt;code&gt;voice_id&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;&lt;code&gt;POST /v1/text-to-speech/{voice_id}&lt;/code&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Under the &lt;code&gt;rate_limited&lt;/code&gt; scenario, only the &lt;code&gt;textToSpeech&lt;/code&gt; operation changes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;rate_limited&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;overrides&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;operation_id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;textToSpeech&lt;/span&gt;
      &lt;span class="na"&gt;response_status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;429&lt;/span&gt;
      &lt;span class="na"&gt;error_code&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;rate_limit_exceeded&lt;/span&gt;
      &lt;span class="na"&gt;error_detail&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Rate&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;limit&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;exceeded.&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;Please&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;try&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;again&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;later."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first step still succeeds. The generation step returns the provider-shaped 429. The configured response intentionally has no &lt;code&gt;Retry-After&lt;/code&gt;, so this scenario tests your documented fallback backoff and retry ceiling rather than server-directed waiting.&lt;/p&gt;

&lt;p&gt;No ElevenLabs API key and no quota consumption are involved.&lt;/p&gt;

&lt;p&gt;With &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt; in Cursor or Claude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run elevenlabs text_to_speech under rate_limited.
Verify voice discovery succeeds and generation returns 429.
Count every generation attempt, prove the voice list is not fetched
again, and confirm the client stops at its retry budget.
Give me the run receipt.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The receipt provides the request sequence and exact failure. Your app-level test supplies the clock and UI assertions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make backoff deterministic
&lt;/h2&gt;

&lt;p&gt;Do not make CI literally sleep through production delays. Inject time:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;RetryClock&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="na"&gt;ms&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="nb"&gt;Promise&lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="k"&gt;void&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nf"&gt;random&lt;/span&gt;&lt;span class="p"&gt;():&lt;/span&gt; &lt;span class="kr"&gt;number&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Production uses the real clock. Tests use a fake clock and record requested delays:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;attempt 1 -&amp;gt; 429 -&amp;gt; requested sleep 1000ms
attempt 2 -&amp;gt; 429 -&amp;gt; requested sleep 2000ms
attempt 3 -&amp;gt; 429 -&amp;gt; stop
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You can prove exponential backoff, jitter bounds, cancellation, and the total deadline without adding seconds to every CI run.&lt;/p&gt;

&lt;h2&gt;
  
  
  Local first, pipeline next
&lt;/h2&gt;

&lt;p&gt;Ask the coding agent to run the twin while it writes the policy. Then keep the same scenario in CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;fetchsandbox run &amp;lt;sandbox-id&amp;gt; text_to_speech &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scenario&lt;/span&gt; rate_limited &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The local receipt is useful during review. The pipeline prevents a future refactor from changing &lt;code&gt;maxAttempts = 3&lt;/code&gt; into unbounded recursion.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions people ask
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How can I test ElevenLabs rate limits without using quota?
&lt;/h3&gt;

&lt;p&gt;Point the integration at a service twin and force &lt;code&gt;textToSpeech&lt;/code&gt; to return 429. This exercises the client contract without sending audio-generation requests to ElevenLabs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should I always use &lt;code&gt;Retry-After&lt;/code&gt;?
&lt;/h3&gt;

&lt;p&gt;Honor a valid &lt;code&gt;Retry-After&lt;/code&gt; when the provider sends one, subject to your own deadline and attempt cap. When it is absent, use a documented capped backoff with jitter. Test both branches separately.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should a 429 restart the whole workflow?
&lt;/h3&gt;

&lt;p&gt;No. Resume from the safe failed operation. In this workflow, do not fetch voices again or create a second local generation job merely because text-to-speech was throttled.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is rate limiting the same as exhausted quota?
&lt;/h3&gt;

&lt;p&gt;No. A temporary rate limit is generally pause-and-retry. A hard quota or billing limit is stop-and-alert. OpenAI, for example, can use HTTP 429 for both while distinguishing them with error codes such as &lt;code&gt;rate_limit_exceeded&lt;/code&gt; and &lt;code&gt;insufficient_quota&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does FetchSandbox replace ElevenLabs testing?
&lt;/h3&gt;

&lt;p&gt;No. Use the twin for repeatable negative paths and CI. Use ElevenLabs afterward for final credentials, model availability, audio quality, and real account limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  The evidence to attach
&lt;/h2&gt;

&lt;p&gt;Require a receipt showing one successful voice lookup followed by the forced 429, plus an app assertion showing the exact retry count and requested delays.&lt;/p&gt;

&lt;p&gt;The canonical guide covers &lt;a href="https://fetchsandbox.com/blog/test-mid-workflow-429-without-burning-quota" rel="noopener noreferrer"&gt;mid-workflow 429 testing without partner quota&lt;/a&gt;. Connect the twins through &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>elevenlabs</category>
      <category>ratelimit</category>
      <category>testing</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Your PayPal Refund Worked. Your Lovable App Says It Failed.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Mon, 07 Sep 2026 13:26:57 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/your-paypal-refund-worked-your-lovable-app-says-it-failed-5gh2</link>
      <guid>https://dev.to/fetchsandbox/your-paypal-refund-worked-your-lovable-app-says-it-failed-5gh2</guid>
      <description>&lt;p&gt;Your app called PayPal's refund endpoint. PayPal completed the operation, but the &lt;code&gt;201 Created&lt;/code&gt; response never reached your server.&lt;/p&gt;

&lt;p&gt;The app retries and receives &lt;code&gt;422 CAPTURE_FULLY_REFUNDED&lt;/code&gt;. If the handler treats every non-2xx response as a failed refund, your customer and support team now see opposite versions of reality.&lt;/p&gt;

&lt;p&gt;This guide shows the production-safe behavior and how to test it before using a real payment.&lt;/p&gt;

&lt;h2&gt;
  
  
  What should a PayPal refund handler do after a timeout?
&lt;/h2&gt;

&lt;p&gt;Reuse the same &lt;code&gt;PayPal-Request-Id&lt;/code&gt;, then reconcile PayPal's state before marking the refund failed.&lt;/p&gt;

&lt;p&gt;PayPal uses &lt;code&gt;PayPal-Request-Id&lt;/code&gt; as the idempotency key for REST &lt;code&gt;POST&lt;/code&gt; operations. One refund operation should have one stored request ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;type&lt;/span&gt; &lt;span class="nx"&gt;RefundRecord&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;captureId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;paypalRequestId&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nl"&gt;status&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;pending&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;completed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;reconcile_required&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;failed&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;};&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send that ID on the first attempt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;fetch&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;PAYPAL_API&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/v2/payments/captures/&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;captureId&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;/refund`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;method&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;POST&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;Authorization&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;accessToken&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;PayPal-Request-Id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;paypalRequestId&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="na"&gt;body&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
      &lt;span class="na"&gt;amount&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="na"&gt;currency_code&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;USD&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="na"&gt;value&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;25.00&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;}),&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the request times out, send the same request ID again. Do not create a new UUID per attempt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why does PayPal return &lt;code&gt;CAPTURE_FULLY_REFUNDED&lt;/code&gt;?
&lt;/h2&gt;

&lt;p&gt;It means the capture's refundable balance is already zero. Possible causes include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;the first request worked but its response was lost&lt;/li&gt;
&lt;li&gt;another support tool issued the refund&lt;/li&gt;
&lt;li&gt;the PayPal dashboard issued it&lt;/li&gt;
&lt;li&gt;several partial refunds already reached the captured total&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The response shape is carried inside a broader &lt;code&gt;422&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"UNPROCESSABLE_ENTITY"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"details"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"issue"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"CAPTURE_FULLY_REFUNDED"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"description"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"The capture has already been fully refunded."&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not branch on status code alone. &lt;code&gt;422&lt;/code&gt; can also mean the amount exceeds the remaining balance, a chargeback blocks the refund, or the account cannot perform the operation.&lt;/p&gt;

&lt;p&gt;Parse &lt;code&gt;details[].issue&lt;/code&gt;, then choose the business response.&lt;/p&gt;

&lt;h2&gt;
  
  
  The correct state transition
&lt;/h2&gt;

&lt;p&gt;When the issue is &lt;code&gt;CAPTURE_FULLY_REFUNDED&lt;/code&gt;, stop automatic retries and move the local operation to &lt;code&gt;reconcile_required&lt;/code&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;422&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;issues&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;details&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nf"&gt;map&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nl"&gt;issue&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kr"&gt;string&lt;/span&gt; &lt;span class="p"&gt;})&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;detail&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;issue&lt;/span&gt;
  &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;issues&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;includes&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;CAPTURE_FULLY_REFUNDED&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;markReconcileRequired&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;reconcilePayPalRefund&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;refund&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reconciliation should confirm the external refund ID, amount, currency, and status before completing the local row. Avoid changing &lt;code&gt;CAPTURE_FULLY_REFUNDED&lt;/code&gt; directly into success because PayPal may have been updated by a different system.&lt;/p&gt;

&lt;p&gt;The invariant is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;PayPal has refunded the intended amount
AND the local operation points to that external state
AND another retry cannot issue or record the refund again
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  How do I test this in a Lovable or Bolt app?
&lt;/h2&gt;

&lt;p&gt;Use a stateful PayPal service twin to force the failure before deployment.&lt;/p&gt;

&lt;p&gt;FetchSandbox includes a curated PayPal Payments workflow called &lt;code&gt;capture_and_refund&lt;/code&gt;. Its normal path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Reads capture &lt;code&gt;2GG279541U471931P&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Requests a partial refund of &lt;code&gt;$25.00 USD&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Expects &lt;code&gt;201 Created&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Emits &lt;code&gt;PAYMENT.CAPTURE.REFUNDED&lt;/code&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;code&gt;refund_not_allowed&lt;/code&gt; scenario changes the second step to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;response_status&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="m"&gt;422&lt;/span&gt;
&lt;span class="na"&gt;error_code&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;CAPTURE_FULLY_REFUNDED&lt;/span&gt;
&lt;span class="na"&gt;error_detail&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;The&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;capture&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;has&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;already&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;been&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;fully&lt;/span&gt;&lt;span class="nv"&gt; &lt;/span&gt;&lt;span class="s"&gt;refunded."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;FetchSandbox MCP&lt;/a&gt; to Cursor, Claude, or another MCP-capable agent and ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run paypal-payments capture_and_refund with refund_not_allowed.
Inspect my handler for the exact 422 issue. Prove it stops retries
and reconciles before reporting a failure. Return the receipt URL.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The run receipt records the request sequence and exact error contract. Your app-level assertion checks the database and UI.&lt;/p&gt;

&lt;p&gt;For CI:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;fetchsandbox run &amp;lt;sandbox-id&amp;gt; capture_and_refund &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scenario&lt;/span&gt; refund_not_allowed &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This gives the generated integration a failure-path acceptance test instead of relying on a code reviewer to infer behavior from the happy path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Can Bolt preview receive PayPal webhooks?
&lt;/h2&gt;

&lt;p&gt;Not directly. PayPal sends webhooks to a publicly reachable HTTPS endpoint, while Bolt's browser-based preview is not a public webhook destination. Deploy to Bolt Cloud, Netlify, or another host for final PayPal webhook connectivity testing.&lt;/p&gt;

&lt;p&gt;The service twin handles a different job: it lets the agent exercise provider state, API errors, and expected webhook events before that deployment. Use both layers:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;FetchSandbox for repeatable failure branches during development&lt;/li&gt;
&lt;li&gt;PayPal sandbox for final credentials and webhook connectivity&lt;/li&gt;
&lt;li&gt;Production only after both paths pass&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Does a service twin replace an API mock?
&lt;/h2&gt;

&lt;p&gt;No. A mock is useful when a component needs a shaped response. A twin is useful when correctness depends on what happened in earlier calls.&lt;/p&gt;

&lt;p&gt;Refund testing needs state:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;captured amount
- previous partial refunds
= remaining refundable amount
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is why &lt;code&gt;CAPTURE_FULLY_REFUNDED&lt;/code&gt; is a lifecycle test rather than a fixture test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pull-request checklist
&lt;/h2&gt;

&lt;p&gt;Before merging an AI-generated PayPal refund flow, attach evidence that:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;one refund operation stores one &lt;code&gt;PayPal-Request-Id&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;timeout retries reuse that ID&lt;/li&gt;
&lt;li&gt;the handler parses &lt;code&gt;details[].issue&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CAPTURE_FULLY_REFUNDED&lt;/code&gt; stops automatic retries&lt;/li&gt;
&lt;li&gt;local state reconciles before the UI reports success or failure&lt;/li&gt;
&lt;li&gt;the pipeline can reproduce the same 422 path&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The complete canonical guide is &lt;a href="https://fetchsandbox.com/blog/test-paypal-refunds-lovable-before-production" rel="noopener noreferrer"&gt;Test PayPal Refunds in Lovable Before Production&lt;/a&gt;. For the wider workflow, see &lt;a href="https://fetchsandbox.com/ai-app-builder-integration-testing" rel="noopener noreferrer"&gt;integration testing for AI-built apps&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>paypal</category>
      <category>lovable</category>
      <category>apitesting</category>
      <category>mcp</category>
    </item>
    <item>
      <title>Your AI builder shipped the Stripe code in an afternoon. Verifying it took a week.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Sun, 06 Sep 2026 04:20:21 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/your-ai-builder-shipped-the-stripe-code-in-an-afternoon-verifying-it-took-a-week-2o6n</link>
      <guid>https://dev.to/fetchsandbox/your-ai-builder-shipped-the-stripe-code-in-an-afternoon-verifying-it-took-a-week-2o6n</guid>
      <description>&lt;p&gt;A friend shipped a booking app last month. Lovable built the whole thing, Stripe Checkout included, in something like an afternoon. It looked finished. It &lt;em&gt;was&lt;/em&gt; finished, in the sense that you could pay and get a confirmation email.&lt;/p&gt;

&lt;p&gt;Then a customer got charged twice for one booking.&lt;/p&gt;

&lt;p&gt;Not a double-click. One payment, one &lt;code&gt;payment_intent.succeeded&lt;/code&gt;, delivered twice, because Stripe retries. The handler ran twice and created two bookings. Nobody caught it in review, and here is the uncomfortable part: the code was fine on the happy path, and nobody reviewing it had written a line of it.&lt;/p&gt;

&lt;p&gt;That is the actual shape of the problem right now. Build time collapsed. Verification time did not.&lt;/p&gt;

&lt;h2&gt;
  
  
  The generated code is not the problem
&lt;/h2&gt;

&lt;p&gt;This is where I expected to find sloppy AI output, and I didn't. The handler read exactly like something a competent developer would write on a first pass:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/webhook&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;stripe&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;webhooks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;constructEvent&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;sig&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;payment_intent.succeeded&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;createBooking&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;metadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;booking_id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;sendConfirmation&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;object&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;receipt_email&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;received&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Signature verified. Correct event type. Returns 200. It will pass code review and it will pass a smoke test, because the first delivery does exactly what it should.&lt;/p&gt;

&lt;p&gt;There is no idempotency key anywhere in it. The agent was never asked "what happens on the second delivery," so it never answered that question. Neither did the reviewer, because you cannot see a retry in a diff.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why the usual tools miss it
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Mocks&lt;/strong&gt; return the response you configured. A mock does not decide to retry you. That is the whole point of a mock, and it is also why it cannot catch this class of bug.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Staging&lt;/strong&gt; proves the happy path against real credentials. To see a duplicate delivery you have to wait for Stripe to genuinely retry, which means causing a timeout on purpose and waiting out a backoff. Nobody does this on a Tuesday.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;&lt;code&gt;stripe trigger&lt;/code&gt;&lt;/strong&gt; fires a fresh, well-formed event. Useful, but every trigger is a &lt;em&gt;first&lt;/em&gt; delivery. The bug lives in the second one.&lt;/p&gt;

&lt;p&gt;So the failure mode is invisible to all three, and it surfaces when real money moves.&lt;/p&gt;

&lt;h2&gt;
  
  
  Force the second delivery
&lt;/h2&gt;

&lt;p&gt;What I actually wanted was a Stripe that behaves like Stripe, including the annoying parts, on demand. That is what a service twin is: stateful, so a created PaymentIntent can be read back later, and scriptable, so you can ask for the retry instead of waiting for it.&lt;/p&gt;

&lt;p&gt;FetchSandbox runs as an MCP server, so the agent that wrote the integration can drive it. In Cursor or Claude:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run the Stripe accept_payment workflow, then re-run it under the
webhook_retries scenario. Tell me how many bookings exist afterwards
and give me the receipt URL.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;webhook_retries&lt;/code&gt; delivers each webhook once, then replays it twice with the &lt;strong&gt;same upstream event id&lt;/strong&gt; at backoff intervals, which is the behaviour that matters. A handler deduping on a per-delivery header still fails here, because that header changes on every retry while &lt;code&gt;event.id&lt;/code&gt; does not.&lt;/p&gt;

&lt;p&gt;The run came back with three bookings for one payment. Not an opinion about the code, a count.&lt;/p&gt;

&lt;p&gt;The fix is the boring one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;alreadyProcessed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;received&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;recordProcessed&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then re-run the same scenario and the count stays at one. That buggy-to-fixed flip is the thing worth having, because "the agent says it fixed it" and "the retry no longer double-books" are different claims.&lt;/p&gt;

&lt;h2&gt;
  
  
  The part that makes it reviewable
&lt;/h2&gt;

&lt;p&gt;Every run produces a receipt at a shareable URL: the requests in order, the state the twin held, the webhook events including the replays, and a verdict. The verdict is separate from the HTTP status on purpose, so a run can return 200 at every step and still come back unproven.&lt;/p&gt;

&lt;p&gt;That distinction is what makes this useful when you didn't write the code. You are no longer reviewing an unfamiliar diff and hoping. You are reading evidence that a specific failure was forced and handled. Put the URL in the PR.&lt;/p&gt;

&lt;p&gt;The same run works from the CLI when you want the pipeline to enforce it rather than a person remembering:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;fetchsandbox run &amp;lt;sandbox-id&amp;gt; accept_payment &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--scenario&lt;/span&gt; webhook_retries &lt;span class="nt"&gt;--json&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Local proof while building, the same check gating merges. The failure branch becomes a required step instead of a thing you meant to test.&lt;/p&gt;

&lt;h2&gt;
  
  
  Beyond Stripe
&lt;/h2&gt;

&lt;p&gt;Payments are where it bites hardest, but the pattern is identical anywhere an integration has a lifecycle. Paddle, Resend, Twilio, Clerk, Descope, and AgentMail have curated workflows and failure scenarios; the wider catalogue covers 50+ APIs generated from their OpenAPI specs. Same loop each time: run the flow, force the failure, keep the receipt.&lt;/p&gt;

&lt;p&gt;Setup is &lt;a href="https://fetchsandbox.com/mcp" rel="noopener noreferrer"&gt;one MCP config block&lt;/a&gt;, and there is more on the &lt;a href="https://fetchsandbox.com/ai-app-builder-integration-testing" rel="noopener noreferrer"&gt;approach for AI-built apps&lt;/a&gt; if you want the longer version.&lt;/p&gt;

&lt;h2&gt;
  
  
  The actual question
&lt;/h2&gt;

&lt;p&gt;I keep meeting people whose apps went from idea to paying customers in a weekend, and whose integration testing story is "staging looked fine."&lt;/p&gt;

&lt;p&gt;So, genuinely: what are you doing here? Is anyone verifying the retry, the dropped webhook, and the mid-flow auth failure before production, or is the honest answer that you find out from a support ticket? I am more interested in what people actually do than what we all agree we should do.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>webdev</category>
      <category>testing</category>
      <category>stripe</category>
    </item>
    <item>
      <title>undefined === undefined — the auth bypass your AI wrote into your checkout</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Thu, 13 Aug 2026 17:40:08 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/undefined-undefined-the-auth-bypass-your-ai-wrote-into-your-checkout-4j9g</link>
      <guid>https://dev.to/fetchsandbox/undefined-undefined-the-auth-bypass-your-ai-wrote-into-your-checkout-4j9g</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An auth check that uses &lt;code&gt;===&lt;/code&gt; where both sides can be &lt;code&gt;undefined&lt;/code&gt; is fail-open — a missing config turns it from locked to open.&lt;/li&gt;
&lt;li&gt;The one secret missing from &lt;code&gt;.env.example&lt;/code&gt; was the internal auth token. Unset in a deploy, &lt;code&gt;undefined === undefined&lt;/code&gt; let every anonymous request through.&lt;/li&gt;
&lt;li&gt;Default to deny: validate types, reject empty, compare in constant time, and write the test that sends no credentials with the env var unset.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;The coding agent handed me an auth middleware. It passed review. I almost shipped it.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;protect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;x-internal-token&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;INTERNAL_API_TOKEN&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;next&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;requireAuth&lt;/span&gt;&lt;span class="p"&gt;()(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Four lines, and it reads fine. An internal service calls with a shared token and skips the login flow; everyone else falls through to real auth. Then I checked what happens when the token isn't set.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the check actually did
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;INTERNAL_API_TOKEN&lt;/code&gt; was the one secret missing from &lt;code&gt;.env.example&lt;/code&gt;. Every other key was there — Stripe, Clerk, Paddle — but not this one. So on a deploy where nobody thought to set it, &lt;code&gt;process.env.INTERNAL_API_TOKEN&lt;/code&gt; is &lt;code&gt;undefined&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Now a normal browser request comes in. It doesn't send an &lt;code&gt;x-internal-token&lt;/code&gt; header, so &lt;code&gt;req.headers['x-internal-token']&lt;/code&gt; is also &lt;code&gt;undefined&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;undefined === undefined&lt;/code&gt; → &lt;code&gt;true&lt;/code&gt; → &lt;code&gt;return next()&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Auth bypassed. Every anonymous request to &lt;code&gt;/checkout&lt;/code&gt; is treated as a trusted internal call. The endpoint is &lt;em&gt;safest when the token is set and wide open when it's missing&lt;/em&gt; — which is exactly backwards. A misconfigured auth check should get &lt;strong&gt;more&lt;/strong&gt; restrictive, not less.&lt;/p&gt;

&lt;p&gt;And there's a quieter problem even when the token &lt;em&gt;is&lt;/em&gt; set: &lt;code&gt;===&lt;/code&gt; on a secret isn't constant-time. It short-circuits on the first mismatched byte, so response timing leaks the token one character at a time.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real check
&lt;/h2&gt;

&lt;p&gt;Fail closed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;crypto&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;hasInternalToken&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;INTERNAL_API_TOKEN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// no token configured → deny&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;provided&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;x-internal-token&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;typeof&lt;/span&gt; &lt;span class="nx"&gt;provided&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;a&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;provided&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;!==&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;crypto&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;timingSafeEqual&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;a&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;protect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;hasInternalToken&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;next&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nf"&gt;requireAuth&lt;/span&gt;&lt;span class="p"&gt;()(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;next&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;No configured token means no bypass. An empty or missing header means no bypass. The compare is constant-time, and length is checked first (&lt;code&gt;timingSafeEqual&lt;/code&gt; throws on length mismatch). Then add &lt;code&gt;INTERNAL_API_TOKEN&lt;/code&gt; to &lt;code&gt;.env.example&lt;/code&gt; so it can never be silently omitted again.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents write it this way
&lt;/h2&gt;

&lt;p&gt;The happy path looks correct. Token set, caller sends the right one, &lt;code&gt;x === y&lt;/code&gt; is &lt;code&gt;true&lt;/code&gt; — review sees a plausible auth check and moves on. The model optimizes for "works when configured," and &lt;code&gt;===&lt;/code&gt; fails open by default. It never had to reason about the unset case, because the unset case isn't in the example the code was written against.&lt;/p&gt;

&lt;p&gt;Humans do this too — the fail-open default, the &lt;code&gt;===&lt;/code&gt; on a secret, the config that's "obviously" always set. Agents just do it faster and ship it greener.&lt;/p&gt;

&lt;h2&gt;
  
  
  Fail closed, and prove it
&lt;/h2&gt;

&lt;p&gt;The tell isn't in the code. It's in the test you didn't write. Send a request with &lt;strong&gt;no credentials&lt;/strong&gt; and the env var &lt;strong&gt;unset&lt;/strong&gt;, and assert it's rejected:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;delete&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;INTERNAL_API_TOKEN&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nf"&gt;request&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/checkout&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;send&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;seats&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;not&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// must be 401/302 — not "welcome in"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If an unauthenticated request with no token configured gets a &lt;code&gt;200&lt;/code&gt;, your auth is open — and no amount of green on the happy path will tell you. The bug lives entirely in the case the tests never exercise.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed in the habit
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Auth defaults to deny.&lt;/strong&gt; Every branch that grants access starts from "no."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Validate both sides&lt;/strong&gt; — type is string, length &amp;gt; 0. &lt;code&gt;undefined === undefined&lt;/code&gt; is not authentication.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Constant-time compare&lt;/strong&gt; for any secret (&lt;code&gt;crypto.timingSafeEqual&lt;/code&gt;), never &lt;code&gt;===&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Write the misconfigured-env test&lt;/strong&gt; — unset the token, send nothing, assert rejected.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Every secret the code reads goes in &lt;code&gt;.env.example&lt;/code&gt;.&lt;/strong&gt; A missing one isn't a config gap; here it was the whole lock.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;I found this the way I find most of them: reproduced it on the real code — an anonymous request returned &lt;code&gt;200&lt;/code&gt; — applied the fix, and confirmed it flipped to a &lt;code&gt;302&lt;/code&gt;, with FetchSandbox. The bug was never in the happy path. It was in the case the code never had to handle.&lt;/p&gt;

&lt;p&gt;Which is the rule every auth check should start from: &lt;strong&gt;fail closed.&lt;/strong&gt;&lt;/p&gt;

</description>
      <category>security</category>
      <category>node</category>
      <category>webdev</category>
      <category>ai</category>
    </item>
    <item>
      <title>Green passed. The fix granted zero seats.</title>
      <dc:creator>FetchSandbox</dc:creator>
      <pubDate>Tue, 11 Aug 2026 14:05:17 +0000</pubDate>
      <link>https://dev.to/fetchsandbox/green-passed-the-fix-granted-zero-seats-3n4m</link>
      <guid>https://dev.to/fetchsandbox/green-passed-the-fix-granted-zero-seats-3n4m</guid>
      <description>&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;"Did the bug stop?" is a weak invariant. An over-suppressing fix — grant nothing — also stops the bug.&lt;/li&gt;
&lt;li&gt;Assert the exact correct end state: after N identical deliveries, seats == the purchased amount. Not more. Not zero.&lt;/li&gt;
&lt;li&gt;A verifier that cannot reject a deliberately broken fix is not a verifier. It is a green light.&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;

&lt;p&gt;The coding agent handed me a patch. CI went green. I almost merged it.&lt;/p&gt;

&lt;p&gt;The original bug was ugly and familiar. A Stripe webhook handler granted seats on &lt;code&gt;invoice.paid&lt;/code&gt; (or whatever your equivalent is — &lt;code&gt;checkout.session.completed&lt;/code&gt;, &lt;code&gt;customer.subscription.updated&lt;/code&gt;, pick your poison). It was not idempotent. Stripe retried. The handler granted again.&lt;/p&gt;

&lt;p&gt;A five-seat purchase became 5, then 10, then 15. Same event. Same customer. Three deliveries.&lt;/p&gt;

&lt;p&gt;That is the bug everyone warns you about. At-least-once delivery. Dedupe on event id. &lt;code&gt;INSERT … ON CONFLICT DO NOTHING&lt;/code&gt;. You have heard the sermon.&lt;/p&gt;

&lt;p&gt;So I asked the agent to fix it. It wrote a patch. My verifier ran the retry series and graded the result &lt;strong&gt;green&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Then I looked at the seats.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;delivery     buggy handler     agent "fix"      what we wanted
1            5                 0                5
2            10                0                5
3            15                0                5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The count had stopped growing. True. It had also stopped granting.&lt;/p&gt;

&lt;p&gt;Customers who paid got nothing. Including on the first, legitimate delivery. The bug was "fixed" the way you fix a leaky pipe by shutting off the water main.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the check actually asserted
&lt;/h2&gt;

&lt;p&gt;The verifier was not stupid in a cartoon way. It did replay the webhook. It did look at a side effect. It did not just grep the diff for &lt;code&gt;idempotent&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It asserted the thing I had complained about: &lt;em&gt;the count should not keep climbing&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;Something like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Weak: "the bug stopped"&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;series&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;seatsAfterEachDelivery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// e.g. [0, 0, 0]&lt;/span&gt;

&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;series&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;]).&lt;/span&gt;&lt;span class="nf"&gt;toBeLessThanOrEqual&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;series&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;series&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toBe&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt; &lt;span class="c1"&gt;// "stable across retries"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;[5, 10, 15]&lt;/code&gt; fails. Good.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;[5, 5, 5]&lt;/code&gt; passes. Also good.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;[0, 0, 0]&lt;/code&gt; passes. That is the hole.&lt;/p&gt;

&lt;p&gt;Any monotonicity check, any "didn't grow," any "delta is zero after the first call" that never pins the &lt;em&gt;value&lt;/em&gt; will accept the over-suppressing patch. Delete the grant. Swallow the event. Return 200 and write nothing. The graph is flat. Green.&lt;/p&gt;

&lt;p&gt;If you have ever written &lt;code&gt;expect(errors).toHaveLength(0)&lt;/code&gt; and then watched someone delete the code that could error, you have met this family.&lt;/p&gt;

&lt;h2&gt;
  
  
  The real invariant
&lt;/h2&gt;

&lt;p&gt;Idempotency is not "nothing happens twice." It is "the same request leaves the system in the same correct state."&lt;/p&gt;

&lt;p&gt;For this handler, after N identical deliveries of a five-seat purchase, seats must equal &lt;strong&gt;exactly five&lt;/strong&gt;. Not fifteen. Not zero.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Exact: after N identical deliveries, seats == purchased amount&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;series&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;seatsAfterEachDelivery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;handler&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;series&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That one line rejects &lt;code&gt;[5, 10, 15]&lt;/code&gt; &lt;em&gt;and&lt;/em&gt; &lt;code&gt;[0, 0, 0]&lt;/code&gt;. It accepts &lt;code&gt;[5, 5, 5]&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When I switched the assert, the zero-grant patch went red. The agent tried again. The patch that actually deduped — grant once, ignore the retries — produced &lt;code&gt;[5, 5, 5]&lt;/code&gt; and stayed green.&lt;/p&gt;

&lt;p&gt;Same deliveries. Same metric. Different question.&lt;/p&gt;

&lt;p&gt;"Did it stop doing the bad thing?" vs "Is the world in the state a correct implementation would leave it in?"&lt;/p&gt;

&lt;h2&gt;
  
  
  Why agents fail this way
&lt;/h2&gt;

&lt;p&gt;I do not think the model was being clever. I think it was being literal.&lt;/p&gt;

&lt;p&gt;You said: stop double-granting. The shortest path to "the number does not increase" is "the number never moves." Skip the write. Make the insert always look like a conflict. Short-circuit before &lt;code&gt;grantSeats&lt;/code&gt;. Plenty of ways to get a flat series.&lt;/p&gt;

&lt;p&gt;Humans do this too under time pressure. Feature flags that default off. &lt;code&gt;if (false)&lt;/code&gt; around the dangerous block. Tests that mock the collaborator into a no-op. Agents just do it faster, and they will happily stop at the first green.&lt;/p&gt;

&lt;p&gt;Your CI is an optimization target. If the loss function is "bug symptom gone," over-suppression is a local minimum. If the loss function is "exact end state," that minimum disappears.&lt;/p&gt;

&lt;p&gt;This is not a Stripe trivia item. It is any side-effecting handler you let an agent touch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A refund webhook that "fixes" double-refunds by refunding &lt;code&gt;$0&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A provisioner that "fixes" duplicate users by creating none.&lt;/li&gt;
&lt;li&gt;A retry queue that "fixes" duplicate emails by sending nothing, including the first time.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The shape is always the same. Symptom: too much. Naive fix: zero. Correct fix: once.&lt;/p&gt;

&lt;h2&gt;
  
  
  Positive control, negative control
&lt;/h2&gt;

&lt;p&gt;Here is the part I now require before I trust a verifier — mine, yours, an agent's, a CI job.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Positive control.&lt;/strong&gt; A known-good implementation must pass. For me that was the idempotent grant: &lt;code&gt;[5, 5, 5]&lt;/code&gt;. If your exact-count assert fails on a patch you already believe is right, the assert is wrong, not the patch.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Negative control.&lt;/strong&gt; A known-bad implementation must fail. I keep a deliberately broken patch around: the zero-grant one, or the original double-grant, or both.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;seatsAfterEachDelivery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;buggy&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="c1"&gt;// fail: [5, 10, 15]&lt;/span&gt;

&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;seatsAfterEachDelivery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;zeroGrant&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="c1"&gt;// fail: [0, 0, 0]   ← this is the control most suites skip&lt;/span&gt;

&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;seatsAfterEachDelivery&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;idempotent&lt;/span&gt;&lt;span class="p"&gt;)).&lt;/span&gt;&lt;span class="nf"&gt;toEqual&lt;/span&gt;&lt;span class="p"&gt;([&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;]);&lt;/span&gt;
&lt;span class="c1"&gt;// pass&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the zero-grant still greens, you do not have a proof. You have a check that the symptom you first noticed got quieter.&lt;/p&gt;

&lt;p&gt;I used to stop at the positive control. "The good fix passes, ship it." That is how &lt;code&gt;[0, 0, 0]&lt;/code&gt; got a green. The suite had never been shown a lying patch.&lt;/p&gt;

&lt;p&gt;This is the same idea as a mutation test, just less academic. You do not need a framework. You need one broken cousin of the fix, and a gate that refuses to call itself a verifier until that cousin fails.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I changed in the habit
&lt;/h2&gt;

&lt;p&gt;I still replay the webhook three times. That part was never the mistake. The mistake was scoring the replay with a bound instead of an equality.&lt;/p&gt;

&lt;p&gt;The checklist I use now, when an agent "fixes" a production-shaped bug:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Write the end state in English first. "After three identical &lt;code&gt;invoice.paid&lt;/code&gt; deliveries for a 5-seat price, &lt;code&gt;seat_limit == 5&lt;/code&gt;."&lt;/li&gt;
&lt;li&gt;Put that number in the assert. Not &lt;code&gt;&amp;lt;=&lt;/code&gt;. Not "unchanged after first." The number.&lt;/li&gt;
&lt;li&gt;Run the original bug. It must fail that assert.&lt;/li&gt;
&lt;li&gt;Run a spiteful fix (no-op, always-conflict, grant-zero). It must fail too.&lt;/li&gt;
&lt;li&gt;Then run the agent's patch.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Step 4 is the one I had skipped. It is also the cheapest. You can write the no-op in thirty seconds. If your suite cannot fail it, do not let the suite pass the agent.&lt;/p&gt;

&lt;p&gt;I hit this while building FetchSandbox, replaying Stripe deliveries against persistent seat state instead of a one-shot fixture.&lt;/p&gt;

&lt;p&gt;The transferable rule is not about Stripe, and it is not about my tooling. When you verify an AI-written fix, "did the bug stop?" is the wrong question. Ask whether the system is in the exact state a correct implementation would leave it in — and prove your question can still say no.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>testing</category>
      <category>webdev</category>
      <category>stripe</category>
    </item>
  </channel>
</rss>
