I ran Pawsome Pantry through an integration check on October 2. Lovable-built app, Paddle and Resend wired up, temporary FetchSandbox twins, some manual setup required. Six checks passed. Two recovery paths were unproven. The app version was declared but not independently attested. The suite verdict was Inconclusive. Both of those are the correct outcomes.
Here's what the six covered: R1 confirmed a first purchase simulated receipt was created. R2 confirmed a duplicate event didn't trigger an extra receipt. R3 confirmed the purchase and recipient matched. R4 confirmed distinct purchases each got their own receipt. R6 confirmed concurrent redelivery produced only one receipt. R8 confirmed a delayed redelivery matched the original. All six were observed inside recorded windows against the real app.
What the two gaps actually required
R5 needed a failed callback combined with a retry during a withheld accepted response. That specific fault and recovery sequence was not observed in this run. R7 needed transient faults to occur and be recorded. That sequence was not observed either.
Not observed is not the same as didn't happen. Absence of evidence in a recorded timeline means those paths remain unproven, not that the recovery failed. You can't mark a check passed when the required failure and recovery weren't independently seen in the timeline.
The emails in this run were simulated, not delivered to real inboxes.
Three questions worth asking before trusting any result
- Did the injected failure really occur in the recorded timeline?
- Was the recovery independently observed, not just inferred from a 200 status?
- Is the final state correct at the data level, not just the response code?
Six greens with two unobserved gaps is a partial picture. The full recorded case, with context on the Paddle and Resend twin setup, is at fetchsandbox.com/blog/lovable-pawsome-pantry-paddle-resend-twins. The run receipt is at fetchsandbox.com/runs/a656a3dbf1.
Before trusting agent-written integration changes, you need inspectable evidence. Not a green status in the chat, a timeline where you can see the failure, the recovery, and the final state. To rerun a fresh validation session against the same app after any identified correction, FetchSandbox MCP is the connect point.
If you've built something in Lovable with Paddle or Resend: what evidence would you want to see before trusting a recovery check?
Top comments (0)