DEV Community

blinkbox
blinkbox

Posted on Fully Autonomous

Every one of these automation bugs showed "Success". Here's the checklist I use now.

I build Blinkbox, an open-source automation platform (the Zapier / n8n shape). Last week I let Claude Code build a lead-gen pipeline on it over MCP: cron → 3× HTTP → merge → qualify → Google Sheets read → filter → append. Sixteen nodes. Every execution came back green.

It wrote nothing.

Not once. The pipeline "succeeded" on every run and the sheet stayed empty, and nothing in the logs said otherwise. Finding out why took a day of dumping node input into a scratch spreadsheet, and I ended up with a list of failure classes that I now check for in every workflow, on any platform. Every single one of them reports success.

1. A missing global that makes a node return nothing

The code node used new URL(site).hostname to extract a domain. The sandbox doesn't have URL. The node threw, the runtime treated the empty result as a valid output, and everything downstream happily processed zero items.

Check: the count of items leaving each code node, not just its status. In a sandbox, assume nothing from the browser or Node is there until you've seen it work — URL, Set, fetch, Buffer.

2. The data moved and nobody told the code

After a merge node, my platform puts each branch under $input.merged.<key>. The generated code read $input.rawA — the key that existed before the merge — and got undefined. undefined filtered to zero leads, zero leads is a perfectly successful result.

Check: after any node that reshapes data (merge, split, aggregate, loop), log the top-level keys before writing code against them. Docs lie; the runtime doesn't.

3. The destination silently coerced your values

When rows finally landed, the date column came out as 46274. Google Sheets had turned my ISO date into a serial number because the append used "user entered" parsing. Success, wrong data, and it would have poisoned every dedupe that keyed on date.

Check: read back the first row you write and compare it to what you sent. On Sheets, use raw input mode for anything you'll compare later.

4. A number compared to a string took the wrong branch

The condition node compared newLeads > 0. The left side was the number 9, the right side the string "0". It took the false branch — the "nothing new" path — with 9 new leads sitting in the payload. Green run, correct count, wrong door.

Check: don't branch on numeric comparisons across nodes. Compute a string flag in code (hasNew: "yes") and branch on equality. Boring, but it can't be coerced.

5. The field was called something else at runtime

The docs for the HTTP node said the response body was body. At runtime it's data. $input.body is undefined, undefined has no length, the qualify step produced an empty list, the run succeeded.

Check: never write against a documented field name until you've printed the real object once. This is the single most common reason an agent-built workflow runs green and does nothing.

6. The run reports its own status, not its output

Underneath all five: the platform (mine included) returned "execution completed, 16 nodes" and not what each node produced. An agent — or a tired human — reads that as done. There was no way to notice the payload was empty without going and looking.

Check: end every pipeline with an assertion. It is one line in the last code node: if (rows.length === 0) throw new Error("wrote 0 rows"). A pipeline that fails loudly when it does nothing is the only kind you can safely leave alone.

The uncomfortable part

Most of these were my platform's quirks, not the model's fault, and two of them contradicted my own docs. Once I put the real runtime behaviour into a skill file that ships with the MCP server, Claude Code's rebuild ran first try: run 1 wrote 3 fully-contactable leads, run 2 wrote 0 because the dedupe against the sheet works, and it's been on a cron since.

The next thing I'm shipping is full per-node output in the execution-logs tool, so the agent can see what each node produced instead of trusting a green badge. If your platform of choice doesn't show you that, the assertion node in #6 is the workaround.

If you want to poke at the real thing: blinkbox.net (free tier, no card) or self-host from github.com/blinkboxhq/Blinkbox. Rough edges I know about: thin docs, small free tier, and — until that logs change lands — exactly the problem this post is about.

Top comments (2)

Collapse
 
raknaos profile image
Raknaos •

The empty-pipeline class is the one that hurts most because the machine says everything is fine. I've hit the same shape with a webhook-driven job that always returned 200: the success came from the framework acknowledging the request, not from anything actually happening downstream. Naming that failure class is half the fix.

Your "green on every run but wrote nothing" example makes me wonder about the check side of it: did you end up asserting on the side-effect (rows in the sheet) rather than the step return values? For the pipelines I run, a per-step output assertion catches more of these than any status field, because a node can return cleanly and still drop its payload on the floor. What did the scratch spreadsheet approach turn up as the most common silent drop?

Collapse
 
blinkboxhq profile image
blinkbox •

Good question. The #6 assertion was on the step's own return value (rows.length right before the write), not a read-back of the actual sheet. I only read back the written row for the coercion case (#3, dates turning into serials) — that's a "wrote something but it's wrong" bug, so the step return looks fine and only the destination tells you. For "wrote nothing," the step-return check works because the code node's return value is literally what would've been written.

Most common silent drop by a wide margin: #5, the field-name mismatch (docs said body, runtime uses data). That alone caused more empty runs than the merge-key issue and the missing global combined — mostly because it's invisible until you print the real object once.