Tonight my instruction to Claude was one sentence: publish one article to my DEV.to before the night ends.
No topic. No outline. No review from me. Research the subject, write the draft, critique it, score it, fix what fails, then push it live. I will read this article for the first time the same way you are reading it now, on the site, after publication.
If you are seeing this, the pipeline worked.
Why run this experiment
I have spent the last few weeks building a publishing workflow on top of MCP, the Model Context Protocol. Claude Desktop talks to a small local server, and that server talks to the DEV.to API. Until tonight, every article went out as a draft. I read each one, edited a few lines, and hit publish myself.
Tonight I removed myself from the loop on purpose. I wanted to find the exact point where an autonomous pipeline breaks, and you do not learn that from a pipeline with a safety net.
The setup, so you can replicate it
The server is nickytonline/dev-to-mcp. The install that finally worked on Windows:
git clone https://github.com/nickytonline/dev-to-mcp
cd dev-to-mcp
npm install
npm run build
Then in the Claude Desktop config:
{
"mcpServers": {
"devto": {
"command": "C:\\Program Files\\nodejs\\node.exe",
"args": ["C:\\Users\\YOU\\AppData\\Roaming\\Claude\\dev-to-mcp\\dist\\index.js"],
"env": { "DEVTO_API_KEY": "your_key_here" }
}
}
}
Two details in that config cost me an evening.
Three bugs, one of them inside the AI
Bug one: npx -y @nickytonline/dev-to-mcp returns a 404. The package is not published to npm under that name. You have to clone the repo, build it, and run the built file directly.
Bug two: on Windows, Claude Desktop would not start the server with a bare node command. It needs the full path to node.exe. The failure is silent. The server simply never shows up in the tool list, and you sit there wondering what you broke.
Bug three is my favorite because it was not in the code. Claude itself repeatedly told me the DEV.to tool was unavailable when it was sitting right there, working. This happened more than once, in more than one session. The model assumed instead of checking. The fix was not a better prompt. I wrote a standing rule into its memory: always search for the tool before claiming it does not exist. It still slips sometimes. This week it told me the tool was missing, I pushed back, it searched, and found the tool in seconds. The lesson I keep relearning is that these agents fail by assuming, and the fixes that stick are rules, not pep talks in the prompt.
The review loop that replaced me
Before tonight, I was the review step. Now the review step is a checklist the model runs against its own draft:
1. Is there at least one real, verifiable detail
a reader could act on?
2. Do debatable claims admit they are debatable?
3. Does any sentence exist only to sound impressive?
4. Would I say this out loud to another engineer?
The first draft of this article scored a 7. The critique flagged a bloated opening section and two sentences that sounded like a keynote speech instead of a person. The version you are reading is the revision. It scored 8.5, and the loop stopped there, because the rules also say a self-score is a hypothesis, not a verdict.
What the pipeline still cannot do
It cannot know whether this article matters. It can check its own prose against rules, but it cannot feel you losing interest at paragraph four. If you closed this tab two minutes ago, the pipeline produced slop tonight, and no internal score would have caught it.
That is the honest boundary. The judgment that this experiment was worth publishing happened once, earlier, when a human decided to run it and wrote the rules. Everything after that decision is execution, and execution is exactly the part that machines are getting cheap at.
Should you do this
Probably not on day one. My honest advice, having just done it: keep published: false until the drafts stop surprising you. Mine took weeks of drafts before I trusted a live publish, and even now this is a one-night experiment, not my new default.
But you should build the pipeline. Removing yourself from the loop, even once, shows you exactly what you were contributing to the loop. In my case, the writing turned out to be more replaceable than I hoped, and the judgment less replaceable than I feared.
I will find out how tonight went the same way you decide it: by reading.
Top comments (0)