DEV Community

Cover image for What My AI Broke This Week #1: Nothing. It Was Me, Three Times
Kalislav Smirnov
Kalislav Smirnov

Posted on AI-assisted

What My AI Broke This Week #1: Nothing. It Was Me, Three Times

This week the agent broke nothing. I broke three things, and it spent the week politely proving it.

The setup: one mini-PC with Proxmox in the corner, a pile of side projects, Claude Code as the pair.

 

DNS is down (DNS was fine)

Saturday night, half past midnight. Nothing at home resolves, so I type the obvious: "something's wrong with DNS on the server, can you look?"

Three minutes later, the verdict. DNS was fine. One name didn't open, and it was the new monitoring dashboard. The reverse proxy had 11 hosts and that one wasn't among them, so TLS died with unrecognized name. The setup plan had a step 5, "add proxy host in the UI". Someone skipped it. (Me.)

It was me all along

The agent didn't add it either. It had no login to the proxy and refused to hand-edit the config, because the proxy rewrites that file from its own database. Correct, and slightly annoying, which is the right amount.

Then the real outage, the five minutes when nothing resolved. The logs:

04:31     Tailscale on the server changed its external addresses
04:35:09  51 DNS connections from the Mac dropped in the same second
          AdGuard: 4 days up, 0 restarts, 1 upstream timeout in 24 h
Enter fullscreen mode Exit fullscreen mode

My Mac had "Use Tailscale DNS" switched on. So every lookup left the laptop, went through the tunnel to a relay in another country, came back through the subnet router, and landed on an AdGuard box three metres from my chair. When the relay hiccuped, DNS went with it. Every other device at home talked to AdGuard directly and never noticed.

My Mac's DNS queries, a Tailscale relay in another country, and the AdGuard box in the same room

What I'd tell another engineer: before you blame the DNS server, draw the path your query actually takes. Mine had a layover.

 


 

Out of memory (we had 16 GB)

Monday evening. "The server has almost no RAM left. Find what's eating it and what we can stop."

The agent stopped nothing. The host showed 25 of 31 GB used, with memory pressure at zero and 5.5 GB free. The number was inflated by config, and the config was mine.

VM            allocated   held on host   guest actually uses
side-project  8 GB        8.3 GB         1.7 GB
render box    6 GB        6.2 GB         0.5 GB   (idle since 2 Oct)
ZFS ARC       up to 3.1   3.0 GB
Enter fullscreen mode Exit fullscreen mode

Every VM had balloon: 0. A guest fills its page cache once, after a render or a build, and never hands the memory back, so the host charges the full allocation forever. That 6 GB box renders video for a side project and had last done anything three days earlier.

Is this out of memory?

The fix took five minutes: smaller allocations with ballooning on, ARC capped at 2 GB, a 7.8 GB zram swap. The host went from 25 GB used to 14, with 16 available. Then I said the words: "better commit and apply."

The plan wanted to create a VM that was already running. The side-project VM wasn't in Terraform state at all, so a full apply would have built a second copy on top of the live one, imported five LXC containers, rewritten the tunnel config and added four DNS records on the way. The agent added an import block, applied with -target on the two VMs it could account for, and got "No changes" on the next plan. The rest of the drift is still there, waiting for a calmer evening.

What I'd tell another engineer: state is a story the repo tells itself. Read the plan like it's lying to you.

 


 

It's the network (it was the uptime)

Some Claude Code sessions worked. Others died with ECONNRESET on every retry. I asked the working one to debug the broken one, which is a sentence I didn't expect to write.

The dead sessions had screenshots in them. Lots. One had 28 images, about 17 MB of base64, and Claude Code resends all of it on every turn. The healthy chats carried 2 to 9 images, around 3 MB. A curl test: 1 MB fine, 20 MB fine, 8 MB died mid-stream with an HTTP/2 error, 35 MB rejected with a 413. The network was fine. Multi-megabyte uploads sometimes weren't.

I asked whether we could report this to Anthropic and maybe get paid. The agent explained, kindly, that the bounty covers security bugs and this was a networking inconvenience. Then it found an open issue with the same picture (46 images), drafted a comment with our numbers, and posted it from my account after I said "go".

Last week I told everyone to write up their bugs instead of their feelings. Here I am doing it in a GitHub thread at nine on a Sunday morning, so at least I'm consistent.


By lunchtime a bigger session died, half a million tokens of context, and /compact died with it, because compaction sends the same giant request. Another issue, another thread of Mac users, and what fixed it for them was a reboot. One had 52 days of uptime; a restart brought back all eight of his dead sessions, one of them at 826k tokens.

The agent then checked mine. 102 days. I had been quietly proud of that number.

Reboot the Mac, or spend a morning proving to GitHub that the API drops my 17 MB of screenshots

What I'd tell another engineer: an agent can prove the network innocent ten different ways. It still won't tell you to turn it off and on again, because you taught it not to be rude.

Hello, IT. Have you tried turning it off and on again?

 

Your turn

What did your AI pair spend this week proving wasn't its fault, and was it right?

Top comments (0)