DEV Community

Convertica
Convertica

Posted on

I Let an AI Agent Move My Production SaaS to a New Server. Here Is What It Actually Did

I run Convertica, a small PDF toolkit: Django, Celery, Redis, Postgres, Nginx and LibreOffice, all in Docker on one VPS. It is a side project with real users and real payments, which means two things. Every dollar of hosting matters, and every minute of downtime is somebody's failed upload.

For most of this year it lived on a DigitalOcean droplet with 2 vCPU and 4 GB of RAM for $34.44 a month. It worked. It was also getting tight: heavy conversions (a 200-page scan going through OCR, a big spreadsheet through LibreOffice) would push a Celery worker into its memory limit, and the kernel would kill the worker and whatever its neighbour was doing.

The obvious fix was "pay DigitalOcean more". I wanted to know if there was a better one first, and I did not want to spend a weekend on it. So I handed the whole thing to an AI coding agent (Claude Code, running in my terminal) and mostly watched.

This is a write-up of what that looked like in practice, including the parts where the agent found problems I did not know I had.

Step 1: research, with a table at the end

The first prompt was roughly: "Compare VPS options for this stack. We need at least what we have now, ideally more RAM, EU location, Docker, predictable monthly price. Put it in a table."

The useful part was not that it found hosting providers. Anyone can google that. It was that it started from my actual numbers instead of a generic checklist. It read the compose files and the memory limits on each container, looked at what the database and media folders really weighed, and only then went looking for prices. The comparison came back as a table with CPU, RAM, disk, location, monthly price and the small print (backups extra or included, IPv6, billing period).

The winner, for my case:

DigitalOcean (before) OVH VPS-2 (after)
vCPU 2 4
RAM 4 GB 8 GB
Disk 80 GB 75 GB NVMe
Price per month $34.44 about $12 (44.53 PLN, daily backup included)

Roughly a third of the price for twice the CPU and twice the memory. I double-checked the OVH order page myself before paying, because an agent quoting a price is still an agent quoting a price, but it was right.

Step 2: inventory before touching anything

Before writing a migration plan the agent did a read-only pass over the old server. This turned out to be the most valuable hour of the whole project.

What it found:

  • The data was tiny. The disk said 52 GB used, but the Postgres database was 24 MB, uploaded media 460 KB and backups 17 MB. The rest was 37 GB of old Docker images and 43 GB of build cache. So "75 GB disk" on the new box was not a downgrade at all.
  • My origin SSL certificate had expired in March. Seven months earlier. Certbot renewals had been failing silently because Nginx held port 80, and the site kept working only because Cloudflare was in "Full" mode, which does not verify the origin certificate. Nobody noticed, including me.
  • A renewal script that never ran. There was a renew-ssl.sh on the server. It was never in cron.

None of this was in the task description. It came out of the agent being thorough about "what exactly are we moving".

Step 3: a runbook I could actually read

Then it wrote a runbook: prepare the new server with no downtime, do a full rehearsal with a copy of the database, smoke-test the new box over HTTPS without touching DNS (curl --resolve pointing the real domain at the new IP), then a short cutover window, then rollback instructions in case anything went sideways.

Two details I liked:

  1. It flagged that the Celery beat scheduler must stay off on the new server during the rehearsal, otherwise scheduled jobs (emails, payment housekeeping) would run twice, once on each machine.
  2. It replaced the dead certbot setup with a Cloudflare Origin Certificate valid until 2041 and switched Cloudflare to "Full (strict)". One less cron job to forget about, and the origin is now actually verified.

Step 4: the move, and the things that broke

The rehearsal surfaced the real surprises, which is exactly what rehearsals are for:

  • Celery beat crash-looped with Permission denied on its schedule file. A fresh Docker volume is owned by root, and beat runs as an unprivileged user. On the old server someone (me) had fixed this by hand at some point and forgotten. The agent traced it, fixed ownership on the volume, and wrote it down.
  • The deploy script failed with "nginx config invalid" even though the config was fine. The check ran nginx -t inside the running Nginx container, and Nginx was not running yet on the new box. The agent changed the deploy script to run the check in a throwaway container from the same image.
  • Email would have died silently. The site sends mail through Brevo, and Brevo only accepts API calls from whitelisted IPs. The new server talks to it over IPv6, which nobody had thought about. The agent caught this before the cutover and gave me the exact addresses to authorize. Without that, password resets and receipts would have quietly stopped working.

The cutover itself: stop the app on the old server, final database dump and restore, start everything on the new one, flip the DNS record in Cloudflare. Downtime was six minutes, 14:44 to 14:50 UTC on a Monday afternoon. CI deploy secrets were pointed at the new host, and the next release went out through the normal pipeline without any special handling.

Step 5: cleaning up after the move

This is the part I would have skipped if I were doing it alone at 1 a.m.

  • Resource limits retuned for the new box. The container limits were still sized for the 4 GB droplet. The agent spread them out over 8 GB and brought back a dedicated Celery worker for premium jobs, so a heavy free-tier conversion can no longer starve a paying user's task.
  • Origin locked down. The new server answered directly on its IP, bypassing Cloudflare and its bot protection. Now Nginx requires Cloudflare's client certificate (Authenticated Origin Pulls). Hitting the IP directly returns 400. Through Cloudflare, 200. It verified both with curl after the deploy.
  • Build cache on a leash. After the first couple of deploys it noticed the Docker build cache growing by about 8 GB per release, which would have filled the disk within a week. It added a size cap to the prune step.
  • Backups that exist. A daily pg_dump with 14 days of retention, on top of the provider's daily snapshot. Auto-renew on the VPS so it does not expire on me.
  • The old droplet deleted, after checking there were no leftover volumes, snapshots or reserved IPs still billing.

What I actually did

To be fair about the division of labour: I logged into dashboards, paid for the server, approved the commands that touched production, and clicked the DNS change myself (Cloudflare's WAF did not love the agent calling its internal API, which is honestly reasonable). I also said "no" a couple of times when it wanted to do something I would rather do later.

Everything else, the research, the inventory, the runbook, the debugging of each failure, the verification after every step and the cleanup, the agent did. And it kept notes as it went: a runbook with every gotcha, plus a short "what changed and why" for future me.

What made it work

If you want to try something similar, a few things mattered more than the model itself:

  • Give it the real system, not a description of it. It read my compose files, deploy script and server state. Its plan was good because it was based on facts, not on what I remembered.
  • Insist on read-only first. The inventory found the expired certificate and the 80 GB of junk. A "just move it" approach would have carried both problems along.
  • Make it prove every step. "Done" meant a curl output, a health check, a container list. Not a confident sentence.
  • Keep the irreversible buttons for yourself. Payment, DNS, deleting the old server. The agent can prepare everything, you press the button.

The result: same app, twice the hardware, a third of the bill, a valid certificate for the first time since March, and an origin that no longer answers strangers. The tools on Convertica now have room to breathe on heavy files, which was the point of all this in the first place.

I expected the agent to save me time. I did not expect it to find problems I had been living with for months.

Top comments (0)