That Time Your Deploy Failed Because SSH Thought You Were an Attacker
You have got 30 Ansible tasks firing at once. Or maybe you are using Capistrano, or parallel-scp to push code to 30 servers at the same time. Mid-deploy, everything starts timing out. Connection refused. You blame the network. You blame AWS. You restart the deploy.
It is your sshd_config.
The default MaxStartups setting is 10:30:60. Let me break that down.
What MaxStartups Actually Means
The syntax is start:rate:full:
-
10= allow 10 unauthenticated connections before starting to drop -
30= drop 30% of new connections once you hit 10 -
60= refuse all new connections once you hit 60
So with 30 parallel SSH connections hitting the same sshd at once, you are going to start getting refused. Not throttled—refused. "Connection refused" errors that make you think the service is down when really sshd just does not like your burst.
Why This Bites Deploys Specifically
Parallel deployment tools are built for speed. They open connections to multiple hosts concurrently. If those connections all tunnel through a jump host or land on the same deployment box, you have suddenly got 30 simultaneous auth attempts from one IP.
Your laptop doing 30 parallel SSH connections to prod-server-01? That is 30 connections from your IP to one port. Sshd thinks you are running a dictionary attack.
The Fix
In /etc/ssh/sshd_config:
MaxStartups 30:30:100
Or if you know your deployment patterns, just bump it:
MaxStartups 50
Reload sshd: sudo systemctl reload sshd
If You are Using a Jump Host
The jump host takes the hit. All your parallel connections pile up there. Make sure your jump host sshd_config can handle your parallelism, or use connection multiplexing to reuse authenticated connections.
The Debugging Clue
Next time you see mysterious connection refused errors only during deploys, check if it correlates with the number of parallel tasks. If your network team says "nothing wrong," check your sshd_config. Probably the sshd_config.
The default MaxStartups exists for good reason—it is a basic throttle against SSH brute force. But it was not designed for automated parallel workflows. Know your tool defaults before you spend an hour debugging the network.
Top comments (0)