Usually, when a company spins up a server for a developer, a DevOps engineer or admin has to manually configure the firewall, SSH access, logging, security updates, and much more. It's usually the same set of steps and tools every time. The process is fairly involved, especially when you have to repeat it regularly for multiple developers. It's easy to forget something or make a mistake, and the cost of that mistake can be high.
In this article I want to share experience I've built up over the years, and after "hundreds of hard-earned lessons."
You need to set up a development environment on a VM in AWS. What tools do you need?
- Basic utilities. git, git-lfs, curl, wget, jq, vim, tmux, htop, make, unzip, ripgrep, fd, and direnv.
- zsh. A convenient shell. Worth setting as the default.
- Node.js 24.
- Docker and Docker Compose.
-
Traefik. A reverse proxy that runs in Docker. It gets an HTTPS certificate from Let's Encrypt and password-protects the site. This lets a developer show off a project at an address like
name.ec2.region.domain. - Kubernetes tools. kubectl, helm, and werf.
- GitHub CLI (gh). For working with GitHub from the terminal.
- Claude Code. An AI assistant for working with code.
- Python tools. uv and uvx.
- yq. Like jq, but for YAML files.
- Tailscale. Private access to the machine over VPN.
- NVIDIA Container Toolkit. Lets Docker use the GPU. Needed only if the machine has one.
Problems you can run into if you don't think about security
Let's look at the main ones.
AWS keys leaked. A developer installed an npm package with malware in it. The package quietly grabbed the keys off the machine. The next morning, the AWS bill shows crypto mining charges.
Countermeasure. Lock down the AWS metadata endpoint so only root can reach it, and enable IMDSv2. Issue short-lived AWS credentials that refresh every 15 minutes and live only in memory. Install Falco so it raises an alert the moment something tries to grab the keys.
The machine got hacked through an open port. Bots scan the internet around the clock. They find an unprotected server within minutes and start brute-forcing passwords.
Countermeasure. Set up an iptables firewall. Block all inbound traffic and open only SSH, HTTP, HTTPS, and Tailscale. Disable password login and root login. Install fail2ban to block an IP after 5 failed attempts. Turn on kernel network hardening settings (sysctl).
Everyone has root. Any mistake or piece of malware instantly gets full control of the machine. Anything can be deleted, and the tracks can be covered.
Countermeasure. Remove sudo rights from the developer. Run Docker rootless, so a compromised container doesn’t hand over control of the machine. Configure Falco to alert on every sudo call.
Nobody knows what happened. There are no logs. Or they lived on the machine that’s already been deleted. There’s nothing left to investigate.
Countermeasure. Ship all logs to Grafana Cloud via Grafana Alloy. The log lives separately and survives after the machine is deleted. Log every sudo attempt separately. Set up chrony for accurate timestamps in the logs.
Configuration drift. Someone tweaks something by hand “just for five minutes.” Six months later, 20 machines are configured 20 different ways.
Countermeasure. Define all settings as code and install the CINC client. Every 30 minutes it checks the machine against the reference config and puts everything back in place.
Forgotten updates. A vulnerability is found in OpenSSH or Docker. Updating everything by hand takes forever, and one machine is bound to get missed.
Countermeasure. Enable unattended-upgrades for daily automatic installation of security patches. Pin critical tools to specific versions and verify them with checksums.
An employee leaves the company. Their SSH key is still sitting on five servers. Nobody knows exactly which ones.
Countermeasure. Keep developers’ SSH keys in a single file alongside the list of machines. That way it’s immediately clear who has access to what and where. Don’t give admins permanent keys. They get a temporary key through EC2 Instance Connect that stops working after 60 seconds.
Everything depends on one person. All the security knowledge lives in one admin’s head. They leave, and the team has no idea how anything is set up.
Countermeasure. Keep the entire configuration as code (CINC). Any team member can read how a machine is configured and reproduce it.
Problems with clients and audits. A client asks how you secure your infrastructure. There’s nothing to show, and the deal stalls.
Countermeasure. Show the audit log in Grafana, ready-made alerts for suspicious events, and Falco’s reports. That’s ready-made proof the security actually works.
An easy, reliable solution
I believe I promised a quick way to do this. Well, here it is.
The Idlefy team built a free, open-source project with a ready-made, up-to-date set of solutions and tools. You can check it out here.
You don't need to worry about your data's security. Idlefy has no direct access to the server. Control happens only through a special tag and AWS access permissions.
So, what's it going to be? Reinvent the wheel, or take one that's already built?

Top comments (0)