I'm Hisashi. I've run my own company for 22 years, and from 2011 to 2023 I ran it while moving between countries. I keep a Mac in my Tokyo office running as an always-on AI server.
This summer in Tokyo, my office Mac started running a multi-day job at 100% GPU, and the more interesting problem became heat. This post is about something every always-on home AI server eventually faces: it gets hot, and you're not there to watch it.
That is fine in an air-conditioned room in spring. It is not fine in a Tokyo summer when I'm in another country and the office aircon is off. Two things can go wrong: the chip thermally throttles and my job quietly slows down, or it just runs hotter than I'd like for weeks on end. I had no way to see either, because I was never in the room.
So I built two things: a dashboard that lets me watch the machine's temperature from anywhere, and an automation that turns the room's air conditioner on when the GPU gets too hot. Total cost: zero, plus a SwitchBot hub I already owned.
1. Reading the sensors without sudo
The first obstacle on Apple Silicon is that the obvious tools need root. powermetrics wants sudo, and I didn't want an automation that hangs forever on a password prompt while I'm on the other side of the planet. So I found a sudo-free source for each metric.
-
CPU usage:
top -l 1 -n 0, then 100 minus the idle percentage. -
GPU usage:
ioreg -r -d 1 -c IOAccelerator. It exposes a "Device Utilization %" field with no privileges required. This is how I confirmed the GPU really was pinned at 97 to 98%. -
Temperatures:
smctemp, a tiny open-source SMC reader you compile with the clang already on your Mac (no Homebrew needed).-cgives CPU temperature,-ggives GPU. -
Fan RPM: the same SMC interface. There's no dedicated flag, but dumping all keys and reading
F0AcandF1Ac(fan 0 and fan 1, actual RPM) works. Mine read about 5,400 and 5,760, against a ceiling near 5,350. In other words, both fans were already flat out.
A 15-line shell script samples all of that once a minute and appends one JSON line to a file. CPU, GPU, both temperatures, both fan speeds, plus the current job's progress. Nothing heavy: the whole sample takes a fraction of a second of CPU and never touches the GPU, so it can't interfere with the work it's measuring.
2. An MRTG-style dashboard I can open from anywhere
Old-school sysadmins will remember MRTG: a web page of rolling time-series graphs you glance at to know your machine is healthy. That's exactly what I wanted, minus the RRDtool setup.
The page is one static HTML file with Chart.js bundled locally (no CDN dependency, so it works even if the office internet hiccups). It fetches the JSON log, draws three graphs (usage, temperature, fan RPM), shows the latest values as cards, and refreshes every 60 seconds. Room temperature and the air-conditioner state sit in the corner.
Serving it is one line: python3 -m http.server. I run both the collector and the server as launchd jobs so they survive reboots and restart themselves if they die.
The part that makes it genuinely useful from a Bangkok hotel is how I reach it. The machine is on my Tailscale network, so from any of my devices I just open:
-
http://m1from anywhere in the world, over the Tailscale tailnet. -
http://m1.localwhen I'm on the office LAN.
Both are plain hostnames, no port number, no IP to remember. Tailscale's MagicDNS gives me the global name; mDNS gives me the local one. I bound the server to port 80 so neither needs a :8765 suffix.
What I see when I open it: a row of live cards (CPU and GPU usage, CPU and GPU temperature, both fan RPMs, room temperature, and the current training step), sitting above three rolling graphs (usage, temperature, fan speed). The temperature cards turn amber as they climb, so a glance from my phone tells me whether the machine is comfortable or cooking. Everything updates every 60 seconds.
Watching the graphs for an evening settled a question I'd assumed I'd have to guess at: was the machine throttling?
It wasn't, yet. Job step time stayed flat while the GPU held 86 to 88°C, which means the chip was holding its clocks. Good. But the fans told the real story: they were already maxed. The cooling system had no headroom left. The only lever I had not pulled was the temperature of the air going into the machine.
That reframed the whole problem. I couldn't make the Mac cool itself harder. I could only make the room cooler. And the room has an air conditioner.
3. Giving the room a thermostat
The details are where you avoid an automation that drives you crazy:
- Hysteresis. Cooling turns on when the GPU stays at or above 92°C for three straight minutes, and off only when it drops to 84°C or below. The 8-degree gap stops it flapping on and off around a single threshold.
- A rate limit. No more than one action every ten minutes, so it can't machine-gun the compressor.
- It only undoes its own actions. The controller turns the aircon off only if it was the one that turned it on. If I switched it on myself for comfort, it leaves it alone. Infrared is fire-and-forget with no state feedback, so the controller tracks what it last commanded rather than pretending to know the real state.
- A floor. It reads room temperature from a SwitchBot sensor and refuses to start cooling if the room is already below 24°C, so it never overcools an empty office.
It runs as one more launchd job on the same one-minute cadence, reading the same temperature log the dashboard uses. Default-off until I explicitly enabled it, so there was never a chance of it firing by surprise during testing.
Right now, with the room at 26°C and the GPU at 88°C, it correctly sits idle. It's waiting for the day the Tokyo heat pushes the chip past 92, at which point the office will cool itself while I'm somewhere else entirely.
4. One mistake worth copying down
When I first stood the dashboard up, I served the whole working directory. That directory also held the file with my SwitchBot API credentials. For about fifteen minutes, anyone on my private network could have fetched http://m1/.env and read my tokens.
It was tailnet-only, so the blast radius was just my own devices, but the lesson generalizes: a static file server serves everything in the folder you point it at. Put only the public files (the HTML, the graph data) in the served directory, and keep secrets and scripts one level up where the server can't reach them. I moved them, confirmed the credentials now return 404, and rotated nothing because the exposure never left my own machines, but it was a clean reminder.
Closing thoughts
An always-on AI server is a small piece of infrastructure, and infrastructure you can't see is infrastructure you'll eventually lose. Two cheap moves fix that: a one-minute sampler feeding a dashboard you can open by name from anywhere, and a temperature-driven rule that lets the room take care of the machine when you're not there.
The Mac in my Tokyo office now watches its own temperature, tells me about it over my tailnet, and cools its own room when the summer gets serious. I don't have to be in the office to keep it safe.



Top comments (0)