A bell on the quarterdeck
At five past three on a Friday afternoon, the VPSVirtual Private Server — a slice of a machine you rent in a data centre, with its own public IP address. A few dollars a month gets you enough to run a small production app. that runs two of my production apps hit 94.8% disk. A cold backup was staging DockerA tool that packages an app together with everything it needs to run into a "container", so it runs the same on any machine and does not collide with anything else installed. images on it and the images were big. There was about 3 GB left before databases start refusing writes, and nobody knew, least of all me, because I was reading about something else at the time. It dropped back five minutes later, climbed to 87.5% again twenty minutes after that, and settled by itself. I found out the next day, reading graphs.
The lookout was already on deck
The annoying part is that the fleet already had a lookout. Every five minutes a sampler on the main server, run by cronThe Linux stopwatch. It runs a command on a schedule — every night at 3am, every ten minutes — without you being there., asks each host over a read-only SSH keyA pair of files — one secret, one public — that let you log in over SSH without a password. You put the public half on the machine you want to reach; the secret half stays on your laptop. how much memory, CPU and disk it’s using, and writes one line of jsonA plain-text format for structured data: names and values in braces and brackets. Easy for programs to read and write, readable enough for people. per host to a daily file. It had written that spike down faithfully: 71.9%, 81.2%, 90.7%, 94.8%. It just had no bell to ring. The numbers went into a chart, and a chart is something you have to go and look at.
The fashionable answer would be another AI coding agentAn AI you give a task and some tools, and it edits files, runs commands and iterates on its own until the task is done. Claude Code and Codex are the two I use., sitting on a schedule, reading the samples and deciding whether to bother me. I’d rather not pay something to squint at a number that a comparison can read. So the sampler rings the bell itself.
A place for the bell to ring
First there had to be somewhere to ring it. The dashboard grew a Needs you list at the top of the Overview, with a count on the sidebar, backed by a plain notifications.yml in the data folder, YAMLA plain-text format for structured settings, readable by people and machines. Docker Compose files, bosun-x records and most config you will touch are YAML. like everything else in there. Anything that can run a command can raise one: npm run notify -- raise --key disk:web-1 --title .... The key names the thing, not the event, so a job that raises the same key every five minutes updates one row instead of stacking up forty. I can snooze a row until tomorrow or next week, or mark it Done, and Done sticks until the source says something different. A lock directory and a write-to-temp-then-rename keep the dashboard and the command line from trampling each other, which a test proves by throwing twenty writers at it at once.
Teaching it when to ring
The rules are a warning at 85%, urgent at 93%, and a warning when a disk climbs fifteen points inside half an hour, even if it’s nowhere near 85 yet, because that’s what something writing far too much looks like. Then I replayed the real Friday through it before letting it near the real inbox. It warned at 3:00, went urgent at 3:05, cleared at 3:10 when the images were deleted, and raised a brand new warning at 3:30 when the next batch arrived. Correct, and useless. A backup that stages and unstages would have had it ringing and falling silent all afternoon.
So an alert that’s standing now holds until usage has stayed under 82% for thirty minutes straight, and it never steps down while it stands. A critical stays critical until it clears. The live figure still moves, “94.8% used, 3.3 GB free”, but that sits in a separate field that doesn’t count as news, so a row I’ve marked Done doesn’t jump back up every time the number twitches. Stepping up to urgent does bring it back, because that’s news. Replayed again, the same afternoon gives one warning, one critical and one clear, which is what I’d have wanted to see at the time.
The problem is not the problem, the problem is nobody’s looking at it. Now something is, every five minutes, for the price of a comparison and a YAML file. Fair winds.
-x