x‑hakt

A bell on the quarterdeck

control

dockeryaml

At five past three on a Friday afternoon, the that runs two of my production apps hit 94.8% disk. A cold backup was staging images on it and the images were big. There was about 3 GB left before databases start refusing writes, and nobody knew, least of all me, because I was reading about something else at the time. It dropped back five minutes later, climbed to 87.5% again twenty minutes after that, and settled by itself. I found out the next day, reading graphs.

100%70%40%85% warn93% urgent94.8% · 3 GB free87.5%2:453:053:253:454:05
The Friday, one sample every five minutes. The sampler saw all of it and told nobody.

The lookout was already on deck

The annoying part is that the fleet already had a lookout. Every five minutes a sampler on the main server, run by , asks each host over a read-only how much memory, CPU and disk it’s using, and writes one line of per host to a daily file. It had written that spike down faithfully: 71.9%, 81.2%, 90.7%, 94.8%. It just had no bell to ring. The numbers went into a chart, and a chart is something you have to go and look at.

The fashionable answer would be another AI , sitting on a schedule, reading the samples and deciding whether to bother me. I’d rather not pay something to squint at a number that a comparison can read. So the sampler rings the bell itself.

A place for the bell to ring

First there had to be somewhere to ring it. The dashboard grew a Needs you list at the top of the Overview, with a count on the sidebar, backed by a plain notifications.yml in the data folder, like everything else in there. Anything that can run a command can raise one: npm run notify -- raise --key disk:web-1 --title .... The key names the thing, not the event, so a job that raises the same key every five minutes updates one row instead of stacking up forty. I can snooze a row until tomorrow or next week, or mark it Done, and Done sticks until the source says something different. A lock directory and a write-to-temp-then-rename keep the dashboard and the command line from trampling each other, which a test proves by throwing twenty writers at it at once.

samplerevery 5 minjudge85 · 93 · +15/30mnotifications.ymlone row per key: disk:<host>Needs you (1)snooze · Doneraise, update or resolve,never a pile of duplicates
No new job and no lookout to pay: the sampler that was already running rings the bell.

Teaching it when to ring

The rules are a warning at 85%, urgent at 93%, and a warning when a disk climbs fifteen points inside half an hour, even if it’s nowhere near 85 yet, because that’s what something writing far too much looks like. Then I replayed the real Friday through it before letting it near the real inbox. It warned at 3:00, went urgent at 3:05, cleared at 3:10 when the images were deleted, and raised a brand new warning at 3:30 when the next batch arrived. Correct, and useless. A backup that stages and unstages would have had it ringing and falling silent all afternoon.

first trywith a holdwarn 3:00urgent 3:05clear 3:10warn 3:30clear 3:35warn 3:00urgent 3:05held: under 82% for 30 min firstclear 4:05
Same afternoon, twice. The first rules were right every five minutes and useless over an hour.

So an alert that’s standing now holds until usage has stayed under 82% for thirty minutes straight, and it never steps down while it stands. A critical stays critical until it clears. The live figure still moves, “94.8% used, 3.3 GB free”, but that sits in a separate field that doesn’t count as news, so a row I’ve marked Done doesn’t jump back up every time the number twitches. Stepping up to urgent does bring it back, because that’s news. Replayed again, the same afternoon gives one warning, one critical and one clear, which is what I’d have wanted to see at the time.

The problem is not the problem, the problem is nobody’s looking at it. Now something is, every five minutes, for the price of a comparison and a YAML file. Fair winds.

-x