Barb is the smallest machine in my fleet — a little cloud droplet named after the woman who owns the trailer park, because her job is to own the place and keep tabs on everything. Monitoring, backups, automation jobs. Barb watches the watchers.
So it was a little awkward to discover that nobody was watching Barb’s watchers.
Six containers, three suspects
While wiring her into a new monitoring setup, I asked a question I apparently hadn’t asked in a while: what’s actually running on this box? The answer was six containers, three of which were a Grafana + InfluxDB + Telegraf stack whose entire job was collecting Barb’s own system metrics and drawing them on dashboards.
Then I checked when anyone had last looked at those dashboards.
Ten months untouched
Grafana’s own database doesn’t lie: three dashboards, last edited October 8th of last year. Admin account’s last login, October 9th. Ten months. In that time, Telegraf had dutifully collected metrics every few seconds, InfluxDB had dutifully stored all of them, and the database had dutifully grown to 756 MB — on the most disk-constrained machine I own, which was sitting at 82% full. A monitoring stack whose only measurable effect on the system was being one of the things you’d want monitoring to warn you about.
Two of the three dashboards were empty drafts. The one real one had seven panels I could not have described to you under oath.
An afternoon of deleting
The teardown took an afternoon. I archived the dashboard definitions and the compose file first — the whole stack can be stood back up verbatim if I ever miss it, which I won’t — then removed all three containers and pruned the leftovers. Disk went from 82% to 78%. Container count went from six to three, all of which earn their keep.
What was actually eating the disk
The good part came while checking what was actually filling the disk, because it wasn’t the monitoring stack. Of 59 GB in use, 48 GB turned out to be backups shipped over from Randy — completely legitimate, actually needed, and the real growth driver nobody had clocked. The flabby containers were a rounding error sitting on top of the actual answer. Classic audit shape: the thing you go in to fix is rarely the biggest thing you find.
A shorter job description
Up/down checks still run, disk usage now has real alerts with real thresholds, and Barb’s job description got shorter and more honest. Deleting things is maintenance. It doesn’t feel like progress because nothing new exists afterward — but “nothing new exists” was the goal. The droplet breathes, the backups have room, and no process on that machine is performing work for an audience of zero.
Ten months of metrics, faithfully collected, never once read. Somewhere in that InfluxDB there was probably a chart showing its own disk usage climbing. It would have been a good chart. Nobody saw it.