This server
2026
What is public is a list, not a side effect
The usual answer is to point Grafana at node_exporter and put a password on it. That publishes whatever the exporter exposes, and an exporter’s job is to expose everything.
Instead the public surface is a checked-in inventory, and the collectors enforce it
in the direction that fails safe. A container appears only if it carries a monitoring.public label. The renderer raises on a name, a status or a
count it does not recognise, rather than passing it through. Adding a number to the
public API is a commit against that inventory.
A collector that raises writes no file at all, the exporter keeps serving the last good one, and the API marks the group stale. Time sources show the pattern most clearly: peer addresses become city labels before the metric is written, so there is no version of the file with an address in it to leak.
Published
- Host counters: CPU, memory, temperature, RAPL power, disk usage and IO
- Per container: status, uptime, CPU and memory against their own limits, for the eight on the list
- Clock offset and stratum, with each upstream peer as a city-level label
- Uptime, and probe latency to the public origin
Never written
- Process lists, command lines, environment, filenames, anything about what the archive holds
- Container IDs, image digests, internal addresses, the name of any container not on the list
- Upstream time-source addresses, turned into labels before the metric exists
- Precise location. The host is published at the middle of the country, flagged approximate
Nothing on the box listens to the internet
There is no port forwarding and no public SSH to find. A tunnel daemon dials out and holds the connection open from the inside.
Two services bind a host port at all, and both bind loopback only. Everything else
talks over a private Docker network. Containers drop all capabilities and run with no-new-privileges; the API also runs as an unprivileged uid on a
read-only root filesystem, with a 16 MB noexec,nosuid tmpfs as the only
writable thing it can see.
Administration goes over a private mesh rather than the public internet. Security patches apply unattended and SMART watches both disks, which is most of what decides whether a box like this is still healthy in a year.
How a request gets in
The router forwards nothing. The tunnel daemon dials out, so the only reachable surface is one hostname on somebody else’s edge.
229 lines instead of cAdvisor
cAdvisor publishes several hundred series per container. This site draws five of them.
Two scripts read docker inspect, docker stats and chronyc, then write five numbers per container and twenty about the
clock. Each writes a temp file and renames it, so the exporter can never read half a
file, and the directory is mounted read-only into the exporter.
The clock sampler has to run on the host, because chronyd binds loopback in the host namespace and its socket directory is root-only. I first ran it as a user unit without lingering enabled, so it died with the login session and the metric stopped without anything going red.
How a number gets out
The two processes meet through a directory the exporter mounts read-only. Nothing pushes, and nothing holds a socket open.
Degradation is a type, not an exception
This page is the client, so its failure modes had to be designed rather than discovered.
Each group in the response carries its own state: live, stale, partial or unavailable. The cache only overwrites itself with a response it trusts, so if Prometheus disappears mid-scrape the API keeps serving the last good snapshot with the affected groups marked, and the dashboard draws dashes in those tiles instead of a spinner that never resolves.
The ceilings sit on the read rather than the write. A metric store answering slowly, or answering with a million points, is the realistic way a read-only API becomes a denial of service against its own host.
The box, while you read about it
Current telemetry from this machine, through the API described above. A dash means the degradation path is doing its job, not that the tile is broken.
Uptime
— Last updated —Temperature
— NO HISTORY YETContainers
CPU & RAM Usage (%)
Storage Overview
Time Offset
— NO HISTORY YETLatency
— NO HISTORY YETWhat it costs to run
Read off Prometheus and node_exporter on 2026-09-02.
The gap is the cpuset rather than luck. Capture and compaction are confined to two cores by the kernel, so the half of the box that serves this page cannot be starved by the half that writes the archive.
Deploys are commit-only
The script refuses to run against a dirty tree and tags the image with the commit SHA, so whatever is running has a name in the history.
A rollback trap is armed before anything changes. Its target defaults to the parent commit and can be named explicitly, which matters the one time you need to skip back over a bad one. The crontab that drives the collectors is captured before it is rewritten and restored with everything else, because a half-rolled-back deploy that leaves new cron lines behind is the failure that would actually have bitten.
One deploy, start to finish
Any failure restores the previous commit’s stack and the crontab it replaced.
Day to day
Containers restart unless stopped. The collector’s check is a dead-man switch that fails if a gating feed goes quiet, so a hung process gets replaced instead of sitting there silent.
Memory, CPU and PID limits on all eight. The archive is pinned by cpuset rather than given a quota, because a quota caps CPU-seconds but still lets the scheduler jitter all four cores.
Security updates apply unattended. Every image is pinned by digest, so nothing moves underneath me until I move it.
chrony against four peers, holding about 100 µs RMS offset. The archive’s timestamps are worth what this is worth, so it is sampled every minute.
The irreplaceable part is small enough to copy nightly and is not being copied nightly. One manual snapshot exists. That is the open gap.
Where it stops
- Nothing alerts: metrics get collected, stored and drawn, and no one is paged. That is how a dead collector ran for six days in August: visible on a dashboard, announced to nobody.
- Prometheus does not watch itself: two scrape targets, neither of them the monitoring stack. If Prometheus degrades, the only thing that notices is a container healthcheck.
- Ten-year retention is a config line: 226 MB after a month says nothing about year three, and nothing downsamples.
- The probe has never gone green: it watches an origin that does not serve the site yet. The check is right and the dashboard says so, which is correct behaviour and still a red light.
- One host: one power supply, one uplink, one pair of disks, no failover. A reboot is an outage, and for the capture a hole in the data.