All projects

This server

2026

boundary

What is public is a list, not a side effect

The usual answer is to point Grafana at node_exporter and put a password on it. That publishes whatever the exporter exposes, and an exporter’s job is to expose everything.

Instead the public surface is a checked-in inventory, and the collectors enforce it in the direction that fails safe. A container appears only if it carries a monitoring.public label. The renderer raises on a name, a status or a count it does not recognise, rather than passing it through. Adding a number to the public API is a commit against that inventory.

A collector that raises writes no file at all, the exporter keeps serving the last good one, and the API marks the group stale. Time sources show the pattern most clearly: peer addresses become city labels before the metric is written, so there is no version of the file with an address in it to leak.

Published

  • Host counters: CPU, memory, temperature, RAPL power, disk usage and IO
  • Per container: status, uptime, CPU and memory against their own limits, for the eight on the list
  • Clock offset and stratum, with each upstream peer as a city-level label
  • Uptime, and probe latency to the public origin

Never written

  • Process lists, command lines, environment, filenames, anything about what the archive holds
  • Container IDs, image digests, internal addresses, the name of any container not on the list
  • Upstream time-source addresses, turned into labels before the metric exists
  • Precise location. The host is published at the middle of the country, flagged approximate
ingress

Nothing on the box listens to the internet

There is no port forwarding and no public SSH to find. A tunnel daemon dials out and holds the connection open from the inside.

Two services bind a host port at all, and both bind loopback only. Everything else talks over a private Docker network. Containers drop all capabilities and run with no-new-privileges; the API also runs as an unprivileged uid on a read-only root filesystem, with a 16 MB noexec,nosuid tmpfs as the only writable thing it can see.

Administration goes over a private mesh rather than the public internet. Security patches apply unattended and SMART watches both disks, which is most of what decides whether a box like this is still healthy in a year.

How a request gets in

Internet
Cloudflare Edge + tunnel
Reverse proxy Caddy
Site + API Svelte, FastAPI

The router forwards nothing. The tunnel daemon dials out, so the only reachable surface is one hostname on somebody else’s edge.

collect

229 lines instead of cAdvisor

cAdvisor publishes several hundred series per container. This site draws five of them.

Two scripts read docker inspect, docker stats and chronyc, then write five numbers per container and twenty about the clock. Each writes a temp file and renames it, so the exporter can never read half a file, and the directory is mounted read-only into the exporter.

The clock sampler has to run on the host, because chronyd binds loopback in the host namespace and its socket directory is root-only. I first ran it as a user unit without lingering enabled, so it died with the login session and the metric stopped without anything going red.

How a number gets out

Host + Docker chronyc, inspect
Collectors cron, 60 s
Textfile atomic rename
node_exporter read-only mount
Prometheus 15 s scrape
Status API cached, capped

The two processes meet through a directory the exporter mounts read-only. Nothing pushes, and nothing holds a socket open.

serve

Degradation is a type, not an exception

This page is the client, so its failure modes had to be designed rather than discovered.

Each group in the response carries its own state: live, stale, partial or unavailable. The cache only overwrites itself with a response it trusts, so if Prometheus disappears mid-scrape the API keeps serving the last good snapshot with the affected groups marked, and the dashboard draws dashes in those tiles instead of a spinner that never resolves.

The ceilings sit on the read rather than the write. A metric store answering slowly, or answering with a million points, is the realistic way a read-only API becomes a denial of service against its own host.

Routes 3 health, system, history. Anything else is a 404, anything but GET a 405, an unexpected query parameter a 422
Cache 15 / 300 s snapshot and history. One Prometheus round trip per window, however many people are reading
Freshness 60 / 180 s per source. Past its threshold a group is marked stale rather than dropped or guessed at
Ceilings 64 / 128k series and samples, with 4 MB on the upstream read and 8 MB on the response. Over the cap is a 503
live

The box, while you read about it

Current telemetry from this machine, through the API described above. A dash means the degradation path is doing its job, not that the tile is broken.

Live from the host

Uptime

— Last updated —

Temperature

— NO HISTORY YET

Containers

no container data

CPU & RAM Usage (%)

no history yet

Storage Overview

NVMe — / — GB
—
SSD — / — GB
—

Time Offset

— NO HISTORY YET

Latency

— NO HISTORY YET
measured

What it costs to run

Read off Prometheus and node_exporter on 2026-09-02.

7-day mean busy, per core
core 0 27.9% archive + compaction, pinned
core 1 31.7% archive + compaction, pinned
core 2 3.3% everything else
core 3 3.3% everything else

The gap is the cpuset rather than luck. Capture and compaction are confined to two cores by the kernel, so the half of the box that serves this page cannot be starved by the half that writes the archive.

Processor draw 3.0 / 15.0 W p50 and peak over 30 days, package plus DRAM by RAPL
CPU temperature 42 / 45 °C now and p99 over 30 days, on passive airflow in a cupboard
Written per day 18.7 / 2.7 GB NVMe and SSD, 7-day means. The NVMe takes the write-ahead tape so the SSD does not
Network moved 603 GiB in August. About 3 Mbit/s average, or 0.35% of the link
Metrics on disk 226 MB Prometheus after a month at 15-second scrapes, across two targets
Public probe 0 in 7 d the blackbox check has not seen a 2xx, because the site is not served at that origin yet
shipping

Deploys are commit-only

The script refuses to run against a dirty tree and tags the image with the commit SHA, so whatever is running has a name in the history.

A rollback trap is armed before anything changes. Its target defaults to the parent commit and can be named explicitly, which matters the one time you need to skip back over a bad one. The crontab that drives the collectors is captured before it is rewritten and restored with everything else, because a half-rolled-back deploy that leaves new cron lines behind is the failure that would actually have bitten.

One deploy, start to finish

Clean tree or it refuses
Build tagged @commit
Compose up config first
Health poll 30 s to answer
Cron rewritten collectors last

Any failure restores the previous commit’s stack and the crontab it replaced.

operate

Day to day

Health

Containers restart unless stopped. The collector’s check is a dead-man switch that fails if a gating feed goes quiet, so a hung process gets replaced instead of sitting there silent.

Isolation

Memory, CPU and PID limits on all eight. The archive is pinned by cpuset rather than given a quota, because a quota caps CPU-seconds but still lets the scheduler jitter all four cores.

Updates

Security updates apply unattended. Every image is pinned by digest, so nothing moves underneath me until I move it.

Time

chrony against four peers, holding about 100 µs RMS offset. The archive’s timestamps are worth what this is worth, so it is sampled every minute.

Backups

The irreplaceable part is small enough to copy nightly and is not being copied nightly. One manual snapshot exists. That is the open gap.

limits

Where it stops

  • Nothing alerts: metrics get collected, stored and drawn, and no one is paged. That is how a dead collector ran for six days in August: visible on a dashboard, announced to nobody.
  • Prometheus does not watch itself: two scrape targets, neither of them the monitoring stack. If Prometheus degrades, the only thing that notices is a container healthcheck.
  • Ten-year retention is a config line: 226 MB after a month says nothing about year three, and nothing downsamples.
  • The probe has never gone green: it watches an origin that does not serve the site yet. The check is right and the dashboard says so, which is correct behaviour and still a red light.
  • One host: one power supply, one uplink, one pair of disks, no failover. A reboot is an outage, and for the capture a hole in the data.