Publishing lab health without publishing the lab.
Case study · a public status page for a private network
PROBLEM
I run a small home lab: one hypervisor, a rack of self-hosted services, a reverse proxy in front of them, a private network holding it together. I wanted a public page that shows — honestly, from measurement — whether that lab is healthy. A status page for something that is not supposed to be reachable.
The tension is obvious. A status page is a list of what exists. For a private network, the list is the sensitive part: hostnames, ports, the shape of the fleet. The page had to say "the reverse proxy is up" without ever being able to say where the reverse proxy is.
CONSTRAINTS
- Nothing addressable leaves the network. No hostname, IP, port, URL, or internal naming convention — not on the page, not in the feed behind it, not in the public repository.
- No inbound exposure. The lab does not accept connections from the internet, and a status page was not a good enough reason to start.
- Static hosting. The site is hand-written HTML on Cloudflare Pages with no framework and no build step. The status page had to fit that, not replace it.
- Honest when broken. A status page whose backend has died must not keep serving a stale "all green". Silence has to look like silence.
- Maintainable by one person. Adding a service should be one edit, not three that can drift.
OPTIONS & TRADEOFFS
Expose an internal status service. Tunnel or proxy an existing monitoring UI out to the internet. Fastest to build, and exactly the thing the constraints forbid: it opens an inbound path, and every such tool renders the targets it monitors. Rejected.
Fetch from the browser. Let the page call into the lab directly. This does not work — the lab is not reachable — and the attempt would put internal addresses in public JavaScript. Rejected.
Push from inside, read at the edge. A collector on the private network builds a sanitised summary and pushes it out to a small Worker at the Cloudflare edge, which stores it. The page reads whatever landed last. The only network path is outbound from the lab, over HTTPS, to a bearer-token endpoint. Chosen.
Publish endpoints, or publish categories? A conventional page lists services by name and lets you click through. This one publishes categories — "Network & DNS", "Compute & AI" — with curated labels inside them, and nothing is a link. Less useful to a visitor than a real dashboard; that is the price of it being public at all.
Blocklist or allowlist? Filtering out things that look sensitive is a losing game — the monitoring tool can be renamed carelessly at any time. Instead the public label for every service is a constant in an explicit map, never derived from the monitor's real name. A service not in the map is not published. Fail closed.
DESIGN
private network Cloudflare edge browser ─────────────── ─────────────── ─────── collector ──POST /api/ingest──► Worker ──► KV GET /api/status ──► page · reads the monitoring tool · bearer token │ · applies the label map · shape + size check └─ fixture, · scrubs itself · refuses address-like values if KV is empty
The collector is one stdlib-Python file. Every minute it reads the internal monitoring tool, rolls each monitor up through the label map, adds host vitals and a static hardware summary, and pushes — but only when something changed, or on a heartbeat so the timestamp stays honest. It scrubs its own output against the same address-shaped pattern the edge enforces, so a bad value drops one row instead of sinking the whole push.
The edge Worker is under a hundred lines. It accepts a push only with the right bearer token, checks the document's size and shape, and then walks every string in it. If any value looks like an IP, a port, a URL, or an internal hostname suffix, the entire push is refused. The page can therefore not publish an address even if the collector is misconfigured — the collector's scrub is hygiene; the edge's refusal is the guarantee.
The page is static. It fetches the latest push, renders it four ways (rows, terminal, table, cards), and if the feed is older than an hour shows "no contact from collector" rather than the last numbers. If nothing has ever been pushed it serves a shipped fixture, clearly labelled as sample data.
One private source file declares every service — its real target, its category, its public label. The public label map is generated from it, and so is the monitoring configuration. The three things that used to be able to disagree now cannot.
VALIDATION
- A headless smoke test drives each of the four renderers against the fixture and counts what they draw — a syntax check proves nothing about whether a view actually renders.
- A CSS audit checks that every class used in the markup has a rule, after a layout was silently lost to a bad edit early on.
- Asset URLs carry a content hash, because the CDN caches stylesheets for hours while HTML updates instantly — an unstamped deploy shows new markup against stale styles.
- Shared scripts are copied between the two hosts by a script that can also just check for drift.
- The refusal filter is exercised against real payloads, not only the fixture — which is how its worst bug was found.
OUTCOME
Live at lab.monk97.me since 2026-09-27, fed by the collector every minute, publishing twenty services across seven categories and nothing that resolves. The page, the Worker, and the collector are all in the public repository; the file that names the real targets is not, and the repository's history was started fresh so it never was.
Uptime is counted, not estimated: one sample per service per minute, a day is its up-samples over its samples, and the figure on the live page is that same sum across every day since measurement began — captioned with the start date rather than a window it has not yet filled. No number is quoted here because it would be stale by tomorrow; the live one is the measurement.
LESSONS
A backstop that has never fired on real data has never been tested. The first genuine push was refused by the edge: the address filter's port rule, :\d{2,5}, cannot tell a port from a time of day, and the push's own ISO-8601 timestamp tripped it. The fixture never went through the filter, so the bug was invisible for as long as the page served sample data. The fix was not to loosen the rule but to exempt that one field by key and validate it with something stricter. Fail-closed needs a named, reviewed exception path, and it is cheaper to design it on day one.
A status page's number has to be a count. For its first week the page built uptime from the lowest rolling reading seen each day. That is not uptime: it showed 16% for a day with 21 minutes of downtime, and it could just as easily miss an outage that fell between readings. History is now a count of samples, and the caption says how long it has actually been counting.
Inside and outside resolve differently, on purpose. The status page's hostname means something else inside the network. The collector's first pushes therefore never left the LAN — they landed on an internal proxy that had never heard of the ingest route. Pushing to the hosting provider's own hostname, which only ever resolves to the edge, sidesteps a DNS design that is otherwise correct.
Three lists that can drift will drift. The monitoring configuration, the public label map, and the lab's launcher were three hand-maintained lists of the same services. Two are now generated from the third.