From 11c02d94ecd8dc4c9276ca033eaa45cd0c2a296b Mon Sep 17 00:00:00 2001 From: jknapp Date: Sat, 22 Aug 2026 19:44:01 -0700 Subject: [PATCH] fix(shared-ols): unmapped Host gets 421, not a 200 that hides a dead site The shared-OLS catch-all (`map _health *`) served html/index.html -- HTTP 200, 11 bytes, "shared-ols" -- to any Host no customer vhost claimed. Three live customer sites (joshuaknapp.net, streamers.channel, blog.anti-social.online) sat in exactly that state for ~2 months on whp01 and no monitor noticed, because every uptime check asks "is it 200?" and it was. A tier-wide catch-all that answers 200 makes a missing vhost indistinguishable from a working site. An unmapped Host now gets 421 Misdirected Request with a short generic body. 421 is semantically exact (the server cannot produce a response for the requested authority) and, unlike 404, cannot be confused with a normal answer from a real site. The discriminator is the request path plus the client address, NOT the Host -- the vhost is selected by the listener map, so by the time these rules run the Host is no longer available to branch on: * `/healthz` from an internal client address (loopback, RFC1918) -> 200 "ok" * everything else, every path, every Host, both listeners -> 421 The 421 for `/` is UNCONDITIONAL: no header, source address or Host talks this vhost into a 200 there, so the property the change exists to guarantee does not rest on anything spoofable. The address gate only hardens /healthz, and X-Forwarded-For cannot be used against it because HAProxy replaces that header with the real client IP. Health probes keep passing unchanged. Both forms were run against a container carrying this change and both exit 0 with "ok": curl -fsSk https://127.0.0.1/healthz (Dockerfile.shared-ols HEALTHCHECK) curl -sfk https://localhost/healthz (WHP setup-shared-ols.sh --health-cmd) `docker inspect` reported healthy with failingStreak=0, on a container with a customer site and on a zero-site container. Measured on the lab VM against OLS 1.8.4 (the production base image): unmapped Host, `/`, :443 and :80 -> 421, 356 bytes, identical for every unmapped Host (no enumeration signal) unmapped Host, any deeper path -> the same 421 configured site, both names, :443/:80 -> 200, served normally litespeed -t -> 0 [ERROR] lines (warnings only, and only about the lab fixture's uid/gid) Two OLS behaviours were measured rather than assumed, and both shaped the implementation -- see the comment block in entrypoint-shared-ols.sh: `context / { type redirect statusCode 421 }` silently degrades to a 302 with an unexpanded Location, and the `errorpage 421` body is fetched as a fresh request through the same rewrite rules (so it needs a %{THE_REQUEST} guard, since %{IS_SUBREQ} and %{ENV:REDIRECT_STATUS} are not populated). The old index.html is removed, not just bypassed: if these rules ever stopped applying, `context /` would fall back to the docRoot index, and with no index.html that is a 403 -- wrong-but-loud, rather than a 200 that is wrong-and-silent. Known consumer to land alongside this: whp-monitoring's probe_shared_ols_catchall() currently detects the catch-all by matching `200` + body `shared-ols`, a signature this change deletes. It must also accept 421, or the detector silently stops detecting. Co-Authored-By: Claude Opus 5 (1M context) --- Dockerfile.shared-ols | 7 ++ scripts/entrypoint-shared-ols.sh | 117 +++++++++++++++++++++++++++- scripts/render-shared-ols-config.sh | 21 ++++- 3 files changed, 139 insertions(+), 6 deletions(-) diff --git a/Dockerfile.shared-ols b/Dockerfile.shared-ols index fea15f6..a9aee5e 100644 --- a/Dockerfile.shared-ols +++ b/Dockerfile.shared-ols @@ -55,6 +55,13 @@ EXPOSE 80 443 ## Health: the entrypoint renders a catch-all _health vhost serving /healthz, so ## this passes from boot (zero customer sites) onward. Self-signed :443. +## +## MUST stay on /healthz, and must stay a LOOPBACK request. That vhost answers +## 421 for every other path/Host so an unmapped customer hostname can never look +## "up" to a monitor; /healthz answers 200 only for an internal client address +## (loopback here). Probing `/` instead would fail the healthcheck and restart +## the whole shared tier. WHP's setup-shared-ols.sh overrides this with the +## equivalent `curl -sfk https://localhost/healthz`; keep the two in step. HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \ CMD curl -fsSk https://127.0.0.1/healthz || exit 1 diff --git a/scripts/entrypoint-shared-ols.sh b/scripts/entrypoint-shared-ols.sh index 4a96ba4..2a8ad7b 100644 --- a/scripts/entrypoint-shared-ols.sh +++ b/scripts/entrypoint-shared-ols.sh @@ -32,18 +32,127 @@ if [ ! -f "$CERT_FILE" ]; then -keyout "$KEY_FILE" -out "$CERT_FILE" -subj "/CN=shared-ols" 2>/dev/null fi -## ---- health vhost (catch-all): valid server with zero customer sites + -## answers HAProxy health checks that hit by IP / unknown Host with a 200 ---- +## ---- health vhost (catch-all) ---- +## This vhost is mapped `map _health *` by render-shared-ols-config.sh, so it +## answers EVERY Host that no customer vhost claims. It exists so the server is +## valid with zero customer sites and so local/edge health probes get a 200. +## +## IT MUST NOT ANSWER 200 FOR AN UNMAPPED CUSTOMER HOST. +## It used to serve html/index.html ("shared-ols", 11 bytes) with HTTP 200 to +## anything that fell through. Measured 2026-08: three live customer sites +## (their vhost had silently stopped being rendered) served that 200 for ~2 +## months and no monitor noticed, because every uptime check asks "is it 200?" +## and the answer was yes. A hostname this server cannot serve now gets +## 421 Misdirected Request -- semantically exact (RFC 7540 s9.1.2: the server is +## not able to produce a response for the combination of scheme and authority in +## the request URI) and unambiguous to monitoring in a way 404 is not, since a +## 404 is a perfectly normal answer from a real, working site. +## +## THE DISCRIMINATOR: request path /healthz AND an INTERNAL client address. +## * Path alone is not enough -- anyone can request /healthz. +## * REMOTE_ADDR is the half an outside caller cannot choose, BECAUSE of +## `useIpInProxyHeader 1` in httpd_config_base.tpl: OLS resolves the client +## IP from X-Forwarded-For, and HAProxy -- the only thing that can reach +## this tier, which has no host-published ports and sits on client-net -- +## SETS (not appends) that header: +## `http-request set-header X-Forwarded-For %[var(txn.real_ip)]` in +## haproxy-manager-base/templates/hap_backend.tpl, which DISCARDS whatever +## the client sent. So a request arriving from outside carries the real +## public client IP. Verified on the lab: `-H 'X-Forwarded-For: 8.8.8.8'` +## on /healthz returns 421. +## * MEASURED LIMIT OF THE IP GATE, stated plainly rather than assumed away: +## OLS takes the FIRST element of a multi-value X-Forwarded-For as +## REMOTE_ADDR. `X-Forwarded-For: 10.0.0.1, 8.8.8.8` returns 200 on /healthz +## here, and anchoring the pattern ^...$ does NOT change that (tested both +## ways) -- because by the time the rule sees REMOTE_ADDR it is already the +## single token `10.0.0.1`. The anchors are kept because they are correct +## and free, not because they close that hole. What closes it is HAProxy: +## `http-request set-header X-Forwarded-For %[var(txn.real_ip)]` REPLACES +## whatever the client sent with one value. +## * AND THE GATE IS NOT LOAD-BEARING ANYWAY. It only guards /healthz. `/`, +## and every other path, is 421 UNCONDITIONALLY -- no header, source +## address or Host can talk this vhost into a 200 there. So even a total +## bypass of the IP gate buys an attacker a 3-byte `ok` on /healthz, never +## a "the site is up" answer on the URL a monitor actually requests. That +## is the property this change exists to guarantee, and it does not rest on +## anything spoofable. +## * The probes that MUST keep passing all originate inside: the Docker +## HEALTHCHECK (`curl -sfk https://127.0.0.1/healthz` in Dockerfile.shared-ols, +## overridden by WHP's setup-shared-ols.sh to `https://localhost/healthz`) +## connects over loopback and sends no X-Forwarded-For, so REMOTE_ADDR falls +## back to the peer, 127.0.0.1. An edge/host probe of the container IP comes +## from the docker gateway (172.16/12), also allowed. +## +## `/` is 421 for EVERY client, internal ones included -- there is deliberately +## no "internal clients still get the old 200 page" escape hatch, because that +## is exactly the response that hid the outage. Anything probing this tier for +## liveness must ask for /healthz. +## +## WHY REWRITE AND NOT A REDIRECT CONTEXT: `context / { type redirect +## statusCode 421 }` was measured on this image (OLS 1.8.4) and does NOT work -- +## 421 is not in OLS's accepted status-code list, so it silently degrades to a +## 302 with a literal, unexpanded `Location: $DOC_ROOT/?`. A rewrite `[R=421,L]` +## does emit a real 421. +## +## WHY THE THE_REQUEST GUARD ON THE ERROR PAGE: a bare [R=421] has no body, and +## a bare 421 with no explanation is a support ticket. `errorpage 421` supplies +## the body, but OLS fetches that URL as a fresh internal request that runs +## through these same rules -- without an exception it is itself 421'd and the +## body comes back empty (measured: content-length 0). %{IS_SUBREQ} and +## %{ENV:REDIRECT_STATUS} are NOT populated by OLS's rewrite engine (both +## measured, both no-ops), but %{THE_REQUEST} keeps the ORIGINAL request line +## across the internal fetch. So: serve misdirected.html when the client did not +## itself ask for it, which lets the error page render while a direct external +## GET /misdirected.html still gets 421 -- no path on this catch-all answers 200 +## to an outside caller. +## +## The body is deliberately generic: no branding, no customer names, nothing +## that reveals which hostnames this server does serve. Every unmapped Host and +## every path gets the byte-identical 421, so the response cannot be used to +## enumerate configured vs unconfigured hostnames. cat > "$HEALTH_DIR/vhconf.conf" <<'EOF' docRoot $VH_ROOT/html enableScript 0 + +errorpage 421 { + url /misdirected.html +} + +rewrite { + enable 1 + rules << "$HEALTH_DIR/html/healthz" -printf 'shared-ols\n' > "$HEALTH_DIR/html/index.html" +cat > "$HEALTH_DIR/html/misdirected.html" <<'EOF' + + +421 Misdirected Request + +

421 Misdirected Request

+

This hostname is not configured on this server.

+

If you own this domain, check that its DNS points to the correct server and +that the site is active in your hosting control panel.

+ + +EOF +## The old catch-all index.html ("shared-ols") is gone on purpose, and actively +## removed so an in-place upgrade of a long-lived container cannot leave it +## behind. If these rewrite rules were ever to stop applying, `context /` would +## fall back to serving the docRoot index -- with no index.html that is a 403, +## which is wrong-but-loud, instead of a 200 that is wrong-and-silent. +rm -f "$HEALTH_DIR/html/index.html" ## ---- ownership: OLS reads conf/ as lsadm. chown the base conf dir + health dir ## NON-recursively (the per-site files under conf/shared-sites are written by the @@ -51,7 +160,7 @@ printf 'shared-ols\n' > "$HEALTH_DIR/html/index.html" ## every container (re)start, delaying first-listen after a crash). The render ## script chowns the httpd_config.conf it produces. ---- chown lsadm:nogroup "$LSWS_CONF" "$HEALTH_DIR" "$HEALTH_DIR/html" 2>/dev/null || true -chown lsadm:nogroup "$HEALTH_DIR/vhconf.conf" "$HEALTH_DIR/html/healthz" "$HEALTH_DIR/html/index.html" 2>/dev/null || true +chown lsadm:nogroup "$HEALTH_DIR/vhconf.conf" "$HEALTH_DIR/html/healthz" "$HEALTH_DIR/html/misdirected.html" 2>/dev/null || true ## ---- assemble httpd_config.conf from the panel's per-site files ---- /scripts/render-shared-ols-config.sh diff --git a/scripts/render-shared-ols-config.sh b/scripts/render-shared-ols-config.sh index 65a2d86..aba490e 100644 --- a/scripts/render-shared-ols-config.sh +++ b/scripts/render-shared-ols-config.sh @@ -156,8 +156,25 @@ for meta in "$SITES_ROOT"/*/site.meta; do done ## --- 5. ALWAYS add a health vhost mapped to the catch-all so the server is -## valid with zero customer sites and HAProxy health checks (which hit by IP / -## unknown Host) get a 200. Exact-domain maps above win over this '*'. --- +## valid with zero customer sites. Exact-domain maps above win over this '*'. +## +## THIS MAP IS WHY AN UNMAPPED HOST GETS AN ANSWER AT ALL. Anything the loop +## above did not emit a `map` for -- a customer domain whose site dir went +## missing, a stale DNS record, a scanner probing by IP -- lands here. It used +## to answer 200 with an 11-byte "shared-ols" body, which is how three live +## customer sites stayed silently broken for ~2 months: every uptime monitor +## asks "is it 200?" and it was. +## +## The health vhost (its vhconf.conf is written by entrypoint-shared-ols.sh, +## which carries the full rationale) now answers 421 Misdirected Request with a +## short generic body for any Host it cannot serve, and keeps 200 ONLY for +## GET /healthz from an internal client address -- the Docker HEALTHCHECK and +## edge liveness probes. Do NOT reintroduce a 200 here for `/`: probe /healthz. +## +## The listener `map` itself is unchanged, deliberately. Dropping the catch-all +## instead would make OLS answer an unmapped Host from whichever vhost it +## considers first, which is worse: an unmapped Host would be served SOMEONE +## ELSE'S SITE. --- { echo "" echo "virtualhost _health {" -- 2.52.0