fix(shared-ols): unmapped Host gets 421, not a 200 that hides a dead site #22

Open
jknapp wants to merge 1 commits from fix/shared-ols-catchall-421 into trunk
3 changed files with 139 additions and 6 deletions
+7
View File
@@ -55,6 +55,13 @@ EXPOSE 80 443
## Health: the entrypoint renders a catch-all _health vhost serving /healthz, so
## this passes from boot (zero customer sites) onward. Self-signed :443.
##
## MUST stay on /healthz, and must stay a LOOPBACK request. That vhost answers
## 421 for every other path/Host so an unmapped customer hostname can never look
## "up" to a monitor; /healthz answers 200 only for an internal client address
## (loopback here). Probing `/` instead would fail the healthcheck and restart
## the whole shared tier. WHP's setup-shared-ols.sh overrides this with the
## equivalent `curl -sfk https://localhost/healthz`; keep the two in step.
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
CMD curl -fsSk https://127.0.0.1/healthz || exit 1
+113 -4
View File
@@ -32,18 +32,127 @@ if [ ! -f "$CERT_FILE" ]; then
-keyout "$KEY_FILE" -out "$CERT_FILE" -subj "/CN=shared-ols" 2>/dev/null
fi
## ---- health vhost (catch-all): valid server with zero customer sites +
## answers HAProxy health checks that hit by IP / unknown Host with a 200 ----
## ---- health vhost (catch-all) ----
## This vhost is mapped `map _health *` by render-shared-ols-config.sh, so it
## answers EVERY Host that no customer vhost claims. It exists so the server is
## valid with zero customer sites and so local/edge health probes get a 200.
##
## IT MUST NOT ANSWER 200 FOR AN UNMAPPED CUSTOMER HOST.
## It used to serve html/index.html ("shared-ols", 11 bytes) with HTTP 200 to
## anything that fell through. Measured 2026-08: three live customer sites
## (their vhost had silently stopped being rendered) served that 200 for ~2
## months and no monitor noticed, because every uptime check asks "is it 200?"
## and the answer was yes. A hostname this server cannot serve now gets
## 421 Misdirected Request -- semantically exact (RFC 7540 s9.1.2: the server is
## not able to produce a response for the combination of scheme and authority in
## the request URI) and unambiguous to monitoring in a way 404 is not, since a
## 404 is a perfectly normal answer from a real, working site.
##
## THE DISCRIMINATOR: request path /healthz AND an INTERNAL client address.
## * Path alone is not enough -- anyone can request /healthz.
## * REMOTE_ADDR is the half an outside caller cannot choose, BECAUSE of
## `useIpInProxyHeader 1` in httpd_config_base.tpl: OLS resolves the client
## IP from X-Forwarded-For, and HAProxy -- the only thing that can reach
## this tier, which has no host-published ports and sits on client-net --
## SETS (not appends) that header:
## `http-request set-header X-Forwarded-For %[var(txn.real_ip)]` in
## haproxy-manager-base/templates/hap_backend.tpl, which DISCARDS whatever
## the client sent. So a request arriving from outside carries the real
## public client IP. Verified on the lab: `-H 'X-Forwarded-For: 8.8.8.8'`
## on /healthz returns 421.
## * MEASURED LIMIT OF THE IP GATE, stated plainly rather than assumed away:
## OLS takes the FIRST element of a multi-value X-Forwarded-For as
## REMOTE_ADDR. `X-Forwarded-For: 10.0.0.1, 8.8.8.8` returns 200 on /healthz
## here, and anchoring the pattern ^...$ does NOT change that (tested both
## ways) -- because by the time the rule sees REMOTE_ADDR it is already the
## single token `10.0.0.1`. The anchors are kept because they are correct
## and free, not because they close that hole. What closes it is HAProxy:
## `http-request set-header X-Forwarded-For %[var(txn.real_ip)]` REPLACES
## whatever the client sent with one value.
## * AND THE GATE IS NOT LOAD-BEARING ANYWAY. It only guards /healthz. `/`,
## and every other path, is 421 UNCONDITIONALLY -- no header, source
## address or Host can talk this vhost into a 200 there. So even a total
## bypass of the IP gate buys an attacker a 3-byte `ok` on /healthz, never
## a "the site is up" answer on the URL a monitor actually requests. That
## is the property this change exists to guarantee, and it does not rest on
## anything spoofable.
## * The probes that MUST keep passing all originate inside: the Docker
## HEALTHCHECK (`curl -sfk https://127.0.0.1/healthz` in Dockerfile.shared-ols,
## overridden by WHP's setup-shared-ols.sh to `https://localhost/healthz`)
## connects over loopback and sends no X-Forwarded-For, so REMOTE_ADDR falls
## back to the peer, 127.0.0.1. An edge/host probe of the container IP comes
## from the docker gateway (172.16/12), also allowed.
##
## `/` is 421 for EVERY client, internal ones included -- there is deliberately
## no "internal clients still get the old 200 page" escape hatch, because that
## is exactly the response that hid the outage. Anything probing this tier for
## liveness must ask for /healthz.
##
## WHY REWRITE AND NOT A REDIRECT CONTEXT: `context / { type redirect
## statusCode 421 }` was measured on this image (OLS 1.8.4) and does NOT work --
## 421 is not in OLS's accepted status-code list, so it silently degrades to a
## 302 with a literal, unexpanded `Location: $DOC_ROOT/?`. A rewrite `[R=421,L]`
## does emit a real 421.
##
## WHY THE THE_REQUEST GUARD ON THE ERROR PAGE: a bare [R=421] has no body, and
## a bare 421 with no explanation is a support ticket. `errorpage 421` supplies
## the body, but OLS fetches that URL as a fresh internal request that runs
## through these same rules -- without an exception it is itself 421'd and the
## body comes back empty (measured: content-length 0). %{IS_SUBREQ} and
## %{ENV:REDIRECT_STATUS} are NOT populated by OLS's rewrite engine (both
## measured, both no-ops), but %{THE_REQUEST} keeps the ORIGINAL request line
## across the internal fetch. So: serve misdirected.html when the client did not
## itself ask for it, which lets the error page render while a direct external
## GET /misdirected.html still gets 421 -- no path on this catch-all answers 200
## to an outside caller.
##
## The body is deliberately generic: no branding, no customer names, nothing
## that reveals which hostnames this server does serve. Every unmapped Host and
## every path gets the byte-identical 421, so the response cannot be used to
## enumerate configured vs unconfigured hostnames.
cat > "$HEALTH_DIR/vhconf.conf" <<'EOF'
docRoot $VH_ROOT/html
enableScript 0
errorpage 421 {
url /misdirected.html
}
rewrite {
enable 1
rules <<<END_rules
RewriteCond %{THE_REQUEST} !\s/+misdirected\.html
RewriteRule ^/?misdirected\.html$ - [L]
RewriteCond %{REMOTE_ADDR} ^(127\.0\.0\.1|::1|10\.[0-9.]+|192\.168\.[0-9.]+|172\.(1[6-9]|2[0-9]|3[01])\.[0-9.]+)$
RewriteRule ^/?healthz$ - [L]
RewriteRule .* - [R=421,L]
END_rules
}
context / {
allowBrowse 1
location $DOC_ROOT/
}
EOF
printf 'ok\n' > "$HEALTH_DIR/html/healthz"
printf 'shared-ols\n' > "$HEALTH_DIR/html/index.html"
cat > "$HEALTH_DIR/html/misdirected.html" <<'EOF'
<!DOCTYPE html>
<html lang="en">
<head><meta charset="utf-8"><title>421 Misdirected Request</title></head>
<body>
<h1>421 Misdirected Request</h1>
<p>This hostname is not configured on this server.</p>
<p>If you own this domain, check that its DNS points to the correct server and
that the site is active in your hosting control panel.</p>
</body>
</html>
EOF
## The old catch-all index.html ("shared-ols") is gone on purpose, and actively
## removed so an in-place upgrade of a long-lived container cannot leave it
## behind. If these rewrite rules were ever to stop applying, `context /` would
## fall back to serving the docRoot index -- with no index.html that is a 403,
## which is wrong-but-loud, instead of a 200 that is wrong-and-silent.
rm -f "$HEALTH_DIR/html/index.html"
## ---- ownership: OLS reads conf/ as lsadm. chown the base conf dir + health dir
## NON-recursively (the per-site files under conf/shared-sites are written by the
@@ -51,7 +160,7 @@ printf 'shared-ols\n' > "$HEALTH_DIR/html/index.html"
## every container (re)start, delaying first-listen after a crash). The render
## script chowns the httpd_config.conf it produces. ----
chown lsadm:nogroup "$LSWS_CONF" "$HEALTH_DIR" "$HEALTH_DIR/html" 2>/dev/null || true
chown lsadm:nogroup "$HEALTH_DIR/vhconf.conf" "$HEALTH_DIR/html/healthz" "$HEALTH_DIR/html/index.html" 2>/dev/null || true
chown lsadm:nogroup "$HEALTH_DIR/vhconf.conf" "$HEALTH_DIR/html/healthz" "$HEALTH_DIR/html/misdirected.html" 2>/dev/null || true
## ---- assemble httpd_config.conf from the panel's per-site files ----
/scripts/render-shared-ols-config.sh
+19 -2
View File
@@ -156,8 +156,25 @@ for meta in "$SITES_ROOT"/*/site.meta; do
done
## --- 5. ALWAYS add a health vhost mapped to the catch-all so the server is
## valid with zero customer sites and HAProxy health checks (which hit by IP /
## unknown Host) get a 200. Exact-domain maps above win over this '*'. ---
## valid with zero customer sites. Exact-domain maps above win over this '*'.
##
## THIS MAP IS WHY AN UNMAPPED HOST GETS AN ANSWER AT ALL. Anything the loop
## above did not emit a `map` for -- a customer domain whose site dir went
## missing, a stale DNS record, a scanner probing by IP -- lands here. It used
## to answer 200 with an 11-byte "shared-ols" body, which is how three live
## customer sites stayed silently broken for ~2 months: every uptime monitor
## asks "is it 200?" and it was.
##
## The health vhost (its vhconf.conf is written by entrypoint-shared-ols.sh,
## which carries the full rationale) now answers 421 Misdirected Request with a
## short generic body for any Host it cannot serve, and keeps 200 ONLY for
## GET /healthz from an internal client address -- the Docker HEALTHCHECK and
## edge liveness probes. Do NOT reintroduce a 200 here for `/`: probe /healthz.
##
## The listener `map` itself is unchanged, deliberately. Dropping the catch-all
## instead would make OLS answer an unmapped Host from whichever vhost it
## considers first, which is worse: an unmapped Host would be served SOMEONE
## ELSE'S SITE. ---
{
echo ""
echo "virtualhost _health {"