From cf6936e225b815965ed8efabf4f94ca89f10e7e9 Mon Sep 17 00:00:00 2001 From: jknapp Date: Thu, 13 Aug 2026 21:29:48 -0700 Subject: [PATCH 1/2] fix(ols): scope the htaccess watcher to docroots, stop lswsctrl status log spam MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ols-htaccess-watcher.sh matched .htaccess by basename only, so ANY .htaccess under a tenant (WordPress plugin guard files, not just the docroot OLS reads) triggered a full graceful restart. Measured on whp01 over 24h: 63 restarts, 0 of them from a docroot .htaccess actually changing — all from Wordfence/W3TC/ WPForms/etc. self-healing files, mostly on tenants that aren't even on this tier. Now matches the full path (%w%f) against */public_html/.htaccess, the only .htaccess OLS ever reads, and logs which path triggered each restart. entrypoint-shared-ols.sh's 3s supervisor poll called `lswsctrl status`, which appends a line to lsrestart.log on every invocation. Measured on whp01: 1,819,286 status lines vs 2,429 real restarts in a 96 MB, never-rotated log. ols_running() now checks the process table directly (pgrep -f 'lshttpd - main', verified against the litespeedtech/openlitespeed base image) instead of shelling out to a logging tool on a fixed timer. Co-Authored-By: Claude Opus 5 (1M context) --- scripts/entrypoint-shared-ols.sh | 52 ++++++++++++++++++++++------- scripts/ols-htaccess-watcher.sh | 57 ++++++++++++++++++++++++-------- 2 files changed, 84 insertions(+), 25 deletions(-) diff --git a/scripts/entrypoint-shared-ols.sh b/scripts/entrypoint-shared-ols.sh index d376b05..4a96ba4 100644 --- a/scripts/entrypoint-shared-ols.sh +++ b/scripts/entrypoint-shared-ols.sh @@ -75,19 +75,47 @@ term_handler() { } trap term_handler TERM INT -## Variable + here-string, not a pipe into `grep -qi` — see the long note on the -## identical function in entrypoint-litespeed.sh: `grep -q` closing the pipe on -## a match can leave the writer dying 141, and `set -o pipefail` (line 14) turns -## that into "OLS is down" *because* the running line matched. The reason is -## structural (a pipefail script must not pipe into an early-exit reader), not -## that this particular output is small; and the here-string is safe here for -## the separate reason that `lswsctrl status` is far below the size at which -## bash spills a here-string to a temp file. A non-zero `lswsctrl` still counts -## as not running, as pipefail made it count before. +## NOT `lswsctrl status` (unlike the otherwise-identical function in +## entrypoint-litespeed.sh). `lswsctrl` appends a timestamped line to +## logs/lsrestart.log on EVERY invocation it makes, including `status` — and +## this loop polls every 3s forever. Measured on whp01: lsrestart.log is 96 MB, +## holding 1,819,286 `status` lines against 2,429 real `restart` lines; at one +## poll per 3s that's ~63 days of continuous polling, which is exactly the +## file's age, and it isn't rotated on any host (whp01/whp02/sdbees all growing +## at ~1.5 MB/day). So: check liveness directly instead of shelling out to a +## tool whose logging is a side effect we don't want on a fixed timer. +## +## Verified (docker run litespeedtech/openlitespeed:1.8.4-lsphp83, the exact +## base this image is built FROM — see Dockerfile.shared-ols): the running main +## process shows in `ps` as `openlitespeed (lshttpd - main)`, one PID, always +## present while OLS is up and absent the instant it is killed (checked via +## `ps aux` immediately after `kill -9` on the main PID). `pgrep -f` matches +## against the full command line, and no other process on this image's `ps` +## output contains that string, so this cannot cross-match an unrelated +## process. It also cannot self-match: pgrep excludes its own PID by default, +## and the invoking process here is bash executing this script file, whose own +## argv never contains the pattern text (only the *source lines* of this script +## do, which `pgrep -f` never sees). +## +## Deliberately NOT the pidfile (/tmp/lshttpd/lshttpd.pid, confirmed present in +## the same probe): pidfiles are known to go stale across a crash (verified — +## after `kill -9` the file still held the dead PID), and treating a stale PID +## as "alive" if the kernel ever reuses that number is a false positive this +## supervisor cannot afford (see below). `pgrep -f` reads the live process +## table, so there is no staleness window to reason about. +## +## Conservative on both failure directions, which matters because this is a +## supervisor predicate, not a metric: a false negative makes start_ols() run +## `lswsctrl start` against an already-running OLS — verified against the same +## probe base image, that is NOT a no-op, it sends SIGUSR1 to the live main +## process, i.e. the same graceful self-restart QUIC.cloud IP refreshes trigger +## (see entrypoint-litespeed.sh's note on that handoff) — a brief, zero- +## downtime blip at worst. A false positive is worse: it leaves a genuinely +## dead OLS un-revived until some later poll happens to notice. So if this +## predicate is ever in doubt it should err toward reporting "not running", not +## "running". ols_running() { - local st - st=$(/usr/local/lsws/bin/lswsctrl status 2>/dev/null) || return 1 - grep -qi 'running with pid' <<<"$st" + pgrep -f 'lshttpd - main' >/dev/null 2>&1 } MAX_STARTS=5 diff --git a/scripts/ols-htaccess-watcher.sh b/scripts/ols-htaccess-watcher.sh index 5b11703..a677563 100644 --- a/scripts/ols-htaccess-watcher.sh +++ b/scripts/ols-htaccess-watcher.sh @@ -15,6 +15,20 @@ ## runs it and the panel monitors it (check-ols-htaccess-watcher.php). set -uo pipefail +## WATCH_ROOT is deliberately left as the host-wide /mnt/users, not narrowed to +## the shared-OLS tenant set, even though that set IS derivable in-container +## (render-shared-ols-config.sh's $SITES_ROOT/*/site.meta VHROOT= is exactly +## that list). Narrowing it would mean handing inotifywait a fixed argv list of +## VHROOT dirs at process start — and inotifywait cannot be told to watch a NEW +## directory once running. The panel provisions sites onto this container live, +## between renders; a site added after the watcher started would then sit +## outside every watch until the next container restart, i.e. exactly the +## silent-failure mode (spec 7) this script exists to prevent, now for brand +## new tenants instead of none. Doing this safely needs a reload path (SIGHUP +## re-exec off the current site.meta list, coordinated with +## render-shared-ols-config.sh) that does not exist yet and is its own change. +## So: WATCH_ROOT stays broad, and correctness comes entirely from the path +## match below, which is sufficient on its own. WATCH_ROOT="${OLS_WATCH_ROOT:-/mnt/users}" DEBOUNCE="${OLS_HTACCESS_DEBOUNCE:-15}" # coalesce window (s) FLOOR="${OLS_HTACCESS_FLOOR:-60}" # min seconds between restarts @@ -24,16 +38,17 @@ last_restart=0 log() { echo "ols-htaccess-watcher: $*" >&2; } do_restart() { + path="$1" now=$(date +%s) if [ $((now - last_restart)) -lt "$FLOOR" ]; then - log "within ${FLOOR}s floor — coalescing, skipping restart" + log "within ${FLOOR}s floor — coalescing, skipping restart ($path)" return fi if "$LSWSCTRL" restart >/dev/null 2>&1; then last_restart=$now - log "graceful restart issued (.htaccess change)" + log "graceful restart issued — $path changed" else - log "WARNING: lswsctrl restart failed" + log "WARNING: lswsctrl restart failed ($path)" fi } @@ -41,18 +56,34 @@ if ! command -v inotifywait >/dev/null 2>&1; then log "FATAL: inotifywait not installed (inotify-tools)"; exit 1 fi mkdir -p "$WATCH_ROOT" -log "watching $WATCH_ROOT for .htaccess changes (debounce=${DEBOUNCE}s floor=${FLOOR}s)" +log "watching $WATCH_ROOT for docroot (public_html) .htaccess changes (debounce=${DEBOUNCE}s floor=${FLOOR}s)" -## -m monitor, -r recursive. We filter to .htaccess in the read loop rather than -## --include so this works on older inotify-tools too. modify/create/delete/move -## all matter (delete of .htaccess also changes rewrite behavior). -inotifywait -m -r -e modify,create,delete,move "$WATCH_ROOT" --format '%f' 2>/dev/null | -while read -r fname; do - case "$fname" in - .htaccess) ;; +## -m monitor, -r recursive. We filter in the read loop rather than --include +## so this works on older inotify-tools too. modify/create/delete/move all +## matter (delete of .htaccess also changes rewrite behavior). +## +## --format '%w%f' (full path), NOT '%f' (basename only). OLS reads .htaccess +## (RewriteFile) only from a vhost's DOCROOT — VHROOT, i.e. +## /mnt/users///public_html (see render-shared-ols-config.sh / +## entrypoint-lsphp.sh) — never anything below it. A basename-only match fires +## for ANY .htaccess anywhere under a tenant, at any depth, and WordPress +## plugins write plenty of those that OLS never opens: measured on whp01 over +## 24h, this watcher fired 63 restarts, of which the docroot .htaccess actually +## changed in 0. All 28 distinct files behind those 63 were plugin guard files +## — Wordfence self-healing waf/views/vendor/tmp/models/lib/.htaccess, W3 Total +## Cache writing one per cached URL under wp-content/cache/page_enhanced/, plus +## WPForms/Gravity Forms/UpdraftPlus/Groundhogg/WP Staging upload guards — and +## most of those tenants are on the shared Apache tier (cac-fpm), not this OLS +## tier at all, so their cache churn was restarting the OLS serving 15 unrelated +## tenants for no reason. Matching the full path down to /public_html/.htaccess +## is what actually ties a change to something OLS will reread. +inotifywait -m -r -e modify,create,delete,move "$WATCH_ROOT" --format '%w%f' 2>/dev/null | +while read -r path; do + case "$path" in + */public_html/.htaccess) ;; *) continue ;; esac - ## A tenant .htaccess changed. Coalesce the save-burst, then restart ONCE. + ## A tenant DOCROOT .htaccess changed. Coalesce the save-burst, then restart ONCE. ## ## The coalesce is HARD-BOUNDED to DEBOUNCE seconds: a previous version blocked ## on `read -t DEBOUNCE` which, on a busy multi-tenant server, never timed out @@ -69,5 +100,5 @@ while read -r fname; do break # ~2s of total quiet — the burst has settled fi done - do_restart + do_restart "$path" done From 77af001af640e066e80508d2749e1f290161c27c Mon Sep 17 00:00:00 2001 From: jknapp Date: Thu, 13 Aug 2026 21:39:14 -0700 Subject: [PATCH 2/2] fix(ols): pin procps explicitly for pgrep dependency MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit entrypoint-shared-ols.sh's ols_running() liveness check now shells out to pgrep, but procps was never in Dockerfile.shared-ols's apt-get install list — pgrep works today only because Ubuntu 24.04's base image pulls procps in transitively. If that stops being true, pgrep: command not found -> exit 127 -> ols_running() false forever -> the crash-loop breaker (MAX_STARTS/WINDOW) escalates to a hard exit 1 at boot. Make the dependency explicit so it can't be pruned as unused. --- Dockerfile.shared-ols | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/Dockerfile.shared-ols b/Dockerfile.shared-ols index 074497d..fea15f6 100644 --- a/Dockerfile.shared-ols +++ b/Dockerfile.shared-ols @@ -24,9 +24,13 @@ FROM litespeedtech/openlitespeed:${OLS_VERSION}-lsphp${PHPVER} ## - gettext-base: envsubst for render-shared-ols-config.sh ## - openssl: self-signed cert for the :443 listener (HAProxy verifies none) ## - curl/ca-certificates: HEALTHCHECK +## - procps: provides pgrep, which entrypoint-shared-ols.sh's ols_running() +## liveness check depends on. Only transitively present via the base image +## today (Ubuntu 24.04 pulls it in) — pin it explicitly so it can't be +## pruned as "unused" and silently break the supervisor's crash detection. RUN apt-get update && \ DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \ - inotify-tools gettext-base openssl ca-certificates curl && \ + inotify-tools gettext-base openssl ca-certificates curl procps && \ apt-get clean && \ rm -rf /var/lib/apt/lists/* /var/cache/apt/archives/*