fix(lsphp): stop SIGPIPE+pipefail reporting the parity extension as missing
`entrypoint-lsphp.sh` decided whether cac_path_parity was loaded with
printf '%s\n' "$LSPHP_INFO" | grep -q '^cac_path_parity support => enabled$'
under `set -euo pipefail`. `grep -q` exits on its first match; printf is still
writing the remaining ~40 KB of `lsphp -i`, takes SIGPIPE, exits 141, and
pipefail prefers 141 over grep's 0. The branch therefore evaluated FALSE
*because the extension was present* — present early enough to stop the reader —
and every affected container fell back to the auto_prepend normaliser that a
customer's own .user.ini silently displaces, i.e. the exact failure the
extension exists to remove. Measured on whp02 against the published
cac-lsphp:php83: 5/5 runs status=141 with pipefail, 0 without.
The race is decided by pipe capacity, which is why it reproduced on whp02 and
not on other daemons: while the payload fits the pipe the writer never blocks
and always finishes first. Forced over the limit it is deterministic — 3x the
same `lsphp -i` (122100 bytes) gives 141 every time in the built image.
Fixed by reading with here-strings, which are not pipelines at all, so there is
no second exit status for pipefail to adopt. Same grep/awk patterns; plumbing
only. Same class fixed everywhere it existed under pipefail:
* entrypoint-lsphp.sh parity probe, and the SCAN_DIR awk probe
* entrypoint-litespeed.sh SCAN_DIR probe (a bare assignment: 141 there does
not degrade, `set -e` kills PID 1), and ols_running
* entrypoint-shared-ols.sh ols_running
* render-shared-ols-config.sh site.meta parsing (`sed | head -1`): measured
141 at 6000 duplicate keys, which under `set -e`
aborts the whole render
* fpm-parity-check.sh the `php-fpm -m` pre-flight, whose whole job is to
stop a harness fault being blamed on the extension
Also: the fallback used to announce "cac_path_parity extension not loadable in
this image" for every reason the branch was reached, including its own plumbing
breaking — a false diagnosis that sends operators to rebuild a good image whose
build gate passed. Verdicts now carry the evidence they rest on, and a probe
that produced nothing is reported as a probe failure that establishes nothing
about the image. Fail-open posture is unchanged: no probe failure is fatal.
Adds scripts/tests/lsphp-info-probe.test.sh, which runs the shipped probes
(extracted verbatim, so they cannot drift from what runs in production) under
`set -euo pipefail` against a realistic ~40 KB phpinfo body, and statically
outlaws the shape repo-wide. Against trunk it fails, naming all 9 offending
lines. Wired into CI as a new Shell-Checks job, because no existing gate ever
executed the entrypoint's branch logic — the .phpt suite and the Dockerfile's
own `lsphp -i | grep -q` probe (which has no pipefail) were both green for the
release whose entrypoint declared that same extension missing.
Verified: PHP 8.3 --no-cache build green, 10/10 .phpt, 9/9 FPM harness; the
built image logs `path parity = extension` and reports `Rewriting => active`
with .from/.to populated; ext-removed and probe-broken variants each produce
their own honest message and still start.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -67,7 +67,19 @@ fi
|
||||
## see lsphp's PHP errors in the exact same file on the new image.
|
||||
## Rendered as a tiny ini in lsphp's scan dir; PHP merges it after the
|
||||
## production-tuning overrides at startup.
|
||||
SCAN_DIR=$(/usr/local/lsws/lsphp${PHPVER}/bin/lsphp -i 2>/dev/null | awk -F'=> ' '/^Scan this dir/ {print $2; exit}')
|
||||
## Captured in two steps on purpose. As a single pipeline this was
|
||||
## `lsphp -i | awk '…{print;exit}'`: awk stops at the "Scan this dir" line,
|
||||
## which sits in the first few hundred bytes of ~40 KB of output, so lsphp can
|
||||
## still be writing when awk closes the pipe. It then dies 141, `set -o
|
||||
## pipefail` (line 12) makes that the pipeline's status, and because this is a
|
||||
## bare assignment `set -e` KILLS PID 1 — the container never starts, on a
|
||||
## machine where the race falls the wrong way. (Its twin in entrypoint-lsphp.sh
|
||||
## chose a degraded fallback instead; this one just exits.) Reading into a
|
||||
## variable first leaves awk's own status as the assignment's, and `|| true`
|
||||
## keeps a genuinely failing lsphp as an empty SCAN_DIR — which the `-n` test
|
||||
## below already handles — rather than as a boot failure.
|
||||
LSPHP_INFO=$(/usr/local/lsws/lsphp"${PHPVER}"/bin/lsphp -i 2>/dev/null || true)
|
||||
SCAN_DIR=$(awk -F'=> ' '/^Scan this dir/ {print $2; exit}' <<<"$LSPHP_INFO")
|
||||
if [ -n "$SCAN_DIR" ]; then
|
||||
cat > "$SCAN_DIR/99-user-error-log.ini" <<EOF
|
||||
; rendered at container start by entrypoint-litespeed.sh
|
||||
@@ -194,7 +206,24 @@ trap term_handler TERM INT
|
||||
## down). We match the running message specifically — a bare grep for "running"
|
||||
## would also match "not running". (This image keeps the pidfile under
|
||||
## /tmp/lshttpd, not logs/, so we never hard-code a pidfile path.)
|
||||
ols_running() { /usr/local/lsws/bin/lswsctrl status 2>/dev/null | grep -qi 'running with pid'; }
|
||||
##
|
||||
## Read into a variable and match with a here-string rather than piping into
|
||||
## `grep -qi`: `grep -q` closes the pipe on its first match, and under the
|
||||
## `set -o pipefail` at the top of this file a writer that is still writing when
|
||||
## that happens dies 141 and the pipeline reports FALSE — i.e. "OLS is down"
|
||||
## precisely because the "running" line matched. (Same defect that shipped in
|
||||
## entrypoint-lsphp.sh's cac_path_parity probe.) `lswsctrl status` prints one
|
||||
## short line, so today it wins the race every time; the bound that makes that
|
||||
## true is a vendor script's output, not something this repo controls, and the
|
||||
## failure it would cause here — a spurious relaunch of a healthy OLS, five of
|
||||
## which trip the crash-loop cap and exit PID 1 — is expensive enough not to
|
||||
## rest on it. A non-zero `lswsctrl` still means "not running", exactly as
|
||||
## pipefail made it mean before.
|
||||
ols_running() {
|
||||
local st
|
||||
st=$(/usr/local/lsws/bin/lswsctrl status 2>/dev/null) || return 1
|
||||
grep -qi 'running with pid' <<<"$st"
|
||||
}
|
||||
|
||||
## Crash-loop cap: if OLS can't stay up, bail out so Docker's restart policy and
|
||||
## the site-health monitor escalate instead of us hot-looping forever.
|
||||
|
||||
Reference in New Issue
Block a user