Files
cloud-apache-container/docs/2026-08-05-boot-posture-and-ci-gating-decision.md
T
shadowdaoandClaude Opus 5 584099ff19 docs: record boot-posture and CI-gating decisions
Two judgement calls arising from the pipefail/SIGPIPE fix, reasoned through
and written up rather than decided silently:

1. Boot posture when 'lsphp -i' genuinely fails. Recommendation is fail-open
   plus a machine-detectable degraded marker and one probe retry, NOT
   fail-closed: a transient fork() failure during this fleet's backup windows
   produces the identical empty-LSPHP_INFO signature, so fail-closed would
   refuse to boot healthy sites under host pressure. Includes the caution
   that degraded must not map onto Docker 'unhealthy', or a watchdog would
   restart-loop a serving site over something restarts cannot fix.

   Corrects a premise the fix rested on: the two entrypoints were NOT already
   consistent — entrypoint-litespeed.sh fails open silently, with no warning.

2. Whether the new static check should gate releases. Recommendation is to
   hard-gate, but only after back-testing the job against trunk and recent
   tags to bound false positives.

Recommendations only; no implementation.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-05 15:43:18 -07:00

20 KiB
Raw Blame History

Decision record: entrypoint boot posture on a failed lsphp -i probe, and CI gating for the pipeline-shape check

Date: 2026-08-05 Status: recommendations for sign-off; nothing here is implemented by this document. Background: the SIGPIPE-under-pipefail fix (here-strings replacing printf | awk in scripts/entrypoint-lsphp.sh and scripts/entrypoint-litespeed.sh) removed a boot-killing 141 and, as a side effect, changed what happens when lsphp -i genuinely produces nothing. Two judgement calls were flagged for a human decision. This document is that decision's working-out.


Call 1 — boot posture when lsphp -i genuinely fails

What actually happens today (post-fix, as-written in the repo)

The two container types are not currently symmetric, which matters for the decision:

  • cac-lsphp (scripts/entrypoint-lsphp.sh): fail-open, but not silent. An empty SCAN_DIR falls to an else-branch that distinguishes "lsphp answered but reported no scan dir" from "lsphp -i produced no usable output at all" (lsphp_info_is_usable, keyed on the PHP Version => banner), prints a distinct WARNING: to stderr for each, and every boot ends with a summary line entrypoint-lsphp: $_SERVER path parity = <MODE> (...) where <MODE> becomes none (no scan dir) or none (lsphp -i unusable). So docker logs carries a greppable verdict — but nothing on the platform consumes it.
  • cac-litespeed (scripts/entrypoint-litespeed.sh, lines 81102): fail-open and completely silent. if [ -n "$SCAN_DIR" ] ... fi with no else. An empty probe means the per-site error_log ini and the opcache override ini are simply never written, and nothing says so anywhere.
  • Shared tiers (entrypoint-shared-ols.sh, entrypoint-shared-httpd.sh): no lsphp -i probe exists in them at all. The question is moot there today, but the blast-radius argument below still applies if one is ever added.

What degradation costs the customer, concretely:

Lost fragment cac-lsphp cac-litespeed Customer-visible symptom
99-user-error-log.ini yes yes PHP errors go to stderr/docker logs instead of /home/$user/logs/php-fpm/error.log. Baked log_errors = On still applies, so errors are not lost — they are somewhere the customer can't see and support doesn't look first. Ticket shape: "my error log is empty."
99-user-opcache.ini yes yes Panel-set opcache overrides silently revert to baked defaults (64 MB / 8000 files). A site tuned up for a big codebase gets cache thrash; hard to attribute.
99-cac-path-parity.ini yes n/a Extension gets no mapping ⇒ $_SERVER path parity with cac-fpm is lost entirely in the "no output" case (the auto_prepend fallback is only written in the "extension missing from module list" branch, which requires the probe to have answered). Symptom: Wordfence/path-sensitive plugins misbehave after migration; extremely hard to diagnose from the outside.

The premise to examine first

The fail-open argument assumes "the site serves either way; only ini fragments differ." That is true when lsphp -i answers but lacks a Scan this dir line (weird packaging, still a working interpreter). It is much weaker when lsphp -i produces nothing: the same $LSPHP_BIN that just failed to print phpinfo is what the entrypoint execs as PID 1 to serve requests. A binary that can't run -i because of a missing shared library will not serve LSAPI either — fail-open in that case doesn't buy a serving container, it buys a container that passes docker ps, then wedges or crash-loops after the entrypoint's diagnostic moment has passed. So the two sub-cases deserve different treatment, and conflating them is the main framing error in a flat "fail-open vs fail-closed" question.

But there is a counterweight, and it is specific to this fleet: lsphp -i can fail transiently on a healthy image. These hosts have documented fork-starvation episodes (Committed_AS pinned against CommitLimit during backup windows; it has broken backups before). A fork()/execve failure at boot time yields exactly "empty LSPHP_INFO" through the || true. Under fail-closed, a container (re)starting during a memory-pressure window would refuse to boot because of the host's state, not its own, and its restart policy would then add restart churn to a host already under pressure. The previous incident has the same shape: the last time "the probe returned nothing useful" happened in production, the image was fine — the probe itself was broken (the SIGPIPE bug), and the fail-closed behavior converted a cosmetic probe defect into a fleet-wide fails-to-boot on busy hosts only. Historically, probe-side failures have outnumbered genuinely-corrupt-image failures 10 (or 20 counting fork starvation as observed class). That base rate is the strongest single argument against pure fail-closed.

Options

A. Fail-open as fixed (status quo).

  • Buys: no false boot failures from probe-side defects or host memory pressure; matches the extension's documented posture; sibling entrypoints consistent (once litespeed gains the warning it currently lacks).
  • Costs: "starts but serves degraded" is precisely the state this platform is worst at noticing — the container watchdog's filter has historically been exited|dead|created only, blind to running-but-wrong, and its post-start verify checks State.Running, so it would even score a restart of a degraded container as success. Nothing alerts on the stderr WARNINGs. Realistic outcome of a genuine failure under Option A: weeks of silent degradation, surfaced eventually as a confusing customer ticket ("where are my errors?" / plugin misbehavior), with the diagnostic clue buried in docker logs of a container nobody had reason to inspect.

B. Fail-closed (exit 1) whenever lsphp -i yields nothing usable.

  • Buys: loud, immediate, operator-visible failure; watchdog catches exited; the diagnostic (the WARNING line) is the last thing in docker logs, exactly where an on-call looks first. A genuinely broken image cannot masquerade as healthy for weeks.
  • Costs: converts transient host-side failures (fork starvation) and any future probe-side defect into customer downtime. On a bad-image push, every container that restarts goes down hard — arguably correct (they weren't going to serve anyway), but it also means a probe regression like the one just fixed would again be a mass boot outage rather than a mass warning. The customer experience of B's false positive is strictly worse than A's true positive: total outage vs. degraded-but-serving.
  • Note: B is not "the old behavior restored" — the old behavior also failed to boot on healthy images because the probe itself was the broken part. Anyone arguing "the old crash-loop caught real problems" should first account for the fact that the only production firing of that crash was a false positive.

C. Fail-open + a machine-detectable degraded marker (middle). Boot and serve, but make the degraded state a first-class, monitorable fact rather than a log line:

  1. On any degraded branch, write a marker file (e.g. /run/cac-degraded containing the mode string) in addition to the existing stderr WARNING, in both entrypoints — and add the currently missing warning + summary line to entrypoint-litespeed.sh so the two siblings really are consistent (today's "consistency" claim is only half true).
  2. Retry the probe once or twice with a short sleep before accepting degradation. This is nearly free and specifically absorbs the fork-starvation transient, which is the most likely real-world cause of an empty probe on this fleet.
  3. Surface the marker: either a periodic platform check (docker exec ... test -f /run/cac-degraded across cac containers, or grepping the boot summary line path parity = none from docker logs --tail) wired into existing monitoring/alerting, or fold it into the health story. Caution on the healthcheck route: the HEALTHCHECKs (all seven Dockerfiles have one) are liveness probes; making "degraded config" report unhealthy is tempting but risky — if the watchdog ever learns to restart unhealthy containers, a persistently-degraded-but-serving site would be restart-looped for a condition a restart cannot fix. Keep degraded ≠ unhealthy; alert on the marker through a channel that pages a human instead of triggering an automated restart.

Recommendation

Option C, with one fail-closed carve-out considered and (narrowly) rejected.

The carve-out I weighed: fail closed only in the lsphp -i is unusable sub-case (probe empty), fail open in the answered but no scan dir sub-case — on the logic that an interpreter that can't print phpinfo probably can't serve. I recommend against it, for two reasons. First, the fork-starvation transient lands in exactly that sub-case and would cause real, host-pressure-correlated boot failures of healthy sites; the retry in C(2) mitigates but can't eliminate it. Second, if the binary truly cannot execute, the final exec "$LSPHP_BIN" -b ... fails on its own a few lines later and PID 1 dies anyway — fail-closed-at-the-probe mostly duplicates a crash the container will have regardless, while adding a new false-positive class. The genuinely dangerous residue — binary runs but is subtly broken — is better caught by the existing HEALTHCHECK (TCP :9000 + pgrep lsphp), which does fail in that world; the platform gap is that the watchdog ignores unhealthy, and that gap is worth closing on its own merits.

Per-type answer: same posture for cac-lsphp and cac-litespeed (their failure domains are equivalent for this probe; the current behavioral gap between them is an oversight, not a design). For shared-tier entrypoints the answer is more strongly fail-open than for per-site containers, should a similar probe ever appear there: a fail-closed shared tier takes down every site on the host, and the marginal diagnostic value of a crash over a marker does not scale with the blast radius.

Why C over A, stated so it can be disagreed with: A and C serve the customer identically; they differ only in whether the platform can notice. The honest failure-likelihood ranking on this fleet is (1) transient probe failure under memory pressure, (2) probe logic regression, (3) genuinely broken image — and only (3) is the case where B outperforms. C handles (1) with the retry, (2) and (3) with the marker + alert, and its worst case is "degraded for as long as the alerting channel takes to be read" instead of A's "degraded for weeks" or B's "down because the host was busy." If you believe a corrupt image push is more likely than a probe/host transient, B becomes defensible — but the recent incident history says otherwise.

What a customer experiences, per option, when the failure is genuine: A/C — site up, error log empty, opcache defaults, possible plugin path weirdness; C additionally gets it noticed and fixed in hours. B — site down until an operator acts, but the operator is paged immediately by existing machinery. Which failure gets silently tolerated for weeks: A's, unambiguously. B's cannot be (crash-loops are loud); C's only if the marker alert is never actually wired up — which is why C is not done until the alerting side ships, and shipping the marker without the consumer would be A with extra steps.

Before committing, measure:

  • Fleet scan: does path parity = none (or an empty scan dir) occur on any live container today? (docker logs grep across cac-lsphp/litespeed containers.) If it already fires, the fix priority changes from "posture" to "why".
  • Frequency of boot-time fork failures during backup windows (correlate container start timestamps with MemAvailable dips) — validates or kills the retry rationale.
  • Whether the watchdog's unhealthy-blindness fix (in progress platform-side) will ever auto-restart on unhealthy; that decides whether degraded may ever map to unhealthy.

Call 2 — should the new CI check gate releases?

What actually exists (.gitea/workflows/build-push.yaml)

A Shell-Checks job runs on every push to trunk: (1) bash -n over every script, (2) shellcheck -S warning over an explicit allow-list of the pipefail-era scripts, (3) scripts/tests/lsphp-info-probe.test.sh, whose section 5 is the structural scan that outlaws early-exit-reader pipelines in any pipefail-enabled script under scripts/ and ext/. None of the six image build jobs has needs: Shell-Checks, and the workflow's own comment says so: "a red job here does not block a push that is otherwise fine." So today a red check produces a red run annotation while all seven images publish anyway.

False-positive surface, honestly assessed:

  • bash -n and the test's sections 14 are deterministic (the pre-fix pipeline is run only informationally; no pass/fail depends on the race). Near-zero FP risk.
  • shellcheck runs against a hand-curated clean set, so it goes red only when a listed file regresses or a new shellcheck version tightens — low FP risk, nonzero.
  • The structural scan is a line-oriented awk text scan with flat quote-stripping. Its own header is candid: it can miss multi-line pipelines (false negatives) and could misread exotic quoting (false positives). It requires the reader to be the first word after | and special-cases case a|b), sed "s|x|y|", ||. FP risk is real but bounded — and an FP is a visible red on a specific line, trivially triaged.

Options

A. Leave it non-blocking (status quo). Buys: zero risk that a scanner quirk halts a release. Costs: this platform's own history says non-blocking signals decay into wallpaper — one-shot installers drifted, the watchdog's blind spot sat unnoticed, and this exact bug class shipped to production through a green build gate and two code reviews. A check that exists precisely because the build gate can't see entrypoint runtime, but that cannot stop a release, protects nothing on the day it matters; it merely timestamps the moment the regression was ignorable. The failure it would have caught kills PID 1 on busy production hosts only — the single most expensive place to discover it.

B. Hard gate now (add needs: Shell-Checks to all six build jobs). Buys: the regression class physically cannot ship again through this pipeline. Costs: an unproven scanner acquires veto power over every release, including emergency ones. Worth sizing that cost concretely rather than abstractly: releases here are operator-driven pushes to trunk, not a high-frequency train; a false red costs the minutes needed to read the offending line, either fix real sloppiness or adjust the scanner, and re-push. There is no mechanism by which a red silently delays an unattended release for days, because releases are attended. The scary version of "blocks all releases" mostly doesn't apply to this repo's release model. The one genuinely bad scenario: an urgent security rebuild blocked at 2 a.m. by a scanner FP — mitigable, since the operator can comment out one needs: line in the same push, an escape hatch that is itself visible in the diff.

C. Staged: soft-fail now, hard gate after N clean runs. Buys: evidence before enforcement; if the scanner has an FP mode, it is discovered while red is cheap. Costs: N pushes of window during which the bug class can ship again, and a follow-up task of exactly the kind this platform historically forgets (see: one-shot installer drift). If staged, the flip must be a dated calendar/scheduler entry or a tracked ticket, not an intention.

D. Split the gate: hard-gate the deterministic steps now, soft-fail the scanner. Put bash -n + the probe regression test (sections 14) in a blocking job immediately — they have effectively no FP surface and section 2's marker check alone prevents the specific regression (reintroducing the inline pipelines) — and leave shellcheck + the structural scan non-blocking until proven. Buys: immediate protection against the exact bug that happened, zero new FP-block risk. Costs: two jobs to maintain; the structural scan (the only piece that generalizes beyond this one probe) stays advisory.

Recommendation

B — make it a hard gate now, with C's discipline applied only to future additions to the check, and one prerequisite: run the full Shell-Checks job as-is against the current trunk and the last ~5 release tags first (locally or via a manual dispatch) to confirm it is green on everything already shipped. If that back-test is green, gate. If it is not, fix or scope the scanner, then gate.

Reasoning, laid out to be disagreeable-with: the decision hinges on comparing the two error costs as they occur in this repo's actual release process. A false red costs an attended operator minutes, with a one-line escape hatch, and every FP found makes the scanner better. A false green — or a true red that doesn't block — reproduces a production incident whose failure mode is "customer containers refuse to boot, but only on hosts busy enough to have collapsed pipe buffers," i.e. undetectable in dev by construction. The asymmetry is roughly minutes-of-operator-time vs. production-boot-outage, and the probability of FP red is bounded by a back-test we can run today instead of estimating. Choosing C is defensible only if you weight "new check might be flaky in ways a back-test can't reveal" highly — plausible for a check with runtime/timing behavior, but sections 14 are deterministic and section 5 is a pure text scan over files that are fully known at back-test time. Flakiness that a back-test can't surface would have to come from the runner environment (shellcheck version drift is the main one — worth pinning the shellcheck version when gating). D is the fallback if the back-test turns up scanner noise that can't be resolved same-day.

Two obligations come with gating, or the gate rots:

  1. The shellcheck allow-list must grow with the repo — a new set -o pipefail script that never gets listed is invisibly exempt from step (2) (the structural scan, by contrast, auto-discovers pipefail scripts). A one-line CI assertion that every pipefail-enabled script appears in the shellcheck list would close that seam.
  2. Green must not be oversold: the scanner's documented limits (single-line pipelines, flat quote-stripping) mean a multi-line pipeline reintroduction passes it. The gate raises the floor; it does not replace review. Record that limit where reviewers see it (it already is, in the test header — keep it there).

Before committing, measure: the back-test above (mandatory); shellcheck version on the Gitea runner vs. pinned (drift is the likeliest future FP source); and after gating, track red-run causes for the first month — if more than ~1 in 10 reds is a scanner FP, revisit with D.


Framing notes (where the question as posed deserves push-back)

  1. "Previously it failed loudly and the watchdog caught it" is only half the history: the only production firing of the loud failure was a false positive (the probe's own SIGPIPE), not a broken image. The old posture's loudness was never demonstrated on a genuine failure; pricing it as a proven detection mechanism overstates its record.
  2. The fail-open "consistency with the sibling" claim is currently aspirational: entrypoint-litespeed.sh fails open silently while entrypoint-lsphp.sh fails open loudly with mode strings. Whatever posture is chosen, making the siblings actually consistent (warning + summary line + marker in both) is part of the work.
  3. "Fail-open vs fail-closed" flattens two distinct sub-cases — probe empty vs. probe answered without a scan dir — that have different likely causes and different serving prognoses. The entrypoint already distinguishes them in its reporting; the decision should too (and above, does).