Author SHA1 Message Date
jknapp 78f2462a75 Merge pull request 'docker: set image.source label to GitHub mirror for ghcr.io linking' (#4) from add-ghcr-source-label into main
Build and push coraza-spoa / Build-and-Push (push) Successful in 51s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m3s
2026-06-03 18:24:41 +00:00
shadowdaoandClaude Opus 4.7 8c652ffef9 docker: set image.source label to GitHub mirror for ghcr.io linking
Adds (Dockerfile) and updates (coraza-spoa/Dockerfile) the OCI
image.source label to point at github.com/shadowdao/haproxy-manager-base.
ghcr.io auto-links a package to a GitHub repo when this label resolves
to a github.com URL whose owner+name match the package's owner — that
makes the published packages show up on the GitHub repo sidebar and
inherit its collaborator settings.

Gitea's registry ignores image.source, so changing the value away from
the previous Gitea URL costs nothing on that side.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-03 11:24:26 -07:00
jknapp 379d929e02 ci: mirror image pushes to ghcr.io/shadowdao (#3)
Build and push coraza-spoa / Build-and-Push (push) Successful in 58s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m13s
2026-06-03 17:08:35 +00:00
shadowdaoandClaude Opus 4.7 09455908c5 ci: mirror image pushes to ghcr.io/shadowdao
Adds a second registry login + tag to both build-push workflows so each
build publishes to ghcr.io alongside the in-house Gitea registry. Single
build, two destinations — docker/build-push-action handles the multi-tag
push in one step.

Requires Gitea Actions secret GHCR_TOKEN (a classic PAT with
write:packages on the shadowdao user).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-03 10:08:19 -07:00
shadowdaoandClaude Opus 4.7 e58454c1cc docs: add haproxy-manager-deploy skill
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m10s
Mirror base images / Mirror-Base (map[dst_path:cloud-hosting-platform/golang src:docker.io/library/golang:1.25 tag:1.25]) (push) Successful in 4s
Mirror base images / Mirror-Base (map[dst_path:cloud-hosting-platform/python src:docker.io/library/python:3.12-slim tag:3.12-slim]) (push) Successful in 3s
Procedural discipline for shipping haproxy-manager-base changes.
The flow differs from WHP's (Gitea Actions auto-build vs.
build-release.sh, docker pull + recreate vs. update.sh) and has
its own foot-guns worth codifying:

- /etc/haproxy is a named volume → baked-in image files under that
  path are shadowed on existing deployments; use /haproxy/ instead
- HAProxy lf-file expansion eats single % → literal CSS percentages
  must be doubled (100%%)
- WAF-block synthetic test ACL must be injected AFTER send-spoe-group
  or the SPOE call overwrites the forced action
- coraza-spoa is distroless (no sh); peek inside with docker create
  + docker cp rather than docker exec sh

Both build paths (build-push.yaml for haproxy-manager-base, build-
push-coraza.yaml for coraza-spoa) are surfaced so a contributor
knows which CI run to watch.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 06:02:56 -07:00
shadowdaoandClaude Opus 4.7 c1331a592a waf-block page: escape literal % as %% (HAProxy lf-file expansion)
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
End-to-end test of the 403 page showed CSS `100%` rendering as `100`
and gradient stops `0%, 100%` rendering as `0, 100` — HAProxy's
`lf-file` directive runs log-format expansion over the file content,
and `%` is the format-escape character. Single `%` is consumed by
the expander.

Doubled every literal CSS percentage (`100%%`, `0%%`, etc.) so HAProxy
emits a single `%` in the rendered body. Format expressions like
`%[unique-id]` and `%[req.hdr(host)]` stay single-`%` — those are the
substitutions we want.

Added a comment block at the top of the file documenting the gotcha for
future editors.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 05:48:14 -07:00
shadowdaoandClaude Opus 4.7 d931ab0dbc waf-block: render a real HTML page on Coraza-denied requests
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m55s
Previously a Coraza block returned an empty 403 with only the
`waf-block: request` header — a legitimate site owner caught in a
false-positive had no idea what happened or how to get help.

Now:
- hap_header.tpl: every request gets a unique-id (uuid()) and that ID
  is injected back into the request as X-Request-Reference for the
  backend, so upstream Apache/PHP logs can correlate too.
- hap_listener.tpl: on a request-phase Coraza deny we use
  `http-request return` with `lf-file` instead of `http-request deny`,
  so HAProxy renders the new errors/403-waf.html page with the
  request reference substituted in. The page tells the visitor a
  request was blocked, displays the reference, and points site owners
  to https://secure.anhonesthost.com/submitticket.php to open a ticket
  rather than exposing a public email address (avoids giving
  attackers a flood target).
- The waf-block header and x-request-reference header are still set
  on the response so curl / monitoring clients can pick them up
  without rendering HTML.
- Response-phase deny stays as the bare 403 — outbound blocks are
  rare in our config and an HTML body could land mid-stream.

Errorfile lives at /haproxy/errors/403-waf.html (NOT under
/etc/haproxy/, because that path is a named volume in deployed
containers and would shadow baked-in files on existing deployments).

Support workflow: visitor quotes the reference → support greps
/var/log/haproxy.log for the uuid → gets timestamp + client IP +
Host + URI → greps /var/log/coraza/audit.log for the matching
transaction → reads the rule_id that fired.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-15 05:41:16 -07:00
shadowdao 220b28f0c4 haproxy: use req.hdr_ip for real-IP resolution (string-IP crashed Coraza SPOA)
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 55s
2026-05-14 08:57:05 -07:00
shadowdao 9770398ab0 coraza: pass var(txn.real_ip) instead of src to Coraza (real client IP in WAF logs)
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 55s
2026-05-14 08:52:01 -07:00
shadowdao 633d9390f2 coraza: pin go.mod to 1.23 (matches go mod tidy output; Dockerfile still uses 1.25 image)
Build and push coraza-spoa / Build-and-Push (push) Successful in 42s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
2026-05-14 08:08:38 -07:00
shadowdaoandClaude Sonnet 4.6 6d43308073 coraza: pre-CRS Include for runtime per-host exemptions (load-order fix)
Build and push coraza-spoa / Build-and-Push (push) Successful in 41s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-05-14 07:55:51 -07:00
shadowdaoandClaude Opus 4.7 489290ed33 coraza: ship rules-catalog.json generated from bundled CRS at build time
Build and push coraza-spoa / Build-and-Push (push) Successful in 44s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 52s
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-14 06:57:42 -07:00
shadowdao b2adcdbed9 coraza: reserve rule-ID range 990000000-990999999 for WHP-generated rules 2026-05-14 06:53:37 -07:00
shadowdao 1f1bc1837e coraza: add second Include for runtime-managed local-overrides.conf 2026-05-14 06:51:24 -07:00
shadowdao 753743de20 coraza: drop 913xxx scanner-UA from enforce list (FP on Mastodon + SiteLock)
Build and push coraza-spoa / Build-and-Push (push) Successful in 40s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
25h whp01 burn-in (2026-05-13) found ~11% FP rate on rule 913100:
ActivityPub federation pulls (Mastodon UA "...Bot" on hackerpublicradio.org
and blog.anti-social.online) and SiteLockSpider scans (a customer-paid
security service hitting greggfranklin.com + suchascream.net). The other
six promoted rule families (930120, 932100-160, 933170-200, 944100-300,
920440, 930130) showed zero FPs across the same window and stay enforced.

Detection-only still feeds the anomaly score, so we lose ~no real
blocking value by demoting this family.
2026-05-13 19:13:22 -07:00
shadowdaoandClaude Opus 4.7 5e5234cb14 refactor(suspension): serve via /suspended route on default-backend, drop bk_suspended
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
The previous design used a separate whp-suspended container (nginx:alpine
serving a static 503 page) reachable via a dedicated bk_suspended backend.
That was over-engineered — haproxy-manager-base already ships a default-app
Flask server on :8080 that serves /default-page and /blocked-ip via
path-rewrite ACLs. Mirroring that pattern lets the suspension page live
in the SAME container, no extra image to build, no extra container to
run/health-monitor.

Changes:
- Add /suspended Flask route on default_app returning 503 + suspended_page.html
- Add templates/suspended_page.html (dark-themed 503 page)
- hap_listener.tpl: 'http-request set-path /suspended' + 'use_backend
  default-backend' when host is in suspended_domains.list (same pattern
  as is_blocked_ip)
- Rename env var from HAPROXY_SUSPENSION_BACKEND (a target hostport) to
  HAPROXY_SUSPENSION_ENABLED (a bool); accepts 1/true/yes/on (case-insensitive)
- Remove hap_suspended_backend.tpl and its rendering in generate_config

Non-WHP deployments (env var unset) see byte-identical haproxy.cfg as before
(verified via jinja2 render diff).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 12:08:45 -07:00
shadowdaoandClaude Opus 4.7 6fd07b4c54 fix(suspended): tolerate startup DNS failure + use docker_dns resolvers
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 52s
If the upstream container isn't up when haproxy-manager starts (e.g. when
haproxy is recreated before whp-suspended), the default `init-addr libc` mode
makes haproxy refuse to start — taking down the whole proxy. Switched to
`init-addr last,none` (use last known address, fall back to 0.0.0.0 = DOWN)
and added `resolvers docker_dns` (defined in hap_header.tpl) so the real IP
is picked up once DNS becomes resolvable.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 11:52:50 -07:00
shadowdaoandClaude Opus 4.7 2ef582a3de feat(suspension): opt-in routing for suspended hosts via bk_suspended backend
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 56s
Adds a new env var HAPROXY_SUSPENSION_BACKEND (default unset). When set
(e.g. "whp-suspended:80"), generate_config() renders:
- A bk_suspended backend pointing at the configured upstream
- An ACL `acl is_suspended_domain hdr(host),lower -f /etc/haproxy/suspended_domains.list`
  + `use_backend bk_suspended if is_suspended_domain` in the frontend,
  sitting after IP-blocking and before any per-domain routing
- An empty /etc/haproxy/suspended_domains.list if missing (haproxy refuses
  to start with -f pointing at a non-existent file)

External tooling (e.g. WHP's site_disable.php) maintains the list via
`docker cp` and HUP-reloads the container.

Non-WHP deployments (home networks, standalone use) leave the env var
unset and see byte-identical haproxy.cfg output. Same opt-in shape as
the existing HAPROXY_CORAZA_SPOE_BACKEND integration.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-13 11:46:18 -07:00
shadowdaoandClaude Opus 4.7 3572c66fb7 coraza: promote 920440 + 930130 to enforce list (empirical detect-only data)
Build and push coraza-spoa / Build-and-Push (push) Successful in 1m17s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
After ~30 min of detect-only on whp01 we have actionable data on what
fires against legitimate customer traffic vs. attacker recon. Two rules
demonstrably catch only the latter and earn promotion to the day-one
enforce list:

  920440 — URL file extension restricted by policy
    Caught 124 events in the sample window, ALL backup/config-file
    disclosure probes (`/wp-config.php.old`, `/db_backup.sql`,
    `/.env.save`, `/releases.sql` ...) from a single GCP-hosted scanner
    hammering joshuaknapp.net. Match patterns: .sql (×62), .bak (×5),
    .old (×3), .save (×2), .backup, .dist. No legitimate URL on
    WP/WooCommerce/Divi/HPR ends in these.

  930130 — Restricted File Access Attempt
    Caught 117 events, ALL dotfile/VCS/config-disclosure probes
    (`/.env`, `/.env.local`, `/.env.bak`, `/.git/config`, `/config.php`,
    `/admin/.env`, `/backend/.env` ...). Spread across joshuaknapp.net,
    cgdannyb.com, onlinesupplements.net. Notably, HPR's
    `/ccdn.php?filename=/eps/...` legitimate audio-delivery URL does NOT
    trigger this rule — verified empirically.

Also documented in the "intentionally detect-only" comment block: 933150
fires on WooCommerce checkout when literal `session_start` appears in
billing form data (alphaoneaminos.com saw 2 such events). That's a
canonical CRS false positive on WooCommerce; left detect-only.

Net effect: existing detect_only deployments stay detect-only (the WHP
apply script bind-mounts an empty overrides over the baked-in file).
When operators next flip a server to enforce, these two extra ranges
activate alongside the original day-one list.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 18:00:21 -07:00
shadowdaoandClaude Opus 4.7 ba4c101135 fix(coraza): add deny rules that act on Coraza's verdict + spop-check on backend
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 55s
Two fixes that complete the SPOE enforcement path:

1. Listener was sending requests to Coraza for inspection but never reading
   the result. Coraza-SPOA sets var(txn.coraza.action) to "deny" / "drop"
   / "redirect" when a rule with that disruptive action fires; HAProxy
   needs explicit rules that READ the variable and apply the action.
   Without them, the audit log shows "Access denied" but the request
   still gets HTTP 200 (verified on staging: sqlmap/JNDI/shellinj all
   detected, all returned 200).

   Added the standard six rules from upstream's example/haproxy/haproxy.cfg
   covering http-request + http-response phases for each of deny/drop/
   redirect. Same set the upstream Coraza-SPOA docs recommend.

   Intentionally did NOT add the upstream's fail-CLOSED rule
   `http-request deny deny_status 500 if { var(txn.coraza.error) -m int gt 0 }`
   — for a hosting platform we want fail-open. Documented inline.

2. Backend health check switched from plain TCP `check` to `option
   spop-check`. The spop-check actually negotiates a SPOE session against
   the agent, so HAProxy detects a half-broken SPOA that's listening on
   :9000 but failing protocol handshakes. Plain `check` would miss that.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 17:16:03 -07:00
shadowdaoandClaude Opus 4.7 f1e9bb2c63 fix(coraza-spoe): match upstream's required spoe shape (groups, arg order, names)
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m18s
Three real bugs in the SPOE config caught when HAProxy validated the
generated file:

1. spoe-agent must declare `groups` not `messages`. The `messages` form
   doesn't make the message reachable via `send-spoe-group`; HAProxy
   complained:
     unable to find SPOE group 'coraza-check' into SPOE engine 'coraza'

2. send-spoe-group references a spoe-GROUP name, which needs its own
   block. Added `spoe-group coraza-req { messages coraza-req }` as
   the indirection layer.

3. Arg names + ORDER are required to match what Coraza-SPOA parses
   positionally. My version had `dest-ip`/`dest-port`; upstream's
   example/haproxy/coraza.cfg (v0.7.1) uses `dst-ip`/`dst-port`.
   Renamed and reordered to match upstream verbatim, including the
   `app=str(haproxy)` literal that matches our config.yaml application
   name.

Also corrected misleading comment about `set-on-error continue`: that
option actually sets a variable on error; the fail-open behavior comes
from us deliberately NOT adding a `http-request deny if errored` rule
in the frontend. Renamed the variable to `error` (matching upstream)
and updated comments to be accurate.

Listener template's send-spoe-group action updated to reference the
new group name `coraza-req`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 17:12:09 -07:00
shadowdaoandClaude Opus 4.7 061309675b fix(coraza-spoe): collapse args to one line + ensure trailing LF on spoe cfg
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 2m11s
Two HAProxy parse errors caught in staging functional test:

1. coraza-spoe.cfg:39 'args': missing fetch method
   The args directive had backslash line continuations. HAProxy doesn't
   support those in SPOE configs — args must be one physical line.
   Collapsed to a single line.

2. coraza-spoe.cfg:50 Missing LF on last line
   Same trailing-LF issue we hit on haproxy.cfg one commit ago. The
   Jinja2 template ends with content rather than a newline, and write()
   doesn't add one. Belt-and-suspenders: explicitly append '\n' before
   writing if not already there.

After this commit HAProxy validates the generated config cleanly. Will
verify on staging now (combined SPOE injection + fail-open + active
attack-detection tests).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 17:07:12 -07:00
shadowdaoandClaude Opus 4.7 4769f67fe9 fix(coraza): ensure haproxy.cfg ends with LF when SPOE backend appended
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
The SPOE backend block from hap_coraza_spoa_backend.tpl was being appended
last to config_parts. The template's render output doesn't end with a
newline (and config_parts is joined with '\n' BETWEEN elements, not after
the last one), so the resulting haproxy.cfg ended on `server coraza-spoa
...` with no trailing LF. HAProxy refuses to parse such files:

    [ALERT] config: parsing [/etc/haproxy/haproxy.cfg:288]: Missing LF
    on last line, file might have been truncated at position 70.

Match the existing pattern at the previous-last config_parts.append
(line 1850 uses `'\n'.join(config_backends) + '\n'`) and add an explicit
'\n' on the coraza block append.

Caught immediately on staging: HTTP 000 to localhost:80 because HAProxy
never started; gunicorn/management API kept serving on :8000 fine.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 17:03:56 -07:00
shadowdaoandClaude Opus 4.7 3e1f9dda2b fix(template): strip Jinja2 whitespace so no-env-var listener is byte-identical
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m23s
Default Jinja2 {% if %}{% endif %} block syntax leaves a trailing newline
even when the conditional doesn't render. Staging verification of PR 2
showed the resulting haproxy.cfg differed from the pre-PR2 version by
exactly 1 blank line — semantically identical but not byte-identical,
which violates the design promise that haproxy-manager-base's default
output stays unchanged for home/standalone deployments.

Use {%- if -%}/{%- endif %} (the whitespace-stripping variants) so the
block contributes zero bytes when coraza_spoe_backend is unset.

Verified locally: without env var = 55 lines, ends cleanly on the
is_blocked_ip rule. With env var = 62 lines, +7 for the SPOE block.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:59:40 -07:00
shadowdaoandClaude Opus 4.7 73b9104565 PR 2/3: opt-in SPOE integration for Coraza WAF
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m59s
Adds the plumbing that lets haproxy-manager talk to the coraza-spoa sidecar
added in PR 1, while keeping the default behavior bit-identical for any
deployment that doesn't set the new env var (the home network / standalone
use cases).

Single gate: HAPROXY_CORAZA_SPOE_BACKEND env var on the haproxy-manager
container. Unset (default) = generate_config() renders zero SPOE-related
output. Set (e.g. "coraza-spoa:9000") = three things happen at config
generation time:

  1. hap_listener.tpl injects 5 lines at the end of the frontend block:
       filter spoe engine coraza config /etc/haproxy/coraza-spoe.cfg
       http-request send-spoe-group coraza coraza-check
     ...placed AFTER rate-limit and IP-block guards so we don't waste WAF
     calls on requests we were going to drop anyway.

  2. A new TCP backend (hap_coraza_spoa_backend.tpl) is appended:
       backend coraza-spoa-backend
           mode tcp
           server coraza-spoa <env-var-target> check ...

  3. The SPOE engine config (hap_coraza_spoe_engine.tpl) is rendered and
     written to /etc/haproxy/coraza-spoe.cfg, defining the spoe-agent
     "coraza" + spoe-message "coraza-check". This sets:
       - option set-on-error continue   (FAIL-OPEN if SPOA is unreachable)
       - timeout processing 100ms       (per-request inspection budget)
       - app=str(haproxy)               (matches sidecar's application name)

Verification (template render only, before staging deploy):
  - hap_listener.tpl with no env var: 55 lines, zero SPOE references
  - hap_listener.tpl with env var:    62 lines, filter + send-spoe-group present
  - Engine cfg + backend block render with correct agent_target substitution

Next: PR 3 wires this into WHP (sidecar deploy via container-manager.sh
extension, server-settings UI for on/off, AI Monitor source for the audit
log). Staging verification of PR 1 + PR 2 together happens after PR 3.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:49:29 -07:00
shadowdaoandClaude Opus 4.7 4e0c22e9c9 ci: mirror golang:1.25 alongside python:3.12-slim, switch coraza-spoa FROM
Build and push coraza-spoa / Build-and-Push (push) Successful in 1m16s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m18s
Cloudflare's bot-management incident on 2026-05-12 took out docker.io blob
pulls twice in one day — first for python:3.12-slim (mirrored in 5a2ebf9),
then again for golang:1.25 when the PR 1 coraza-spoa build hit the same
R2-via-Cloudflare failure on the build stage's base image.

Restructure .gitea/workflows/mirror-base-image.yaml into a matrix that
iterates over a list of (src, dst_path, tag) entries. Adding a new base
image is now a one-line matrix entry. fail-fast: false so one image's
upstream being down doesn't block refreshing the others.

Switch coraza-spoa/Dockerfile's build stage FROM to the in-house golang
mirror. Runtime FROM (gcr.io/distroless/static-debian12:nonroot) stays
on upstream — distroless is on Google's registry, separate from Docker
Hub's Cloudflare R2 setup, and didn't fail during today's incident.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:40:42 -07:00
shadowdaoandClaude Opus 4.7 e4c506bcd9 PR 1/3: add coraza-spoa sidecar image
Build and push coraza-spoa / Build-and-Push (push) Failing after 24s
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 55s
Self-contained sidecar that runs Coraza-SPOA v0.7.1 (latest upstream as of
2026-05-08, with OWASP CRS bundled in the binary). HAProxy will consult it
per-request via SPOE in PR 2; for now this PR ships the image only.

Defines:
- coraza-spoa/Dockerfile       — multi-stage build (golang:1.25 -> distroless),
                                 pinned to v0.7.1, ARG-overridable
- coraza-spoa/config.yaml      — single application "haproxy", JSON audit log
                                 to /var/log/coraza/audit.log, SecRuleEngine
                                 DetectionOnly globally
- coraza-spoa/overrides.conf   — day-one enforce list: scanner UAs (913xxx),
                                 RCE shell injection (932100-932160),
                                 webshell paths (933170-933200), targeted LFI
                                 (930120), Log4Shell/JNDI (944100-944300).
                                 Rationale per-range documented inline.
                                 Detect-only for XSS/SQLi/protocol (high FP
                                 on WP/WooCommerce/Divi customer mix).
- coraza-spoa/README.md        — deployment shape, audit log location, pin
                                 upgrade procedure, false-positive tuning.
- .gitea/workflows/build-push-coraza.yaml — Gitea Action triggered on
                                 coraza-spoa/** changes, publishes
                                 repo.anhonesthost.net/cloud-hosting-platform/
                                 coraza-spoa:latest. Path-scoped so it
                                 doesn't fire on every haproxy-manager push.

No changes to haproxy-manager-base itself in this PR — the existing image
stays bit-identical, used standalone in home networks and other projects
without dependency on this sidecar. PR 2 will add the OPT-IN template
plumbing that lets haproxy-manager call out to this agent when an env var
is set.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:28:44 -07:00
shadowdaoandClaude Opus 4.7 55670daf5b ci: add weekly Gitea Action to mirror python:3.12-slim into in-house registry
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m16s
Companion to the Dockerfile change in 5a2ebf9. The previous manual refresh
note in the Dockerfile becomes automated: a workflow_dispatch + weekly cron
that pulls python:3.12-slim from docker.io and re-pushes it to
repo.anhonesthost.net/cloud-hosting-platform/python:3.12-slim.

Workflow can also be triggered manually from the Gitea UI when Python
publishes patches between cron firings. Logs the upstream and mirror digests
so it's easy to verify "did the mirror really update" after a run.

If more base images need mirroring later (haproxy itself, alpine, etc.),
this workflow should be promoted to a matrix or moved to a dedicated infra
repo — keeping it co-located with haproxy-manager-base for now since it's
the only consumer.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:18:32 -07:00
shadowdaoandClaude Opus 4.7 5a2ebf991c ci: mirror python:3.12-slim into in-house registry
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m58s
docker.io serves image blobs from Cloudflare R2. The 2026-05-12 Cloudflare
incident took out blob pulls for hours and broke this image's Gitea CI
build mid-way through the haproxy-manager gunicorn migration (commit
bdd7d2f). With the base image mirrored at repo.anhonesthost.net,
CI builds no longer depend on docker.io reachability.

Refresh procedure documented in the Dockerfile comment block. Manual
re-push monthly or when Python patches drop. A future Gitea Action could
automate the pull-tag-push so we always have a current base.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 16:08:44 -07:00
shadowdaoandClaude Opus 4.7 bdd7d2f098 swap werkzeug dev server for gunicorn + accept all HTTP methods on default/blocked pages
HAProxy Manager Build and Push / Build-and-Push (push) Failing after 13s
Two related fixes for the issues the AI Monitor surfaced on whp01 on
2026-05-12 (haproxy-manager going "healthy but stalled" after long
uptime, and noise from POST /blocked-ip returning 405):

1. Production WSGI server. The Flask app was running on werkzeug's
   built-in dev server (the one that prints "WARNING: This is a
   development server" on every startup). werkzeug is single-threaded
   and accumulates worker state over long uptimes; after ~24h on whp01
   the health endpoint stops responding while the container still
   reports "healthy" because Docker's HEALTHCHECK uses an HTTP probe
   from inside the same werkzeug process that's stalled.

   Replace with gunicorn (gthread worker class, --max-requests=1000
   with jitter so workers recycle periodically). Two gunicorn instances,
   one per Flask app — port 8000 for the management API, port 8080 for
   the default/blocked-ip page server. Both lift their app objects from
   the haproxy_manager module so gunicorn can import them.

   Required structural change: default_app was created INSIDE the
   __name__ == '__main__' block at module bottom, where gunicorn could
   never reach it. Moved to module level. The __main__ block now stays
   only for `python haproxy_manager.py` local-dev workflow.

   Container init (init_db, certbot register, generate_config,
   start_haproxy) extracted into a do_initial_setup() function called
   from a new scripts/init.py. start-up.sh runs init.py to completion
   before either gunicorn binds, which keeps HAProxy startup off the
   WSGI workers' fork paths (no race between workers all trying to
   start_haproxy() at once).

2. /blocked-ip and / accept ALL methods. HAProxy proxies blocked-IP
   traffic to default_app preserving the original verb, so a blocked
   POST request used to hit Flask's GET-only route and get a 405 +
   the AI Monitor flagged the noise. Adding the full method list lets
   the 403 page render regardless of verb.

Gunicorn settings tunable via env (workers, timeout, max-requests).
API gets --timeout 120 because ACME cert issuance can be slow; the
default page server stays on the gunicorn default 30s.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-12 15:24:28 -07:00
shadowdaoandClaude Opus 4.7 8a86beac73 feat: clear stale certbot lock files before each ACME run + at startup
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
certbot uses fasteners (fcntl-based locking) to serialize concurrent
invocations. The kernel auto-releases fcntl locks when the holding
process exits, but the .certbot.lock FILES persist on disk — and we've
seen real cases where subsequent runs report "Another instance of
Certbot is already running" even when no certbot process is alive.
Observed during the 2026-05-09 bundling rollout when a hung worker
held a lock across container-internal Python crashes.

When SSL is blocked on a customer site, this is high-impact: the
certbot lock can sit stale until somebody manually deletes it.

clear_stale_certbot_locks():
  - probes each known lock path with fcntl.LOCK_NB
  - if the lock is unheld → file is stale → delete it
  - if the lock IS held → leave it alone (real certbot is running)

Wired in:
  - container startup (init block)
  - /api/ssl single-domain handler
  - /api/ssl/bundle handler
  - /api/certificates/renew handler

Safe to call repeatedly; never deletes a lock a real process holds, so
can never trigger concurrent certbot runs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 12:09:19 -07:00
shadowdaoandClaude Opus 4.7 f7ef34b988 feat(api/ssl/bundle): clean up superseded lineages after issuance
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
The bundle endpoint correctly issued multi-SAN certs but left old
single-SAN .pem files (e.g. <name>-0001.pem) in /etc/haproxy/certs/.
HAProxy's `bind ... ssl crt /etc/haproxy/certs` loads everything in the
directory and picked the alphabetically-first matching file — typically
the older single-SAN one — so the new bundle had no effect on what was
served. Repro on peptidesaver.net: bundle covered 4 SANs but HAProxy
kept serving peptidesaver.net-0001.pem (single SAN, April-issued).

After a successful bundle write, walk SSL_CERTS_DIR and remove any
.pem whose CN is in the new bundle's name list (excluding the bundle's
own combined file). Drop the matching certbot lineage with
`certbot delete --cert-name <X> -n` so `certbot renew` stops touching
the dead lineage too.

Returns a `cleanup` summary in the API response so callers can log /
display what was deleted.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 11:58:21 -07:00
shadowdaoandClaude Opus 4.7 90255cc4b3 feat(api): add /api/ssl/bundle for per-site SAN cert issuance
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m1s
WHP's renewal orchestrator now bundles a site's domains into one cert
covering all SANs, instead of N separate single-domain orders. Single
ACME order = better behavior under Let's Encrypt's 50/hour orders limit
when many domains need attention at once.

Endpoint: POST /api/ssl/bundle
Body: {"primary": "example.com", "sans": ["www.example.com", ...]}

- Uses --cert-name <primary> so the lineage stays stable across renewals
  (no -0001/-0002 proliferation seen with the legacy single-domain flow).
- Single combined .pem at /etc/haproxy/certs/<primary>.pem; HAProxy SNI-
  matches against the cert's SAN list, so one file serves all included
  hostnames.
- Updates the domains table for every SAN in the bundle.
- Hard cap at 100 SANs (LE limit).

Existing /api/ssl single-domain endpoint kept for backwards compat.
The WHP haproxy_manager::bundleSSL() helper falls back to a per-domain
loop if /api/ssl/bundle returns 404, so the WHP side keeps working
during the rolling image upgrade window.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-05-09 11:32:15 -07:00
shadowdaoandClaude Opus 4.7 b731feab12 Self-heal trusted IP whitelist files at startup
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 3m26s
Volume-mounted /etc/haproxy can shadow the image-baked
trusted_ips.list/trusted_ips.map, causing HAProxy to fail
config validation with "failed to open pattern file" on
non-WHP deployments. Touch empty files if they don't exist
so the ACLs always parse.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-04-30 10:02:16 -07:00
shadowdaoandClaude Opus 4.6 615044fa14 Fix resolvers block placement — must be outside global section
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
The resolvers section was inserted inside the global section, causing
HAProxy to parse global directives (pidfile, maxconn, etc.) as
resolver keywords. Moved resolvers to its own top-level section
between global and defaults where HAProxy expects it.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-02 05:18:48 -07:00
shadowdaoandClaude Opus 4.6 cf4eb5092c Add DNS resolver for automatic container IP re-resolution
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
When Docker containers restart, they can get new IPs on the bridge
network. HAProxy caches DNS at config load time, so stale IPs cause
503s until config is regenerated.

Added a 'docker_dns' resolvers section pointing to Docker's embedded
DNS (127.0.0.11) with 10s hold time. Backend servers now use
'resolvers docker_dns init-addr last,libc,none' so HAProxy:
- Re-resolves container names every 10 seconds
- Falls back to last known IP if DNS is temporarily unavailable
- Starts even if a backend can't be resolved yet (init-addr none)

This eliminates 503s from container restarts, scaling, and recreation
without requiring a HAProxy config regeneration.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 22:27:07 -07:00
shadowdaoandClaude Opus 4.6 ecf891ff02 Don't abort cert renewal when a single domain fails
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 2m11s
The renewal script was exiting immediately when certbot returned a
non-zero exit code, which happens when ANY cert fails to renew. A
single dead domain (e.g., DNS no longer pointed here) would block
ALL other certificates from being processed and combined for HAProxy.

Now logs the failures but continues to copy/combine successfully
renewed certificates and reload HAProxy.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 15:17:15 -07:00
shadowdaoandClaude Opus 4.6 3da5df67d0 Update CLAUDE.md with HAProxy hardening and AI log monitor docs
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 2m36s
Documents HAProxy health checks, watchdog, rate limiting, trusted IP
whitelist, timeout hardening, HTTP/2 protection, and the AI-powered
log monitor system with two-tier analysis, auto-remediation, and
notification support.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-04-01 08:16:44 -07:00
shadowdaoandClaude Opus 4.6 da40328438 Fix: remove comments from trusted IP files breaking HAProxy startup
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 54s
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 14:19:29 -07:00
shadowdaoandClaude Opus 4.6 13a5be636e Raise rate limits further for media-heavy sites
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 53s
Generous thresholds that accommodate sites with many images/assets
while still catching obvious automated floods:
- Request rate: tarpit at 300 req/s, block at 500 req/s
- Connection rate: 500/10s
- Concurrent connections: 500
- Error rate: 100/30s

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 14:12:24 -07:00
shadowdaoandClaude Opus 4.6 5390ebb8a6 Raise rate limit thresholds to avoid false positives on normal traffic
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m22s
Previous thresholds (200/500 req/10s) were too aggressive — WordPress
login pages with their CSS/JS/image assets can easily burst 30-50
requests per page load, triggering tarpits and blocks on legitimate
users.

New thresholds:
- Request rate: tarpit at 1000/10s (100 req/s), block at 2000/10s (200 req/s)
- Connection rate: 300/10s (was 150)
- Concurrent connections: 200 (was 100)
- Error rate: 50/30s (was 20)

These still catch real floods and scanners while giving normal web
traffic plenty of headroom.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 14:10:53 -07:00
shadowdaoandClaude Opus 4.6 53d259bd3f Add trusted IP whitelist for rate limit bypass
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m25s
Adds trusted_ips.list and trusted_ips.map files that exempt specific
IPs from all rate limiting rules. Supports both direct source IP
matching (is_trusted_ip) and proxy-header real IP matching
(is_whitelisted). Files are baked into the image and can be updated
by editing and rebuilding.

Adds phone system IP 172.116.197.166 to the whitelist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-31 13:39:41 -07:00
21 changed files with 109 additions and 1273 deletions
+9 -37
View File
@@ -23,32 +23,6 @@ Do not skip the verify step. The container can come up "healthy" while still ser
--- ---
## Step 0a — Resolve the target host (never hardcoded)
This skill deliberately does **not** bake in a server hostname — this repo is mirrored to a public remote, so a real FQDN in the skill would leak into commits. Instead, resolve the deploy target into a `DEPLOY_HOST` shell variable that every `ssh` command below uses.
```bash
HOST_FILE=".claude/skills/haproxy-manager-deploy/target-host.local"
DEPLOY_HOST="$(cat "$HOST_FILE" 2>/dev/null)"
```
- **If `$DEPLOY_HOST` is non-empty**, use it — that's the user's saved target. The file is gitignored, so the real hostname never lands in a commit.
- **If it's empty**, ask the user which server this deploy targets (e.g. production vs. staging) and what its hostname or SSH alias is. Then offer to save it so future deploys don't have to ask:
```bash
echo 'the-host-they-gave.example' > "$HOST_FILE" # gitignored — safe to store the real FQDN here
```
Confirm `$DEPLOY_HOST` is set before running any `ssh` step:
```bash
[ -n "$DEPLOY_HOST" ] || echo "DEPLOY_HOST not set — ask the user for the target server"
```
All commands below assume the variable is set in the same shell session (`ssh root@"$DEPLOY_HOST" ...`).
---
## Step 0 — Confirm before pushing ## Step 0 — Confirm before pushing
If the user just said "deploy" or "ship the haproxy fix", confirm what's actually changing: a template, the Python manager, the coraza-spoa subdir (separate image, separate workflow), or a static asset. Look at `git status` and `git diff` and read the diff back to the user if it's non-trivial. If the user just said "deploy" or "ship the haproxy fix", confirm what's actually changing: a template, the Python manager, the coraza-spoa subdir (separate image, separate workflow), or a static asset. Look at `git status` and `git diff` and read the diff back to the user if it's non-trivial.
@@ -98,7 +72,7 @@ Pushing immediately triggers the Gitea Actions build.
The Go build inside coraza-spoa takes ~2-3 minutes; the haproxy-manager-base build is faster (~1-2 min). Don't bother polling the runs UI — just pull on the target server until the digest changes: The Go build inside coraza-spoa takes ~2-3 minutes; the haproxy-manager-base build is faster (~1-2 min). Don't bother polling the runs UI — just pull on the target server until the digest changes:
```bash ```bash
ssh root@"$DEPLOY_HOST" 'until docker pull -q repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest 2>&1 | tail -1 | grep -qE "Image is up to date|Status: Downloaded"; do sleep 15; done' ssh root@whp01.cloud-hosting.io 'until docker pull -q repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest 2>&1 | tail -1 | grep -qE "Image is up to date|Status: Downloaded"; do sleep 15; done'
``` ```
`-q` suppresses the noisy layer progress so the grep can match cleanly. If you started this command before the CI build finished, it'll loop until the new image lands; once the digest matches, it exits. `-q` suppresses the noisy layer progress so the grep can match cleanly. If you started this command before the CI build finished, it'll loop until the new image lands; once the digest matches, it exits.
@@ -106,7 +80,7 @@ ssh root@"$DEPLOY_HOST" 'until docker pull -q repo.anhonesthost.net/cloud-hostin
To confirm you got the new image, check the image-creation time vs your push: To confirm you got the new image, check the image-creation time vs your push:
```bash ```bash
ssh root@"$DEPLOY_HOST" 'docker images repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base --format "{{.CreatedSince}}"' ssh root@whp01.cloud-hosting.io 'docker images repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base --format "{{.CreatedSince}}"'
``` ```
It should say "X minutes ago" matching the build wait, not "yesterday". It should say "X minutes ago" matching the build wait, not "yesterday".
@@ -119,10 +93,10 @@ The image is `gcr.io/distroless/static-debian12:nonroot`-based, no shell. To pee
```bash ```bash
# haproxy-manager-base (has sh): # haproxy-manager-base (has sh):
ssh root@"$DEPLOY_HOST" 'docker run --rm --entrypoint sh repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest -c "ls /haproxy/errors/ && head -5 /haproxy/errors/403-waf.html"' ssh root@whp01.cloud-hosting.io 'docker run --rm --entrypoint sh repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest -c "ls /haproxy/errors/ && head -5 /haproxy/errors/403-waf.html"'
# coraza-spoa (distroless, no sh) — use docker create + docker cp instead: # coraza-spoa (distroless, no sh) — use docker create + docker cp instead:
ssh root@"$DEPLOY_HOST" 'docker create --name _peek repo.anhonesthost.net/cloud-hosting-platform/coraza-spoa:latest && docker cp _peek:/etc/coraza/config.yaml - | tar xO; docker rm _peek' ssh root@whp01.cloud-hosting.io 'docker create --name _peek repo.anhonesthost.net/cloud-hosting-platform/coraza-spoa:latest && docker cp _peek:/etc/coraza/config.yaml - | tar xO; docker rm _peek'
``` ```
This step exists because the CI build can succeed but ship the wrong file (wrong commit pulled, build cache issue, etc.). Catching it here is one step earlier than catching it from a customer report. This step exists because the CI build can succeed but ship the wrong file (wrong commit pulled, build cache issue, etc.). Catching it here is one step earlier than catching it from a customer report.
@@ -132,12 +106,12 @@ This step exists because the CI build can succeed but ship the wrong file (wrong
## Step 6 — Recreate the container ## Step 6 — Recreate the container
```bash ```bash
ssh root@"$DEPLOY_HOST" '/root/whp/scripts/container-manager.sh recreate haproxy-manager' ssh root@whp01.cloud-hosting.io '/root/whp/scripts/container-manager.sh recreate haproxy-manager'
``` ```
For coraza-spoa changes: For coraza-spoa changes:
```bash ```bash
ssh root@"$DEPLOY_HOST" '/root/whp/scripts/container-manager.sh recreate coraza-spoa' ssh root@whp01.cloud-hosting.io '/root/whp/scripts/container-manager.sh recreate coraza-spoa'
``` ```
`container-manager.sh recreate` does: stop, remove, docker pull (idempotent if already pulled), start with the right flags from settings.json. **It reads `/docker/whp/settings.json` for things like `coraza_waf.mode`**, so if the user has toggled mode while you were building, the recreated container reflects the current setting — not whatever it was when you started. `container-manager.sh recreate` does: stop, remove, docker pull (idempotent if already pulled), start with the right flags from settings.json. **It reads `/docker/whp/settings.json` for things like `coraza_waf.mode`**, so if the user has toggled mode while you were building, the recreated container reflects the current setting — not whatever it was when you started.
@@ -149,7 +123,7 @@ ssh root@"$DEPLOY_HOST" '/root/whp/scripts/container-manager.sh recreate coraza-
For haproxy-manager: For haproxy-manager:
```bash ```bash
ssh root@"$DEPLOY_HOST" ' ssh root@whp01.cloud-hosting.io '
echo "=== container ===" echo "=== container ==="
docker ps --filter name=haproxy-manager --format "image: {{.Image}} status: {{.Status}}" docker ps --filter name=haproxy-manager --format "image: {{.Image}} status: {{.Status}}"
echo "=== healthy ===" echo "=== healthy ==="
@@ -180,20 +154,18 @@ If your change affects what a visitor sees (block pages, redirects, security res
```bash ```bash
# Inject a temporary ACL that forces the WAF deny path on a custom header, # Inject a temporary ACL that forces the WAF deny path on a custom header,
# fire one request, observe the rendered response, then revert + reload. # fire one request, observe the rendered response, then revert + reload.
ssh root@"$DEPLOY_HOST" ' ssh root@whp01.cloud-hosting.io '
docker exec haproxy-manager cp /etc/haproxy/haproxy.cfg /tmp/cfg-bak docker exec haproxy-manager cp /etc/haproxy/haproxy.cfg /tmp/cfg-bak
docker exec haproxy-manager sh -c "sed -i \"/http-request send-spoe-group coraza coraza-req/a\\\\ http-request set-var(txn.coraza.action) str(deny) if { req.hdr(x-force-waf-block) -m str yes }\" /etc/haproxy/haproxy.cfg" docker exec haproxy-manager sh -c "sed -i \"/http-request send-spoe-group coraza coraza-req/a\\\\ http-request set-var(txn.coraza.action) str(deny) if { req.hdr(x-force-waf-block) -m str yes }\" /etc/haproxy/haproxy.cfg"
docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null
sleep 1 sleep 1
curl -sSk -D - -H "x-force-waf-block: yes" -H "Host: <live-vhost>" "https://localhost/" | head -40 curl -sSk -D - -H "x-force-waf-block: yes" -H "Host: hub.hackerpublicradio.org" "https://localhost/" | head -40
# revert # revert
docker exec haproxy-manager cp /tmp/cfg-bak /etc/haproxy/haproxy.cfg docker exec haproxy-manager cp /tmp/cfg-bak /etc/haproxy/haproxy.cfg
docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null
' '
``` ```
**Pick a real `<live-vhost>`.** The `Host:` header must match a domain currently served by this haproxy-manager, or the request won't route to the WAF path. Don't hardcode a customer hostname in this skill — pull a live one at test time (any entry from the panel's domain list, or `docker exec haproxy-manager ls /etc/letsencrypt/live`) and substitute it.
**The injection point matters.** Insert AFTER `http-request send-spoe-group coraza coraza-req`, because the SPOE call overwrites `txn.coraza.action` based on the real Coraza verdict — if you inject before it, your override is wiped. **The injection point matters.** Insert AFTER `http-request send-spoe-group coraza coraza-req`, because the SPOE call overwrites `txn.coraza.action` based on the real Coraza verdict — if you inject before it, your override is wiped.
**The reload mechanism matters.** Use `echo reload | socat stdio /tmp/haproxy-cli` — the container is python-based but doesn't have `kill` in PATH, and `docker kill --signal=HUP` signals the python manager (PID 1), not haproxy. **The reload mechanism matters.** Use `echo reload | socat stdio /tmp/haproxy-cli` — the container is python-based but doesn't have `kill` in PATH, and `docker kill --signal=HUP` signals the python manager (PID 1), not haproxy.
-15
View File
@@ -36,26 +36,11 @@ jobs:
username: shadowdao username: shadowdao
password: ${{ secrets.GHCR_TOKEN }} password: ${{ secrets.GHCR_TOKEN }}
# Read the human-readable release version from the VERSION file so every
# build is pinnable for rollback (alongside the immutable git SHA). Bump
# VERSION (YYYY.MM.N) in the same commit as a release-worthy change.
- name: Read version
id: ver
run: echo "version=$(cat VERSION)" >> "$GITHUB_OUTPUT"
- name: Build Image - name: Build Image
uses: docker/build-push-action@v6 uses: docker/build-push-action@v6
with: with:
platforms: linux/amd64 platforms: linux/amd64
push: true push: true
build-args: |
VERSION=${{ steps.ver.outputs.version }}
# Three tags per registry: :latest (moving), :<version> (human-readable
# release), :<sha> (immutable, guaranteed-unique rollback target).
tags: | tags: |
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:${{ steps.ver.outputs.version }}
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:${{ gitea.sha }}
ghcr.io/shadowdao/haproxy-manager-base:latest ghcr.io/shadowdao/haproxy-manager-base:latest
ghcr.io/shadowdao/haproxy-manager-base:${{ steps.ver.outputs.version }}
ghcr.io/shadowdao/haproxy-manager-base:${{ gitea.sha }}
-4
View File
@@ -37,7 +37,3 @@ ENV/
# OS # OS
.DS_Store .DS_Store
Thumbs.db Thumbs.db
# Local-only deploy config (never commit real hostnames)
.claude/skills/haproxy-manager-deploy/target-host.local
*.local
+1 -1
View File
@@ -94,7 +94,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
- `trusted_ips.list` — Source IP whitelist for rate limit bypass (one CIDR/IP per line) - `trusted_ips.list` — Source IP whitelist for rate limit bypass (one CIDR/IP per line)
- `trusted_ips.map` — Real IP whitelist for proxy-header matching (format: `<IP> 1`) - `trusted_ips.map` — Real IP whitelist for proxy-header matching (format: `<IP> 1`)
- Both files are baked into the Docker image via `COPY` in the Dockerfile - Both files are baked into the Docker image via `COPY` in the Dockerfile
- Ship as comment-only templates (no real IPs). Add trusted IPs locally and do **not** commit them — this repo is mirrored publicly. Entries persist in the `/etc/haproxy` named volume across recreates - Currently contains phone system IP `172.116.197.166`
### Timeout Hardening (hap_header.tpl) ### Timeout Hardening (hap_header.tpl)
+1 -7
View File
@@ -14,13 +14,9 @@ FROM repo.anhonesthost.net/cloud-hosting-platform/python:3.12-slim
# sidebar; pointing at the public GitHub mirror enables that linking. The # sidebar; pointing at the public GitHub mirror enables that linking. The
# canonical source-of-truth git remote is still Gitea, but Gitea's registry # canonical source-of-truth git remote is still Gitea, but Gitea's registry
# doesn't consume this label, so there's no contention. # doesn't consume this label, so there's no contention.
# Stamped from the VERSION file by CI (build-arg) so `docker inspect` reports
# what's running on any host. Defaults to "dev" for local/manual builds.
ARG VERSION=dev
LABEL org.opencontainers.image.title="haproxy-manager-base" \ LABEL org.opencontainers.image.title="haproxy-manager-base" \
org.opencontainers.image.description="HAProxy management API with Let's Encrypt automation, Coraza WAF integration, and template-driven config" \ org.opencontainers.image.description="HAProxy management API with Let's Encrypt automation, Coraza WAF integration, and template-driven config" \
org.opencontainers.image.source="https://github.com/shadowdao/haproxy-manager-base" \ org.opencontainers.image.source="https://github.com/shadowdao/haproxy-manager-base" \
org.opencontainers.image.version="${VERSION}" \
org.opencontainers.image.licenses="MIT" org.opencontainers.image.licenses="MIT"
RUN apt update -y && apt dist-upgrade -y && apt install socat haproxy cron certbot curl jq net-tools -y && apt clean && rm -rf /var/lib/apt/lists/* RUN apt update -y && apt dist-upgrade -y && apt install socat haproxy cron certbot curl jq net-tools -y && apt clean && rm -rf /var/lib/apt/lists/*
@@ -47,9 +43,7 @@ RUN mkdir -p /var/spool/cron/crontabs && \
echo '0 */12 * * * /haproxy/scripts/renew-certificates.sh >> /var/log/haproxy-manager.log 2>&1' >> /var/spool/cron/crontabs/root && \ echo '0 */12 * * * /haproxy/scripts/renew-certificates.sh >> /var/log/haproxy-manager.log 2>&1' >> /var/spool/cron/crontabs/root && \
chmod 600 /var/spool/cron/crontabs/root && \ chmod 600 /var/spool/cron/crontabs/root && \
chown root:crontab /var/spool/cron/crontabs/root chown root:crontab /var/spool/cron/crontabs/root
# 443/udp carries HTTP/3 (QUIC). EXPOSE is documentation only — the container EXPOSE 80 443 8000
# must still be run with `-p 443:443/udp` for the UDP listener to be reachable.
EXPOSE 80 443 443/udp 8000
# Add health check # Add health check
HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3 \ HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3 \
CMD curl -sf --max-time 5 http://localhost:8000/health && curl -s --max-time 5 -o /dev/null http://localhost/ || exit 1 CMD curl -sf --max-time 5 http://localhost:8000/health && curl -s --max-time 5 -o /dev/null http://localhost/ || exit 1
+6 -6
View File
@@ -6,10 +6,10 @@ A Flask-based API service for managing HAProxy configurations with dynamic SSL c
To run the container: To run the container:
```bash ```bash
# Without API key authentication (default) # Without API key authentication (default)
docker run -d -p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy --name haproxy-manager your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest docker run -d -p 80:80 -p 443:443 -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy --name haproxy-manager repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
# With API key authentication (recommended for production) # With API key authentication (recommended for production)
docker run -d -p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy -e HAPROXY_API_KEY=your-secure-api-key-here --name haproxy-manager your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest docker run -d -p 80:80 -p 443:443 -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy -e HAPROXY_API_KEY=your-secure-api-key-here --name haproxy-manager repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
``` ```
## Features ## Features
@@ -394,7 +394,7 @@ You can customize the default page by setting environment variables:
```bash ```bash
docker run -d \ docker run -d \
-p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 \ -p 80:80 -p 443:443 -p 8000:8000 \
-v lets-encrypt:/etc/letsencrypt \ -v lets-encrypt:/etc/letsencrypt \
-v haproxy:/etc/haproxy \ -v haproxy:/etc/haproxy \
-e HAPROXY_API_KEY=your-secure-api-key-here \ -e HAPROXY_API_KEY=your-secure-api-key-here \
@@ -402,7 +402,7 @@ docker run -d \
-e HAPROXY_DEFAULT_MAIN_MESSAGE="This website is currently under construction and will be available soon." \ -e HAPROXY_DEFAULT_MAIN_MESSAGE="This website is currently under construction and will be available soon." \
-e HAPROXY_DEFAULT_SECONDARY_MESSAGE="Please check back later or contact us for more information." \ -e HAPROXY_DEFAULT_SECONDARY_MESSAGE="Please check back later or contact us for more information." \
--name haproxy-manager \ --name haproxy-manager \
your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
``` ```
## Example Usage ## Example Usage
@@ -411,12 +411,12 @@ docker run -d \
```bash ```bash
# Start container with API key # Start container with API key
docker run -d \ docker run -d \
-p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 \ -p 80:80 -p 443:443 -p 8000:8000 \
-v lets-encrypt:/etc/letsencrypt \ -v lets-encrypt:/etc/letsencrypt \
-v haproxy:/etc/haproxy \ -v haproxy:/etc/haproxy \
-e HAPROXY_API_KEY=your-secure-api-key-here \ -e HAPROXY_API_KEY=your-secure-api-key-here \
--name haproxy-manager \ --name haproxy-manager \
your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
# Add a domain # Add a domain
curl -X POST http://localhost:8000/api/domain \ curl -X POST http://localhost:8000/api/domain \
-1
View File
@@ -1 +0,0 @@
2026.08.1
+1 -1
View File
@@ -32,7 +32,7 @@ docker run -d \
--network client-net \ --network client-net \
--restart unless-stopped \ --restart unless-stopped \
-v /var/log/coraza:/var/log/coraza \ -v /var/log/coraza:/var/log/coraza \
your-registry.example.com/cloud-hosting-platform/coraza-spoa:latest repo.anhonesthost.net/cloud-hosting-platform/coraza-spoa:latest
``` ```
Then on the `haproxy-manager` container, add the env var: Then on the `haproxy-manager` container, add the env var:
+1 -1
View File
@@ -100,7 +100,7 @@
<div class="owner"> <div class="owner">
<h2>Site owner?</h2> <h2>Site owner?</h2>
<p>If you operate this site and believe this block is incorrect, please contact your hosting provider's support team and include the request reference above. They can look up exactly which rule fired and adjust it if it's a false positive.</p> <p>If you operate this site and believe this block is incorrect, please <a href="https://secure.anhonesthost.com/submitticket.php" rel="noopener">open a support ticket</a> and include the request reference above. Our team can look up exactly which rule fired and adjust it if it's a false positive.</p>
</div> </div>
<p class="small">Reference IDs expire from our active logs after 14 days, so please open a ticket promptly if you'd like this investigated.</p> <p class="small">Reference IDs expire from our active logs after 14 days, so please open a ticket promptly if you'd like this investigated.</p>
+62 -419
View File
@@ -12,45 +12,12 @@ from datetime import datetime, timedelta
import json import json
import ipaddress import ipaddress
import shutil import shutil
import stat
import tempfile import tempfile
import threading import threading
import time import time
import re import re
import fcntl import fcntl
# ---------------------------------------------------------------------------
# Bounded subprocess execution (incident 2026-07-07)
# ---------------------------------------------------------------------------
# Every external command this manager runs — certbot ACME issuance/renewal,
# `socat` reloads over the haproxy admin socket, `haproxy -c` validation — is a
# potential hang. The management API runs under gunicorn gthread workers, and a
# subprocess.run() with NO timeout blocks its worker thread forever if the
# command stalls (e.g. an ACME/upstream that stops responding mid-read).
# gunicorn's --timeout does not rescue this: for gthread it only kills a worker
# whose *main* thread stops heart-beating, but the main thread keeps polling
# while pool threads are wedged. Enough stalled calls exhaust the 4-thread pool
# and the whole API stops responding — "healthy" health-check, every request
# 30s-timeouts — which is exactly what stalled WHP site updates on 2026-07-07.
#
# Fix: give EVERY subprocess.run() a default timeout unless the caller passes
# one explicitly. On expiry Python kills the child and raises
# subprocess.TimeoutExpired (a subclass of Exception); the existing per-endpoint
# try/except turns that into a clean error AND releases the worker thread.
# Bounding by default (instead of editing ~30 call sites) means no site can be
# missed and any future call is protected automatically.
DEFAULT_SUBPROCESS_TIMEOUT = int(os.environ.get('HAPROXY_MGR_SUBPROCESS_TIMEOUT', '180'))
_unbounded_subprocess_run = subprocess.run
def _bounded_subprocess_run(*args, **kwargs):
if kwargs.get('timeout') is None:
kwargs['timeout'] = DEFAULT_SUBPROCESS_TIMEOUT
return _unbounded_subprocess_run(*args, **kwargs)
subprocess.run = _bounded_subprocess_run
app = Flask(__name__) app = Flask(__name__)
# Default page server (port 8080) — served to HAProxy clients whose request hit # Default page server (port 8080) — served to HAProxy clients whose request hit
@@ -106,18 +73,8 @@ HAPROXY_CONFIG_PATH = '/etc/haproxy/haproxy.cfg'
HAPROXY_BACKUP_PATH = '/etc/haproxy/haproxy.cfg.backup' HAPROXY_BACKUP_PATH = '/etc/haproxy/haproxy.cfg.backup'
BLOCKED_IPS_MAP_PATH = '/etc/haproxy/blocked_ips.map' BLOCKED_IPS_MAP_PATH = '/etc/haproxy/blocked_ips.map'
BLOCKED_IPS_MAP_BACKUP_PATH = '/etc/haproxy/blocked_ips.map.backup' BLOCKED_IPS_MAP_BACKUP_PATH = '/etc/haproxy/blocked_ips.map.backup'
# Coraza SPOE engine file. `haproxy -c` parses this too (the frontend's
# `filter spoe engine coraza config <path>` line points at it), so it is part
# of the same restorable config set as haproxy.cfg — rolling back haproxy.cfg
# while leaving a broken coraza-spoe.cfg behind still fails validation.
CORAZA_SPOE_CONFIG_PATH = '/etc/haproxy/coraza-spoe.cfg'
CORAZA_SPOE_BACKUP_PATH = '/etc/haproxy/coraza-spoe.cfg.backup'
HAPROXY_SOCKET_PATH = '/var/run/haproxy.sock' HAPROXY_SOCKET_PATH = '/var/run/haproxy.sock'
SSL_CERTS_DIR = '/etc/haproxy/certs' SSL_CERTS_DIR = '/etc/haproxy/certs'
# Stable per-host secret for QUIC Retry/address-validation tokens. Lives in the
# /etc/haproxy named volume so it survives container recreates; self-healed on
# first config render. See get_or_create_cluster_secret().
CLUSTER_SECRET_PATH = '/etc/haproxy/cluster-secret'
API_KEY = os.environ.get('HAPROXY_API_KEY') # Optional API key for authentication API_KEY = os.environ.get('HAPROXY_API_KEY') # Optional API key for authentication
# Setup logging # Setup logging
@@ -850,12 +807,10 @@ def renew_certificates():
# Defensive: clear any stale lock left by a SIGKILLed prior run. # Defensive: clear any stale lock left by a SIGKILLed prior run.
clear_stale_certbot_locks() clear_stale_certbot_locks()
# Run certbot renew. Explicit long timeout (overrides the module # Run certbot renew
# default): `renew` walks every lineage and can legitimately make many
# ACME round-trips when several certs are actually due.
result = subprocess.run([ result = subprocess.run([
'certbot', 'renew', '--quiet' 'certbot', 'renew', '--quiet'
], capture_output=True, text=True, timeout=900) ], capture_output=True, text=True)
if result.returncode == 0: if result.returncode == 0:
# Check if any certificates were renewed # Check if any certificates were renewed
@@ -1732,45 +1687,6 @@ def dns_challenge_verify():
log_operation('dns_challenge_verify', False, str(e)) log_operation('dns_challenge_verify', False, str(e))
return jsonify({'success': False, 'error': str(e)}), 500 return jsonify({'success': False, 'error': str(e)}), 500
def get_or_create_cluster_secret():
"""Return a stable secret for QUIC token derivation, generating it once.
HAProxy uses `cluster-secret` to key QUIC Retry/address-validation tokens.
Without a stable value it picks a random one each (re)start and logs a
notice; tokens then don't survive reloads. We persist one in the
/etc/haproxy named volume so it's stable across container recreates.
Exclusive-create avoids a race if two renders run concurrently. Failure to
read/write is non-fatal: we fall back to an empty string and the template
simply omits the directive (HAProxy reverts to its random-per-process
behaviour), so QUIC still works.
"""
try:
if os.path.exists(CLUSTER_SECRET_PATH):
with open(CLUSTER_SECRET_PATH, 'r') as f:
secret = f.read().strip()
if secret:
return secret
# Generate and persist exclusively (0600). hex => config-safe charset.
secret = os.urandom(32).hex()
fd = os.open(CLUSTER_SECRET_PATH, os.O_CREAT | os.O_EXCL | os.O_WRONLY, 0o600)
try:
os.write(fd, secret.encode())
finally:
os.close(fd)
logger.info("Generated new QUIC cluster-secret at %s", CLUSTER_SECRET_PATH)
return secret
except FileExistsError:
# Lost the create race — another render just wrote it; read it back.
try:
with open(CLUSTER_SECRET_PATH, 'r') as f:
return f.read().strip()
except Exception as e:
logger.error("Failed to read cluster-secret after race: %s", e)
return ''
except Exception as e:
logger.error("Failed to get/create cluster-secret: %s", e)
return ''
def generate_config(): def generate_config():
try: try:
conn = sqlite3.connect(DB_FILE) conn = sqlite3.connect(DB_FILE)
@@ -1802,21 +1718,6 @@ def generate_config():
config_parts = [] config_parts = []
# Snapshot the last-known-good config BEFORE anything below touches a
# file in /etc/haproxy. Everything this function writes (haproxy.cfg,
# blocked_ips.map, coraza-spoe.cfg) is validated as one set by
# `haproxy -c`, so the rollback point has to predate the first of them.
# Taking it here (rather than inside reload_haproxy_safely(), which runs
# after the writes) is what makes rollback real - see create_backup().
backup_ok, backup_status = create_backup()
if not backup_ok:
# Could not even attempt a snapshot (I/O error). Writing a new
# config now would leave us with no way back, so refuse.
raise Exception(
"Refusing to regenerate config: failed to back up the current "
"configuration, so a failed change could not be rolled back"
)
# Optional Coraza WAF integration. When HAPROXY_CORAZA_SPOE_BACKEND is # Optional Coraza WAF integration. When HAPROXY_CORAZA_SPOE_BACKEND is
# set on the haproxy-manager container, we render an extra TCP backend # set on the haproxy-manager container, we render an extra TCP backend
# pointing at a coraza-spoa sidecar AND inject a `filter spoe ...` line # pointing at a coraza-spoa sidecar AND inject a `filter spoe ...` line
@@ -1848,9 +1749,7 @@ def generate_config():
logger.error(f"Failed to create {suspended_list_path}: {e}") logger.error(f"Failed to create {suspended_list_path}: {e}")
# Add Haproxy Default Headers # Add Haproxy Default Headers
default_headers = template_env.get_template('hap_header.tpl').render( default_headers = template_env.get_template('hap_header.tpl').render()
cluster_secret = get_or_create_cluster_secret(),
)
config_parts.append(default_headers) config_parts.append(default_headers)
# Update blocked IPs map file first # Update blocked IPs map file first
@@ -1911,11 +1810,7 @@ def generate_config():
# First pass: exact domain ACLs (higher priority - evaluated first) # First pass: exact domain ACLs (higher priority - evaluated first)
for domain in exact_domains: for domain in exact_domains:
if not domain['backend_name']: if not domain['backend_name']:
# Expected for domains registered without a proxy backend (e.g. the logger.warning(f"Skipping domain {domain['domain']} - no backend name")
# panel's own hostname, present only for certificate management).
# Log at INFO — not WARNING — so it doesn't trip log monitors as an
# error; it recurs on every generate_config by design.
logger.info(f"Skipping domain {domain['domain']} - no proxy backend (cert/management-only)")
continue continue
try: try:
@@ -1934,8 +1829,7 @@ def generate_config():
# Second pass: wildcard domain ACLs (lower priority - evaluated after exact matches) # Second pass: wildcard domain ACLs (lower priority - evaluated after exact matches)
for domain in wildcard_domains: for domain in wildcard_domains:
if not domain['backend_name']: if not domain['backend_name']:
# See note above — INFO, not WARNING; expected for cert/management-only domains. logger.warning(f"Skipping wildcard domain {domain['domain']} - no backend name")
logger.info(f"Skipping wildcard domain {domain['domain']} - no proxy backend (cert/management-only)")
continue continue
try: try:
@@ -2006,21 +1900,25 @@ backend default-backend
# how the file was authored. # how the file was authored.
if not coraza_spoe_cfg.endswith('\n'): if not coraza_spoe_cfg.endswith('\n'):
coraza_spoe_cfg += '\n' coraza_spoe_cfg += '\n'
write_config_atomically(CORAZA_SPOE_CONFIG_PATH, coraza_spoe_cfg) coraza_spoe_path = '/etc/haproxy/coraza-spoe.cfg'
logger.info(f"Coraza SPOE engine config written to " with open(coraza_spoe_path, 'w') as f:
f"{CORAZA_SPOE_CONFIG_PATH} " f.write(coraza_spoe_cfg)
logger.info(f"Coraza SPOE engine config written to {coraza_spoe_path} "
f"(SPOA target: {coraza_spoe_backend})") f"(SPOA target: {coraza_spoe_backend})")
# Write complete configuration to tmp
temp_config_path = "/etc/haproxy/haproxy.cfg"
config_content = '\n'.join(config_parts) config_content = '\n'.join(config_parts)
logger.debug("Generated HAProxy configuration") logger.debug("Generated HAProxy configuration")
# Write new configuration to file (atomically - a truncated haproxy.cfg # Write complete configuration to tmp
# is as fatal as an invalid one). The rollback point was taken above, # Write new configuration to file
# before this write. with open(HAPROXY_CONFIG_PATH, 'w') as f:
write_config_atomically(HAPROXY_CONFIG_PATH, config_content) f.write(config_content)
# Use safe reload with validation and rollback # Use safe reload with validation and rollback
success, message = reload_haproxy_safely(backup_status=backup_status) success, message = reload_haproxy_safely()
if success: if success:
logger.info("Configuration generated and HAProxy reloaded safely") logger.info("Configuration generated and HAProxy reloaded safely")
log_operation('generate_config', True, 'Configuration generated and HAProxy reloaded safely') log_operation('generate_config', True, 'Configuration generated and HAProxy reloaded safely')
@@ -2037,304 +1935,61 @@ backend default-backend
traceback.print_exc() traceback.print_exc()
raise raise
# --------------------------------------------------------------------------- def create_backup():
# Config backup / rollback """Create backup of current config and map files"""
# ---------------------------------------------------------------------------
# Rollback only works if the backup predates the write it is supposed to undo.
# Until 2026-08 create_backup() ran from inside reload_haproxy_safely(), i.e.
# AFTER generate_config() had already overwritten haproxy.cfg — so the "backup"
# was a copy of the new (possibly broken) config and restore_backup() restored
# the same broken bytes. The advertised rollback was a no-op and a fatal
# haproxy.cfg persisted on disk, where start_haproxy() refuses to launch (the
# June 2026 missing-template incident). create_backup() must now be called by
# the writer, BEFORE the first byte is written.
# Statuses returned by create_backup() that mean a rollback target exists.
_ROLLBACK_AVAILABLE_STATUSES = ('created', 'kept_previous')
def _files_identical(path_a, path_b):
"""Byte-compare two files.
Deliberately not filecmp.cmp(): it memoises on (size, mtime), and
shutil.copy2() preserves mtime, so a stale cache entry could report a
changed config as unchanged. These files are small; read them.
"""
try: try:
if os.path.getsize(path_a) != os.path.getsize(path_b): if os.path.exists(HAPROXY_CONFIG_PATH):
return False shutil.copy2(HAPROXY_CONFIG_PATH, HAPROXY_BACKUP_PATH)
with open(path_a, 'rb') as fa, open(path_b, 'rb') as fb: if os.path.exists(BLOCKED_IPS_MAP_PATH):
while True: shutil.copy2(BLOCKED_IPS_MAP_PATH, BLOCKED_IPS_MAP_BACKUP_PATH)
chunk_a = fa.read(65536) logger.info("Backups created successfully")
chunk_b = fb.read(65536)
if chunk_a != chunk_b:
return False
if not chunk_a:
return True
except OSError:
return False
def _config_set_matches_backup():
"""True if every live config file is byte-identical to its backup copy.
After a successful reload the live set has already been recorded as
known-good (see promote_current_config_to_backup()), which is the common
case at the start of the next generation. Recognising it lets create_backup()
skip both the re-validation and the copy - worth doing because
`haproxy -c` on an edge with hundreds of certificates is not free and
generate_config() runs synchronously inside customer-facing API calls.
"""
for live_path, backup_path in _config_backup_pairs():
if os.path.exists(live_path) != os.path.exists(backup_path):
return False
if (os.path.exists(live_path)
and not _files_identical(live_path, backup_path)):
return False
return True
def _config_backup_pairs():
"""(live, backup) pairs forming one restorable config set.
Built at call time rather than at import so the module-level path constants
stay patchable (tests, alternate deployments).
"""
return (
(HAPROXY_CONFIG_PATH, HAPROXY_BACKUP_PATH),
(BLOCKED_IPS_MAP_PATH, BLOCKED_IPS_MAP_BACKUP_PATH),
(CORAZA_SPOE_CONFIG_PATH, CORAZA_SPOE_BACKUP_PATH),
)
def write_config_atomically(path, content):
"""Write content to path via temp file + rename.
A half-written haproxy.cfg (disk full, container killed mid-write) is just
as fatal as an invalid one and is invisible to the caller. os.replace() is
atomic within a filesystem, so the file on disk is always either the whole
old config or the whole new one — never a truncated hybrid. This also keeps
the "existing config is already broken" case from being self-inflicted.
"""
directory = os.path.dirname(path) or '.'
# Preserve the mode of the file we are replacing; mkstemp defaults to 0600
# and HAProxy config files are conventionally 0644.
try:
mode = stat.S_IMODE(os.stat(path).st_mode)
except OSError:
mode = 0o644
fd, tmp_path = tempfile.mkstemp(
dir=directory, prefix=os.path.basename(path) + '.', suffix='.tmp'
)
try:
with os.fdopen(fd, 'w') as f:
f.write(content)
f.flush()
os.fsync(f.fileno())
os.chmod(tmp_path, mode)
os.replace(tmp_path, path)
except Exception:
try:
os.unlink(tmp_path)
except OSError:
pass
raise
def create_backup(require_valid=True):
"""Snapshot the CURRENT on-disk config set as the rollback point.
MUST be called BEFORE the new configuration is written — see the module
comment above. Calling it afterwards silently disarms rollback.
require_valid=True (default) refuses to promote a config that HAProxy
already rejects. Backing up a broken config would make "rollback" mean
"restore a different broken config"; keeping the older, validated backup
instead means a rollback always lands on something HAProxy will actually
start with. Cost is one `haproxy -c` run per config generation.
Returns (ok, status):
ok=False, status='error' - the copy itself failed; caller decides.
status='created' - backup now holds the current config.
status='kept_previous' - current config missing or invalid; the
existing (older, good) backup was kept.
status='unavailable' - nothing to roll back to at all (first
run, or broken config and no prior
backup). Rollback is NOT possible.
"""
try:
snapshot_ok = True
reason = None
if not os.path.exists(HAPROXY_CONFIG_PATH):
snapshot_ok = False
reason = 'no existing HAProxy config on disk (first run?)'
elif _config_set_matches_backup():
# The backup already IS the current config, recorded when it last
# loaded successfully. Nothing to copy and nothing to re-validate.
logger.debug("Config backup already matches the live config")
return True, 'created'
elif require_valid:
status, msg = validate_config_file(HAPROXY_CONFIG_PATH)
if status == 'invalid':
snapshot_ok = False
reason = f'current config on disk does not validate: {msg}'
elif status == 'unavailable':
# The validator itself could not run (no haproxy binary, etc).
# That is NOT evidence the config is bad, and refusing to back
# up would leave us with no rollback target at all, so fall
# back to last-written semantics and say so loudly.
logger.warning(
f"Could not verify current config before backup ({msg}); "
"backing it up unverified"
)
if not snapshot_ok:
if os.path.exists(HAPROXY_BACKUP_PATH):
logger.warning(
f"Not refreshing config backup: {reason}. Keeping the "
f"existing backup at {HAPROXY_BACKUP_PATH} as the rollback "
"target."
)
return True, 'kept_previous'
logger.error(
f"No config backup could be taken: {reason}, and no previous "
f"backup exists at {HAPROXY_BACKUP_PATH}. ROLLBACK IS NOT "
"AVAILABLE for this configuration change."
)
return True, 'unavailable'
for live_path, backup_path in _config_backup_pairs():
if os.path.exists(live_path):
shutil.copy2(live_path, backup_path)
logger.info("Backup of last-known-good config created successfully")
return True, 'created'
except Exception as e:
logger.error(f"Failed to create backup: {e}")
return False, 'error'
def promote_current_config_to_backup():
"""Record the live config as the known-good rollback target.
Called ONLY after the config has both validated and been loaded by HAProxy,
so "backup" really means "the last configuration this box was running".
Must never be called before a reload attempt: doing so would make the
backup a copy of the config we may still have to roll back from - the same
class of bug as backing up after the write.
Without this, a box whose very first generation succeeded has no rollback
target at all until its second successful generation, and any corruption of
haproxy.cfg in between leaves nothing to recover to.
"""
try:
for live_path, backup_path in _config_backup_pairs():
if os.path.exists(live_path):
shutil.copy2(live_path, backup_path)
logger.debug("Known-good config backup updated after successful reload")
return True return True
except Exception as e: except Exception as e:
# Non-fatal: the config is live and working, we just failed to record logger.error(f"Failed to create backup: {e}")
# it. Loud, because the next change now has a staler rollback target.
logger.error(f"Failed to record known-good config backup: {e}")
return False return False
def restore_backup(): def restore_backup():
"""Restore the backed-up config set over the live files. """Restore from backup files"""
Returns (restored, message). restored=False means NOTHING was rolled back
and the live config is still whatever the failed change left on disk —
callers MUST surface that difference, it is the difference between "we
recovered" and "this edge is sitting on a config HAProxy will not load".
"""
if not os.path.exists(HAPROXY_BACKUP_PATH):
msg = (f"No config backup at {HAPROXY_BACKUP_PATH} - cannot roll back; "
f"{HAPROXY_CONFIG_PATH} still holds the failed configuration")
logger.critical(msg)
return False, msg
try: try:
for live_path, backup_path in _config_backup_pairs(): if os.path.exists(HAPROXY_BACKUP_PATH):
if os.path.exists(backup_path): shutil.copy2(HAPROXY_BACKUP_PATH, HAPROXY_CONFIG_PATH)
shutil.copy2(backup_path, live_path) if os.path.exists(BLOCKED_IPS_MAP_BACKUP_PATH):
msg = f"Configuration restored from backup ({HAPROXY_BACKUP_PATH})" shutil.copy2(BLOCKED_IPS_MAP_BACKUP_PATH, BLOCKED_IPS_MAP_PATH)
logger.info(msg) logger.info("Backups restored successfully")
return True, msg return True
except Exception as e: except Exception as e:
msg = (f"Failed to restore backup: {e} - {HAPROXY_CONFIG_PATH} may hold " logger.error(f"Failed to restore backup: {e}")
"a broken configuration") return False
logger.critical(msg)
return False, msg
def validate_config_file(config_path):
"""Run `haproxy -c` against config_path.
Returns (status, message) with status one of:
'valid' - HAProxy parsed the file successfully
'invalid' - HAProxy rejected it (message carries stderr)
'unavailable' - the validator could not be run at all (binary missing,
timeout, ...). Deliberately distinct from 'invalid':
it tells us nothing about the config.
"""
try:
result = subprocess.run(['haproxy', '-c', '-f', config_path],
capture_output=True, text=True)
except Exception as e:
return 'unavailable', f"Error validating HAProxy config: {e}"
if result.returncode == 0:
return 'valid', None
return 'invalid', f"HAProxy configuration validation failed: {result.stderr}"
def validate_haproxy_config(): def validate_haproxy_config():
"""Validate the live HAProxy configuration file. Returns (is_valid, error).""" """Validate HAProxy configuration file"""
status, message = validate_config_file(HAPROXY_CONFIG_PATH)
if status == 'valid':
logger.info("HAProxy configuration validation passed")
return True, None
logger.error(message)
return False, message
def reload_haproxy_safely(backup_status=None):
"""Safely reload HAProxy with validation and rollback.
PRECONDITION: the caller must already have called create_backup() BEFORE
writing the new config, and pass the status it returned. This function runs
after the new config is on disk, so it cannot take a meaningful backup
itself — doing so is exactly the bug this contract exists to prevent.
backup_status=None means the caller did not take a pre-write backup. We do
NOT create one here (that would overwrite a genuinely good backup with the
unverified new config); we log it and fall back to whatever backup already
exists on disk.
"""
try: try:
if backup_status is None: result = subprocess.run(['haproxy', '-c', '-f', HAPROXY_CONFIG_PATH],
logger.error( capture_output=True, text=True)
"reload_haproxy_safely() called without a pre-write backup " if result.returncode == 0:
"status - rollback will fall back to whatever backup already " logger.info("HAProxy configuration validation passed")
"exists on disk. Callers must call create_backup() BEFORE " return True, None
"writing the new configuration." else:
) error_msg = f"HAProxy configuration validation failed: {result.stderr}"
elif backup_status not in _ROLLBACK_AVAILABLE_STATUSES: logger.error(error_msg)
logger.warning( return False, error_msg
f"Proceeding with reload without a rollback target " except Exception as e:
f"(backup status: {backup_status})" error_msg = f"Error validating HAProxy config: {e}"
) logger.error(error_msg)
return False, error_msg
def reload_haproxy_safely():
"""Safely reload HAProxy with validation and rollback"""
try:
# Create backup before changes
if not create_backup():
return False, "Failed to create backup"
# Validate new configuration # Validate new configuration
is_valid, error_msg = validate_haproxy_config() is_valid, error_msg = validate_haproxy_config()
if not is_valid: if not is_valid:
# Restore backup on validation failure # Restore backup on validation failure
restored, restore_msg = restore_backup() restore_backup()
if not restored:
logger.critical(
"Config validation failed AND rollback was not possible - "
f"{HAPROXY_CONFIG_PATH} holds an invalid configuration that "
"HAProxy will refuse to start with"
)
return False, (f"Config validation failed: {error_msg} | "
f"ROLLBACK FAILED: {restore_msg}")
return False, f"Config validation failed: {error_msg}" return False, f"Config validation failed: {error_msg}"
# Attempt reload # Attempt reload
@@ -2355,28 +2010,20 @@ def reload_haproxy_safely(backup_status=None):
if reload_result.returncode == 0: if reload_result.returncode == 0:
logger.info("HAProxy reloaded successfully") logger.info("HAProxy reloaded successfully")
# Now - and only now - is this config known good.
promote_current_config_to_backup()
return True, "HAProxy reloaded successfully" return True, "HAProxy reloaded successfully"
else: else:
# Reload failed, restore backup # Reload failed, restore backup
restored, restore_msg = restore_backup() restore_backup()
if restored: # Try to reload with backup config
# Try to reload with the restored (known-good) config subprocess.run('echo "reload" | socat stdio /tmp/haproxy-cli',
subprocess.run( shell=True, capture_output=True)
'echo "reload" | socat stdio /tmp/haproxy-cli',
shell=True, capture_output=True)
error_msg = f"HAProxy reload failed: {reload_result.stderr}" error_msg = f"HAProxy reload failed: {reload_result.stderr}"
if not restored:
error_msg += f" | ROLLBACK FAILED: {restore_msg}"
logger.error(error_msg) logger.error(error_msg)
return False, error_msg return False, error_msg
except Exception as e: except Exception as e:
# Critical error during reload, restore backup # Critical error during reload, restore backup
restored, restore_msg = restore_backup() restore_backup()
error_msg = f"Critical error during reload: {e}" error_msg = f"Critical error during reload: {e}"
if not restored:
error_msg += f" | ROLLBACK FAILED: {restore_msg}"
logger.error(error_msg) logger.error(error_msg)
return False, error_msg return False, error_msg
else: else:
@@ -2387,15 +2034,11 @@ def reload_haproxy_safely(backup_status=None):
check=True, capture_output=True, text=True check=True, capture_output=True, text=True
) )
logger.info("HAProxy started successfully") logger.info("HAProxy started successfully")
# Now - and only now - is this config known good.
promote_current_config_to_backup()
return True, "HAProxy started successfully" return True, "HAProxy started successfully"
except subprocess.CalledProcessError as e: except subprocess.CalledProcessError as e:
# Start failed, restore backup # Start failed, restore backup
restored, restore_msg = restore_backup() restore_backup()
error_msg = f"Failed to start HAProxy: {e.stderr}" error_msg = f"Failed to start HAProxy: {e.stderr}"
if not restored:
error_msg += f" | ROLLBACK FAILED: {restore_msg}"
logger.error(error_msg) logger.error(error_msg)
return False, error_msg return False, error_msg
except Exception as e: except Exception as e:
-52
View File
@@ -1,52 +0,0 @@
#!/usr/bin/env python3
"""Idempotent haproxy liveness check — driven by the in-container supervisor loop.
Why this exists
---------------
haproxy runs as a *background child of PID 1* (gunicorn) — it is started once at
container init (scripts/init.py -> do_initial_setup -> start_haproxy) and then
left running. Nothing supervises it after that. If the haproxy master process
dies mid-life (SIGABRT -> exit 134, segfault, or an OOM of the haproxy master),
the container stays "up" because gunicorn is still PID 1, so Docker's
`--restart` policy never fires. haproxy then stays down until the *external*
host watchdog (haproxy-watchdog.sh) notices port 80 is dead for ~3 minutes and
does a full `docker restart` — which drops every in-flight connection.
This script closes that gap: called on a short interval by the supervisor loop
in start-up.sh, it re-launches haproxy *in place* within one interval.
Safety
------
start_haproxy() is guarded by `is_process_running('haproxy')` (psutil-based, so
it works in this container which has no `ps`), so calling this while haproxy is
healthy is a cheap no-op. It only ever acts when haproxy is genuinely gone.
"""
import sys
sys.path.insert(0, '/haproxy')
import haproxy_manager # noqa: E402 (sys.path manipulation must come first)
def main():
if haproxy_manager.is_process_running('haproxy'):
return 0
haproxy_manager.logger.warning(
"[haproxy-supervisor] haproxy process not found — attempting in-place restart"
)
# start_haproxy() validates the config (and regenerates it if invalid)
# before launching, and swallows its own errors, so it will not raise here.
haproxy_manager.start_haproxy()
if haproxy_manager.is_process_running('haproxy'):
haproxy_manager.logger.info("[haproxy-supervisor] haproxy restarted in place")
return 0
haproxy_manager.logger.error(
"[haproxy-supervisor] haproxy restart FAILED — still not running after start_haproxy()"
)
return 1
if __name__ == '__main__':
sys.exit(main())
+20 -40
View File
@@ -6,40 +6,22 @@
SOCKET="/tmp/haproxy-cli" SOCKET="/tmp/haproxy-cli"
MAP_FILE="/etc/haproxy/blocked_ips.map" MAP_FILE="/etc/haproxy/blocked_ips.map"
# HAProxy runs in master-worker mode here, and /tmp/haproxy-cli is the MASTER
# socket. Data-plane commands (map/table manipulation) are NOT accepted on the
# master socket — they must be routed to a worker with the "@<n>" prefix. "@1"
# targets the current active worker. (A bare "add map ..." on the master socket
# fails with "Unknown command: 'add'".)
cli() { printf '@1 %s\n' "$*" | socat stdio "$SOCKET"; }
# Map lookup in haproxy.cfg is `map_ip(...,0) -m int gt 0`, so each entry MUST be
# "<ip_or_cidr> 1" — a bare IP yields an empty value (0) and is NOT blocked once
# the map file is re-read on reload. The runtime map and the file must agree.
MAP_VALUE=1
# Ensure map file exists # Ensure map file exists
if [ ! -f "$MAP_FILE" ]; then if [ ! -f "$MAP_FILE" ]; then
echo "# Blocked IPs - Format: <ip_or_cidr> 1 (one per line)" > "$MAP_FILE" touch "$MAP_FILE"
echo "# Blocked IPs - Format: IP_ADDRESS" > "$MAP_FILE"
fi fi
# Escape regex metacharacters (notably dots) in an IP/CIDR for anchored matching.
esc_re() { printf '%s' "$1" | sed 's/[.[\*^$/]/\\&/g'; }
case "$1" in case "$1" in
block) block)
if [ -z "$2" ]; then if [ -z "$2" ]; then
echo "Usage: $0 block IP_ADDRESS" echo "Usage: $0 block IP_ADDRESS"
exit 1 exit 1
fi fi
re="$(esc_re "$2")" # Add IP to map file
# Persist (idempotent, anchored so 1.2.3.4 doesn't match 1.2.3.45), grep -q "^$2" "$MAP_FILE" || echo "$2" >> "$MAP_FILE"
# always with the trailing value so the block survives a reload. # Add to runtime map
if ! grep -qE "^${re}([[:space:]]|$)" "$MAP_FILE"; then echo "add map /etc/haproxy/blocked_ips.map $2 1" | socat stdio "$SOCKET"
echo "$2 $MAP_VALUE" >> "$MAP_FILE"
fi
# Apply at runtime immediately (no reload).
cli "add map $MAP_FILE $2 $MAP_VALUE"
echo "Blocked IP: $2" echo "Blocked IP: $2"
;; ;;
@@ -48,33 +30,31 @@ case "$1" in
echo "Usage: $0 unblock IP_ADDRESS" echo "Usage: $0 unblock IP_ADDRESS"
exit 1 exit 1
fi fi
re="$(esc_re "$2")" # Remove from map file
# Remove from map file (match "<ip>" optionally followed by a value). sed -i "/^$2$/d" "$MAP_FILE"
sed -i -E "/^${re}([[:space:]]|$)/d" "$MAP_FILE" # Remove from runtime map
# Remove from runtime map. echo "del map /etc/haproxy/blocked_ips.map $2" | socat stdio "$SOCKET"
cli "del map $MAP_FILE $2"
echo "Unblocked IP: $2" echo "Unblocked IP: $2"
;; ;;
list) list)
echo "Currently blocked IPs:" echo "Currently blocked IPs:"
# `show map` output is "<ptr> <key> <value>" — the IP is field 2. echo "show map /etc/haproxy/blocked_ips.map" | socat stdio "$SOCKET" | awk '{print $1}'
cli "show map $MAP_FILE" | awk 'NF>=2 {print $2}'
;; ;;
clear) clear)
echo "Clearing all blocked IPs..." echo "Clearing all blocked IPs..."
cli "clear map $MAP_FILE" echo "clear map /etc/haproxy/blocked_ips.map" | socat stdio "$SOCKET"
echo "# Blocked IPs - Format: <ip_or_cidr> 1 (one per line)" > "$MAP_FILE" echo "# Blocked IPs - Format: IP_ADDRESS" > "$MAP_FILE"
echo "All IPs unblocked" echo "All IPs unblocked"
;; ;;
stats) stats)
echo "=== HAProxy 3.0.11 Threat Intelligence Dashboard ===" echo "=== HAProxy 3.0.11 Threat Intelligence Dashboard ==="
cli "show table web" | awk 'NR<=21' echo "show table web" | socat stdio "$SOCKET" | awk 'NR<=21'
echo "" echo ""
echo "=== Top Threat Scores ===" echo "=== Top Threat Scores ==="
cli "show table web" | awk ' echo "show table web" | socat stdio "$SOCKET" | awk '
NR>1 { NR>1 {
ip = $1 ip = $1
auth_fail = 0 auth_fail = 0
@@ -104,7 +84,7 @@ case "$1" in
exit 1 exit 1
fi fi
# Add to manual blacklist using GPC(13) # Add to manual blacklist using GPC(13)
cli "set table web key $2 data.gpc(13) 1" echo "set table web key $2 data.gpc(13) 1" | socat stdio "$SOCKET"
echo "Manually blacklisted IP: $2 (GPC(13) = 1)" echo "Manually blacklisted IP: $2 (GPC(13) = 1)"
;; ;;
@@ -114,7 +94,7 @@ case "$1" in
exit 1 exit 1
fi fi
# Clear manual blacklist flag # Clear manual blacklist flag
cli "set table web key $2 data.gpc(13) 0" echo "set table web key $2 data.gpc(13) 0" | socat stdio "$SOCKET"
echo "Removed manual blacklist for IP: $2" echo "Removed manual blacklist for IP: $2"
;; ;;
@@ -124,7 +104,7 @@ case "$1" in
exit 1 exit 1
fi fi
# Add to auto-blacklist using GPC(14) # Add to auto-blacklist using GPC(14)
cli "set table web key $2 data.gpc(14) 1" echo "set table web key $2 data.gpc(14) 1" | socat stdio "$SOCKET"
echo "Auto-blacklisted IP: $2 (GPC(14) = 1)" echo "Auto-blacklisted IP: $2 (GPC(14) = 1)"
;; ;;
@@ -135,14 +115,14 @@ case "$1" in
fi fi
# Show detailed threat breakdown for specific IP # Show detailed threat breakdown for specific IP
echo "Threat analysis for $2:" echo "Threat analysis for $2:"
cli "show table web key $2" echo "show table web key $2" | socat stdio "$SOCKET"
;; ;;
*) *)
echo "Usage: $0 {block|unblock|list|clear|blacklist|unblacklist|auto-blacklist|threat-score|stats} [IP_ADDRESS]" echo "Usage: $0 {block|unblock|list|clear|blacklist|unblacklist|auto-blacklist|threat-score|stats} [IP_ADDRESS]"
echo "" echo ""
echo "HAProxy 3.0.11 Enhanced Security Commands:" echo "HAProxy 3.0.11 Enhanced Security Commands:"
echo " block IP - Block IP via map file (immediate + persisted)" echo " block IP - Block IP via map file (immediate)"
echo " unblock IP - Unblock IP from map file" echo " unblock IP - Unblock IP from map file"
echo " blacklist IP - Manual blacklist via GPC(13) array" echo " blacklist IP - Manual blacklist via GPC(13) array"
echo " unblacklist IP - Remove manual blacklist flag" echo " unblacklist IP - Remove manual blacklist flag"
+1 -23
View File
@@ -27,33 +27,11 @@ cron &
# Phase 1: container init # Phase 1: container init
python /haproxy/scripts/init.py python /haproxy/scripts/init.py
# Phase 1.5: in-container haproxy supervisor.
# haproxy runs as a background child of PID 1 (gunicorn) with NOTHING watching
# it after init. If the haproxy master dies mid-life (e.g. SIGABRT -> exit 134,
# segfault), the container stays "up" (gunicorn is PID 1), Docker's --restart
# policy never fires, and haproxy is down until the external host watchdog
# full-restarts the whole container minutes later (dropping every connection).
# This loop revives haproxy in place within one interval. ensure_haproxy.py is
# idempotent — a cheap no-op whenever haproxy is already running.
HAPROXY_SUPERVISOR_INTERVAL="${HAPROXY_SUPERVISOR_INTERVAL:-15}"
(
while true; do
sleep "${HAPROXY_SUPERVISOR_INTERVAL}"
python /haproxy/scripts/ensure_haproxy.py 2>&1 || true
done
) &
# Phase 2: WSGI servers # Phase 2: WSGI servers
# Tunable via env: HAPROXY_MGR_API_WORKERS (default 1), HAPROXY_MGR_API_TIMEOUT # Tunable via env: HAPROXY_MGR_API_WORKERS (default 1), HAPROXY_MGR_API_TIMEOUT
# (default 120 — API can do slow ACME calls), HAPROXY_MGR_MAX_REQUESTS (default # (default 120 — API can do slow ACME calls), HAPROXY_MGR_MAX_REQUESTS (default
# 1000 — worker recycle frequency). # 1000 — worker recycle frequency).
# API_WORKERS="${HAPROXY_MGR_API_WORKERS:-1}"
# API_WORKERS default is 2 (was 1). A single worker is a single point of
# failure: if its gthread pool ever wedges (see the 2026-07-07 subprocess-hang
# incident — now bounded by DEFAULT_SUBPROCESS_TIMEOUT in haproxy_manager.py),
# the entire management API goes dark. A second worker keeps the API answering
# (config regenerate, health, SSL) while the other recycles via --max-requests.
API_WORKERS="${HAPROXY_MGR_API_WORKERS:-2}"
API_TIMEOUT="${HAPROXY_MGR_API_TIMEOUT:-120}" API_TIMEOUT="${HAPROXY_MGR_API_TIMEOUT:-120}"
MAX_REQ="${HAPROXY_MGR_MAX_REQUESTS:-1000}" MAX_REQ="${HAPROXY_MGR_MAX_REQUESTS:-1000}"
MAX_REQ_JITTER="${HAPROXY_MGR_MAX_REQUESTS_JITTER:-100}" MAX_REQ_JITTER="${HAPROXY_MGR_MAX_REQUESTS_JITTER:-100}"
-467
View File
@@ -1,467 +0,0 @@
#!/usr/bin/env python3
"""Regression tests for HAProxy config backup / rollback ordering.
Why this file exists
--------------------
generate_config() used to write the new haproxy.cfg and only THEN call
reload_haproxy_safely() -> create_backup(), so the "backup" was a copy of the
config that had just been written. On a validation failure restore_backup()
restored the identical broken bytes: the advertised rollback was a no-op and a
fatal haproxy.cfg stayed on disk, where start_haproxy() refuses to launch.
These tests pin the ordering invariant (backup predates the write) and the
observable end-to-end behaviour (after a failed validation the file on disk is
the previous working config and HAProxy will start with it).
Running
-------
python3 scripts/test-config-rollback.py # tests the repo checkout
HAPROXY_MANAGER_DIR=/some/other/tree \
python3 scripts/test-config-rollback.py # tests another tree
The repo has no Python test framework (scripts/test-*.sh are curl-based
integration scripts against a running API), so this is a self-contained
stdlib-unittest script - no pytest, no venv, no new dependencies beyond the
application's own requirements.txt (Flask/Jinja2/psutil), which are already
present in the container image.
No HAProxy binary is required: a stub `haproxy` is put on PATH that mimics
`haproxy -c -f <file>` by rejecting any config containing the token
__BROKEN__, which is how the tests inject an invalid configuration.
"""
import os
import sys
import shutil
import sqlite3
import logging
import tempfile
import textwrap
import unittest
BROKEN_TOKEN = '__BROKEN__'
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
# haproxy_manager builds its Jinja2 environment from the relative path
# Path('templates'), so it has to be imported with the module dir as cwd.
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
# The module opens /var/log/haproxy-manager.log at import time via
# logging.FileHandler. Redirect that one call so the suite runs unprivileged.
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda fn, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(fn)), *a, **kw)
)
try:
import haproxy_manager as hm
except ImportError as exc: # pragma: no cover - environment problem, not a failure
sys.stderr.write(
f"SKIP: cannot import haproxy_manager ({exc}).\n"
"Install the application requirements first: pip install -r requirements.txt\n"
)
raise SystemExit(77)
finally:
logging.FileHandler = _real_file_handler
logging.getLogger('haproxy_manager').setLevel(logging.CRITICAL)
FAKE_HAPROXY = textwrap.dedent(f"""\
#!/bin/sh
# Test stub for the haproxy binary.
# haproxy -c -f FILE -> exit 1 if FILE contains {BROKEN_TOKEN}, else 0
# haproxy -W -S ... -f FILE (start) -> same validation, then exit 0
cfg=""
while [ $# -gt 0 ]; do
case "$1" in -f) cfg="$2"; shift ;; esac
shift
done
if [ -n "$cfg" ] && grep -q '{BROKEN_TOKEN}' "$cfg" 2>/dev/null; then
echo "[ALERT] parsing [$cfg:1] : unknown keyword '{BROKEN_TOKEN}'" >&2
exit 1
fi
exit 0
""")
class RollbackTestCase(unittest.TestCase):
"""Base fixture: an isolated fake /etc/haproxy plus a stub haproxy binary."""
def setUp(self):
self.tmp = tempfile.mkdtemp(prefix='haproxy-rollback-test-')
self.addCleanup(shutil.rmtree, self.tmp, True)
bindir = os.path.join(self.tmp, 'bin')
os.makedirs(bindir)
stub = os.path.join(bindir, 'haproxy')
with open(stub, 'w') as fh:
fh.write(FAKE_HAPROXY)
os.chmod(stub, 0o755)
self._old_path = os.environ['PATH']
os.environ['PATH'] = bindir + os.pathsep + self._old_path
self.addCleanup(lambda: os.environ.__setitem__('PATH', self._old_path))
self.etc = os.path.join(self.tmp, 'etc')
os.makedirs(self.etc)
overrides = {
'DB_FILE': os.path.join(self.etc, 'haproxy_config.db'),
'HAPROXY_CONFIG_PATH': os.path.join(self.etc, 'haproxy.cfg'),
'HAPROXY_BACKUP_PATH': os.path.join(self.etc, 'haproxy.cfg.backup'),
'BLOCKED_IPS_MAP_PATH': os.path.join(self.etc, 'blocked_ips.map'),
'BLOCKED_IPS_MAP_BACKUP_PATH': os.path.join(self.etc, 'blocked_ips.map.backup'),
'CLUSTER_SECRET_PATH': os.path.join(self.etc, 'cluster-secret'),
'SSL_CERTS_DIR': os.path.join(self.etc, 'certs'),
'HAPROXY_SOCKET_PATH': os.path.join(self.etc, 'haproxy.sock'),
# Added by the rollback fix; older trees do not have it.
'CORAZA_SPOE_CONFIG_PATH': os.path.join(self.etc, 'coraza-spoe.cfg'),
'CORAZA_SPOE_BACKUP_PATH': os.path.join(self.etc, 'coraza-spoe.cfg.backup'),
}
self._saved = {}
for name, value in overrides.items():
self._saved[name] = getattr(hm, name, None)
setattr(hm, name, value)
self.addCleanup(self._restore_globals)
os.makedirs(hm.SSL_CERTS_DIR)
# log_operation() appends to a hardcoded /var/log path. Injecting `open`
# into the module namespace shadows the builtin for that module only
# (module globals are searched before builtins), so the real
# log_operation code still runs.
real_open = open
log_dir = self.tmp
def _redirecting_open(path, *args, **kwargs):
if isinstance(path, str) and path.startswith('/var/log/'):
path = os.path.join(log_dir, os.path.basename(path))
return real_open(path, *args, **kwargs)
hm.open = _redirecting_open
self.addCleanup(lambda: hm.__dict__.pop('open', None))
hm.init_db()
def _restore_globals(self):
for name, value in self._saved.items():
if value is None:
hm.__dict__.pop(name, None)
else:
setattr(hm, name, value)
# -- helpers ---------------------------------------------------------
def add_domain(self, domain, backend_name, address='10.0.0.1'):
with sqlite3.connect(hm.DB_FILE) as conn:
cur = conn.cursor()
cur.execute('INSERT INTO domains (domain, ssl_enabled) VALUES (?, 0)',
(domain,))
domain_id = cur.lastrowid
cur.execute('INSERT INTO backends (name, domain_id) VALUES (?, ?)',
(backend_name, domain_id))
backend_id = cur.lastrowid
cur.execute(
'INSERT INTO backend_servers '
'(backend_id, server_name, server_address, server_port) '
'VALUES (?, ?, ?, ?)',
(backend_id, 'srv1', address, 8080))
conn.commit()
def block_ip(self, ip):
with sqlite3.connect(hm.DB_FILE) as conn:
conn.execute('INSERT INTO blocked_ips (ip_address, reason) VALUES (?, ?)',
(ip, 'test'))
conn.commit()
def read(self, path):
with open(path) as fh:
return fh.read()
def config_is_loadable(self):
"""True if HAProxy would accept the config currently on disk."""
import subprocess
return subprocess.run(
['haproxy', '-c', '-f', hm.HAPROXY_CONFIG_PATH],
capture_output=True).returncode == 0
def generate_good_config(self):
self.add_domain('good.example.com', 'good_backend')
hm.generate_config()
self.assertTrue(self.config_is_loadable(),
'fixture precondition: first generated config must be valid')
return self.read(hm.HAPROXY_CONFIG_PATH)
def break_the_config(self):
"""Queue a domain whose rendered backend the validator rejects."""
self.add_domain('bad.example.com', BROKEN_TOKEN + '_backend', '10.0.0.2')
class TestBackupOrdering(RollbackTestCase):
def test_backup_is_taken_before_the_new_config_is_written(self):
"""The ordering invariant, asserted directly.
Whatever create_backup() sees on disk must be the OLD config; if the
write happens first the backup is a copy of the new config and rollback
is meaningless.
"""
good = self.generate_good_config()
seen = {}
real_create_backup = hm.create_backup
def spy(*args, **kwargs):
seen['config_on_disk'] = self.read(hm.HAPROXY_CONFIG_PATH)
return real_create_backup(*args, **kwargs)
hm.create_backup = spy
self.addCleanup(setattr, hm, 'create_backup', real_create_backup)
self.add_domain('second.example.com', 'second_backend', '10.0.0.3')
hm.generate_config()
self.assertIn('config_on_disk', seen,
'create_backup() was never called during generate_config()')
self.assertEqual(
seen['config_on_disk'], good,
'create_backup() ran AFTER the new config was written - the backup '
'is a copy of the new config, so rollback cannot undo anything')
def test_backup_tracks_the_last_known_good_config(self):
"""After a change that validated AND loaded, the backup is that config.
The rollback target is "the last configuration HAProxy actually ran",
not "the file that happened to be there last time".
"""
good = self.generate_good_config()
self.add_domain('second.example.com', 'second_backend', '10.0.0.3')
hm.generate_config()
live = self.read(hm.HAPROXY_CONFIG_PATH)
self.assertNotEqual(live, good, 'fixture sanity: the new config should differ')
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), live,
'the successful config was not recorded as known-good')
def test_backup_is_not_promoted_when_the_change_fails(self):
"""A config that never loaded must not become the rollback target."""
good = self.generate_good_config()
self.break_the_config()
with self.assertRaises(Exception):
hm.generate_config()
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
'a config that failed validation was promoted to backup')
class TestRollbackEndToEnd(RollbackTestCase):
def test_failed_validation_leaves_the_last_good_config_on_disk(self):
good = self.generate_good_config()
self.break_the_config()
with self.assertRaises(Exception):
hm.generate_config()
on_disk = self.read(hm.HAPROXY_CONFIG_PATH)
self.assertNotIn(BROKEN_TOKEN, on_disk,
'the rejected config is still on disk - rollback was a no-op')
self.assertEqual(on_disk, good,
'on-disk config is not byte-identical to the last good one')
def test_haproxy_would_still_start_after_a_failed_change(self):
"""The operational consequence: the edge can still come up."""
self.generate_good_config()
self.break_the_config()
with self.assertRaises(Exception):
hm.generate_config()
self.assertTrue(self.config_is_loadable(),
'HAProxy would refuse to start with the config left on disk')
with self.assertLogs('haproxy_manager', level='INFO') as captured:
hm.start_haproxy()
self.assertTrue(
any('HAProxy started successfully' in line for line in captured.output),
f'start_haproxy() did not succeed after rollback: {captured.output}')
def test_blocked_ips_map_is_rolled_back_too(self):
"""generate_config() rewrites the map file before writing haproxy.cfg."""
self.block_ip('192.0.2.10')
self.generate_good_config()
good_map = self.read(hm.BLOCKED_IPS_MAP_PATH)
self.block_ip('198.51.100.20')
self.break_the_config()
with self.assertRaises(Exception):
hm.generate_config()
self.assertEqual(self.read(hm.BLOCKED_IPS_MAP_PATH), good_map,
'blocked IPs map was not rolled back with the config')
def test_first_run_failure_reports_that_rollback_was_impossible(self):
"""No prior config: there is nothing to restore, and that must be said.
A missing backup must never be reported as a successful restore, and it
must never be turned into "restore an empty file".
"""
self.break_the_config()
with self.assertRaises(Exception) as ctx:
hm.generate_config()
self.assertIn('ROLLBACK FAILED', str(ctx.exception),
'a failed change with no backup was not reported as such')
self.assertFalse(os.path.exists(hm.HAPROXY_BACKUP_PATH),
'a backup was fabricated from the broken config')
# The broken config is deliberately left in place: start_haproxy() can
# then detect it and try to regenerate. It must not be blanked.
self.assertGreater(os.path.getsize(hm.HAPROXY_CONFIG_PATH), 0,
'config file was emptied instead of left for diagnosis')
class TestBackupPrimitives(RollbackTestCase):
def test_restore_backup_distinguishes_missing_backup_from_success(self):
restored, message = hm.restore_backup()
self.assertFalse(restored,
'restore_backup() reported success with no backup present')
self.assertIn('cannot roll back', message.lower())
good = self.generate_good_config()
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
fh.write('scribbled over\n')
restored, message = hm.restore_backup()
self.assertTrue(restored, message)
self.assertEqual(self.read(hm.HAPROXY_CONFIG_PATH), good)
def test_a_successful_generation_records_a_rollback_target(self):
"""Even the first-ever generation must leave something to roll back to."""
good = self.generate_good_config()
self.assertTrue(
os.path.exists(hm.HAPROXY_BACKUP_PATH),
'after a successful reload there is still no known-good backup')
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good)
def test_a_broken_current_config_does_not_replace_a_good_backup(self):
"""The known-good marker.
If the config already on disk is broken (previous failed write, manual
edit), snapshotting it would make "rollback" mean "restore a different
broken config". The older validated backup must survive.
"""
good = self.generate_good_config()
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
'fixture: a good backup should exist by now')
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
fh.write(f'garbage {BROKEN_TOKEN} config\n')
ok, status = hm.create_backup()
self.assertTrue(ok)
self.assertEqual(status, 'kept_previous')
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
'a broken config overwrote the known-good backup')
def test_reload_does_not_take_its_own_backup(self):
"""reload_haproxy_safely() runs after the write, so it must not back up."""
good = self.generate_good_config()
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
fh.write(f'broken {BROKEN_TOKEN}\n')
success, message = hm.reload_haproxy_safely(backup_status='created')
self.assertFalse(success)
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
'reload_haproxy_safely() overwrote the good backup')
self.assertEqual(self.read(hm.HAPROXY_CONFIG_PATH), good,
'reload_haproxy_safely() did not roll the config back')
def test_unchanged_config_is_not_revalidated(self):
"""Fast path: if the backup already is the live config, do no work.
generate_config() runs inside customer-facing API calls and
`haproxy -c` is expensive on an edge with hundreds of certificates.
"""
self.generate_good_config()
calls = []
real_validate = hm.validate_config_file
hm.validate_config_file = lambda path: (calls.append(path),
real_validate(path))[1]
self.addCleanup(setattr, hm, 'validate_config_file', real_validate)
ok, status = hm.create_backup()
self.assertTrue(ok)
self.assertEqual(status, 'created')
self.assertEqual(calls, [],
'the unchanged live config was re-validated needlessly')
def test_fast_path_does_not_hide_a_drifted_broken_config(self):
"""If the live config drifted from the backup, the gate must still run."""
good = self.generate_good_config()
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
fh.write(f'hand edited {BROKEN_TOKEN}\n')
ok, status = hm.create_backup()
self.assertTrue(ok)
self.assertEqual(status, 'kept_previous',
'a drifted broken config was silently accepted')
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good)
def test_backup_set_covers_every_file_generate_config_writes(self):
pairs = dict(hm._config_backup_pairs())
for path in (hm.HAPROXY_CONFIG_PATH, hm.BLOCKED_IPS_MAP_PATH,
hm.CORAZA_SPOE_CONFIG_PATH):
self.assertIn(path, pairs,
f'{path} is written by generate_config() but is not '
'part of the backed-up config set')
def test_coraza_spoe_config_round_trips(self):
self.generate_good_config()
with open(hm.CORAZA_SPOE_CONFIG_PATH, 'w') as fh:
fh.write('spoe-good\n')
hm.create_backup()
with open(hm.CORAZA_SPOE_CONFIG_PATH, 'w') as fh:
fh.write('spoe-broken\n')
restored, message = hm.restore_backup()
self.assertTrue(restored, message)
self.assertEqual(self.read(hm.CORAZA_SPOE_CONFIG_PATH), 'spoe-good\n')
class TestAtomicWrite(RollbackTestCase):
def test_write_is_atomic_and_preserves_mode(self):
path = os.path.join(self.etc, 'atomic.cfg')
with open(path, 'w') as fh:
fh.write('old')
os.chmod(path, 0o644)
hm.write_config_atomically(path, 'new content\n')
self.assertEqual(self.read(path), 'new content\n')
self.assertEqual(oct(os.stat(path).st_mode & 0o777), oct(0o644))
leftovers = [n for n in os.listdir(self.etc) if n.endswith('.tmp')]
self.assertEqual(leftovers, [], f'temp files left behind: {leftovers}')
def test_failed_write_leaves_the_previous_file_intact(self):
path = os.path.join(self.etc, 'atomic.cfg')
with open(path, 'w') as fh:
fh.write('old content\n')
# Anything that makes f.write() blow up mid-flight stands in for a full
# disk / killed container.
with self.assertRaises(Exception):
hm.write_config_atomically(path, object())
self.assertEqual(self.read(path), 'old content\n',
'a failed write clobbered the previous config')
leftovers = [n for n in os.listdir(self.etc) if n.endswith('.tmp')]
self.assertEqual(leftovers, [], f'temp files left behind: {leftovers}')
if __name__ == '__main__':
print(f"testing haproxy_manager from: {MODULE_DIR}")
unittest.main(verbosity=2)
-41
View File
@@ -1,41 +0,0 @@
# Long-lived backend for {{ name }} (template_override='hap_backend_longlived').
# Use for apps whose PRIMARY traffic holds connections open: media streaming,
# large up/downloads, or persistent viewer/streaming sessions. Both the primary
# and the SSE backend are tuned long-lived here (no http-server-close,
# http-no-delay, 6h server/tunnel/keep-alive timeouts).
#
# Compare hap_backend_websocket.tpl, which keeps the PRIMARY backend standard
# and only makes the -sse-backend long-lived. Pick this one when the main path
# itself needs long-lived connections, not just an SSE side-channel.
backend {{ name }}-backend
no option http-server-close
option http-no-delay
timeout server 6h
timeout tunnel 6h
timeout http-keep-alive 6h
option forwardfor
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
http-request set-header X-Real-IP %[var(txn.real_ip)]
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
http-request set-header X-Forwarded-Proto https if { ssl_fc }
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
{% for server in servers %}
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
{% endfor %}
# SSE variant (Accept: text/event-stream / ?action=stream auto-routes here)
backend {{ name }}-sse-backend
no option http-server-close
option http-no-delay
timeout server 6h
timeout tunnel 6h
timeout http-keep-alive 6h
option forwardfor
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
http-request set-header X-Real-IP %[var(txn.real_ip)]
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
http-request set-header X-Forwarded-Proto https if { ssl_fc }
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
{% for server in servers %}
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
{% endfor %}
-34
View File
@@ -1,34 +0,0 @@
# Long-lived / websocket-safe backend for {{ name }} (template_override)
# For apps with persistent WebSocket/streaming connections (e.g. Jitsi /xmpp-websocket, /colibri-ws).
backend {{ name }}-backend
no option http-server-close
option http-no-delay
timeout server 6h
timeout tunnel 6h
timeout http-keep-alive 6h
option forwardfor
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
http-request set-header X-Real-IP %[var(txn.real_ip)]
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
http-request set-header X-Forwarded-Proto https if { ssl_fc }
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
{% for server in servers %}
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
{% endfor %}
# SSE variant (Accept: text/event-stream / ?action=stream auto-routes here)
backend {{ name }}-sse-backend
no option http-server-close
option http-no-delay
timeout server 6h
timeout tunnel 6h
timeout http-keep-alive 6h
option forwardfor
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
http-request set-header X-Real-IP %[var(txn.real_ip)]
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
http-request set-header X-Forwarded-Proto https if { ssl_fc }
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
{% for server in servers %}
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
{% endfor %}
-17
View File
@@ -27,23 +27,6 @@ global
# SSL and Performance # SSL and Performance
tune.ssl.default-dh-param 2048 tune.ssl.default-dh-param 2048
# HTTP/3 over QUIC. The Debian haproxy package is built against system
# OpenSSL via the compatibility shim (USE_QUIC_OPENSSL_COMPAT), which is
# not a native QUIC TLS stack. HAProxy therefore rejects `quic*@` binds
# unless this opt-in is set. `limited-quic` enables QUIC through the compat
# layer (no 0-RTT — that needs quictls/aws-lc or native OpenSSL 3.5 QUIC).
# Without this, the quic bind in the frontend fails to start: "this SSL
# library does not support the QUIC protocol".
limited-quic
{%- if cluster_secret %}
# Stable secret keying QUIC Retry/address-validation tokens. Self-healed
# to /etc/haproxy/cluster-secret (named volume) by the manager so it
# survives recreates; without it haproxy picks a random one per process
# and tokens don't survive reloads (benign, just a startup notice).
cluster-secret "{{ cluster_secret }}"
{%- endif %}
# HTTP/2 protection against Rapid Reset (CVE-2023-44487) and stream abuse # HTTP/2 protection against Rapid Reset (CVE-2023-44487) and stream abuse
tune.h2.fe.max-total-streams 2000 tune.h2.fe.max-total-streams 2000
tune.h2.fe.glitches-threshold 50 tune.h2.fe.glitches-threshold 50
-76
View File
@@ -4,21 +4,6 @@ frontend web
# crt can now be a path, so it will load all .pem files in the path # crt can now be a path, so it will load all .pem files in the path
bind 0.0.0.0:443 ssl crt {{ crt_path }} alpn h2,http/1.1 bind 0.0.0.0:443 ssl crt {{ crt_path }} alpn h2,http/1.1
# HTTP/3 over QUIC (UDP/443). Same cert path as the TCP listener above.
# The Debian haproxy package is built +QUIC (QUIC_OPENSSL_COMPAT), so this
# is config-only — no source build. Requires UDP/443 published on the
# container (`-p 443:443/udp`) and open at the host firewall. `h3` is the
# only ALPN QUIC negotiates; h2/http1 stay on the TCP bind above. Sharing
# the frontend means all the real-IP, rate-limit, IP-block and Coraza
# rules below apply identically to H3 traffic.
bind quic4@0.0.0.0:443 ssl crt {{ crt_path }} alpn h3
# Advertise H3 so browsers upgrade their existing TCP (h2) connection to
# QUIC on the next request. `ma` is how long (seconds) the client may
# cache the advertisement. http-after-response applies it to every
# response, including haproxy-generated ones (blocks, default page).
http-after-response set-header alt-svc "h3=\":443\"; ma=86400"
# Capture Host header so it appears in httplog output (in %hr field) # Capture Host header so it appears in httplog output (in %hr field)
http-request capture req.hdr(Host) len 64 http-request capture req.hdr(Host) len 64
@@ -64,67 +49,6 @@ frontend web
# High error rate: >100 errors in 30s (scanner/fuzzer behavior) # High error rate: >100 errors in 30s (scanner/fuzzer behavior)
http-request tarpit deny_status 403 if { sc_http_err_rate(0) gt 100 } !is_local !is_trusted_ip !is_whitelisted !is_health_check http-request tarpit deny_status 403 if { sc_http_err_rate(0) gt 100 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
# --- WordPress wp-login.php brute-force protection ---
# The generic limits above are deliberately high (media-heavy sites), so a
# slow credential-stuffing run (dozens of login POSTs/min) slips under them.
# Track POSTs to wp-login.php per real client IP in a DEDICATED 60s table
# (sc1 / backend wp_bruteforce, defined in hap_security_tables.tpl) and
# tarpit once an IP exceeds 30/min. Only login POSTs are counted — GETs of
# the login form, normal browsing, and the handful of POSTs a legit user
# makes are unaffected; an offending IP can still browse, just not keep
# hammering login. path_end also covers subdirectory WP installs. Honors the
# same whitelist (RFC1918 / trusted_ips.list / trusted_ips.map).
acl wp_login_path path_end /wp-login.php
http-request track-sc1 var(txn.real_ip) table wp_bruteforce if METH_POST wp_login_path
http-request tarpit deny_status 429 if METH_POST wp_login_path { sc_http_req_rate(1) gt 30 } !is_local !is_trusted_ip !is_whitelisted
# --- WordPress wp-login.php "must-load-the-form-first" cookie challenge ---
# Defeats DISTRIBUTED credential-stuffing (hundreds of thousands of unique
# IPs, each low-and-slow, so the per-IP rule above can't see them). Such
# bots POST straight to /wp-login.php without ever GETting the form — on
# these sites the login POST:GET ratio is ~15:1. We hand out a cookie when
# the form is actually fetched (GET) and require it on POST; direct-POST
# bots lack it and are denied AT THE EDGE before reaching PHP. Real logins
# are unaffected — WordPress login already requires loading the page and
# accepting cookies. Immediate deny (NOT tarpit) — under a 300k-POST flood,
# holding tarpit connections would exhaust HAProxy. Honors the whitelist.
# Mark login-form GETs at REQUEST time (method/path are reliably evaluable
# here; in the response phase they are not) so the cookie is emitted on the
# form's own response.
http-request set-var(txn.wp_login_form) int(1) if METH_GET wp_login_path
http-after-response add-header set-cookie "whplc=1; Path=/; Max-Age=1800; HttpOnly; Secure; SameSite=Lax" if { var(txn.wp_login_form) -m found }
acl has_login_cookie req.cook(whplc) -m found
http-request deny deny_status 403 if METH_POST wp_login_path !has_login_cookie !is_local !is_trusted_ip !is_whitelisted
# WordPress REST batch endpoint lockdown ("wp2shell": CVE-2026-63030 +
# CVE-2026-60137). Chaining a core SQL injection with REST batch-route
# confusion gives unauthenticated RCE on WP 6.9.0-6.9.4 and 7.0.0-7.0.1
# (fixed in 6.9.5 / 7.0.2). Exploits are public and were used against this
# fleet on 2026-07-19/20; one site was compromised via this path before
# patching. This is a virtual patch: it does not repair the vulnerable
# application logic, it only removes reachability, so it stays until every
# site is confirmed on a fixed release.
#
# Both routing forms must be covered -- a rule matching only the pretty
# permalink path leaves the ?rest_route= fallback wide open, and urlp()
# does not URL-decode, hence the third ACL for the %2F spelling.
#
# Anonymous-only. batch/v1 is used legitimately by the block editor for
# multi-entity saves, so a blanket deny would break wp-admin for real
# users; requiring a wordpress_logged_in_* cookie costs them nothing.
# req.cook() needs an exact name and WordPress suffixes a per-site hash,
# so this substring-matches the raw Cookie header instead.
#
# Immediate deny, not tarpit -- holding connections open helps an attacker
# who is already scripting this. Honors the same whitelist as above.
acl wp_batch_path path_beg /wp-json/batch/v1
acl wp_batch_route urlp(rest_route) -i -m beg /batch/v1
acl wp_batch_route_enc query -i -m sub rest_route=%2Fbatch%2Fv1
acl has_wp_logged_in req.hdr(Cookie) -i -m sub wordpress_logged_in_
http-request deny deny_status 403 if wp_batch_path !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
http-request deny deny_status 403 if wp_batch_route !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
http-request deny deny_status 403 if wp_batch_route_enc !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
# IP blocking using map file (manual blocks only) # IP blocking using map file (manual blocks only)
# Map file format: /etc/haproxy/blocked_ips.map contains "<ip_or_cidr> 1" per line # Map file format: /etc/haproxy/blocked_ips.map contains "<ip_or_cidr> 1" per line
# Runtime updates: echo "add map #0 IP_ADDRESS 1" | socat stdio /var/run/haproxy.sock # Runtime updates: echo "add map #0 IP_ADDRESS 1" | socat stdio /var/run/haproxy.sock
-8
View File
@@ -6,11 +6,3 @@ frontend stats
stats refresh 30s stats refresh 30s
stats show-legends stats show-legends
stats show-node stats show-node
# Dedicated stick-table for WordPress wp-login.php brute-force tracking.
# Tracked via track-sc1 from the `web` frontend (hap_listener.tpl); counts only
# login POSTs per real client IP over a 60s window. Separate from the generic
# sc0 connection/rate table so the login-attempt threshold is independent of
# the (much higher) flood thresholds.
backend wp_bruteforce
stick-table type ip size 100k expire 30m store http_req_rate(60s)
+1 -9
View File
@@ -1,9 +1 @@
# Source-IP whitelist — exempt from HAProxy rate limits (one IP or CIDR per line). 172.116.197.166
# Referenced by templates/hap_listener.tpl:
# acl is_trusted_ip src -f /etc/haproxy/trusted_ips.list
#
# Add trusted source IPs below. Do NOT commit real/personal IPs to this repo —
# it is mirrored publicly. Keep real entries in an untracked local copy, or add
# them directly on the server (the file lives in the /etc/haproxy named volume
# and persists across container recreates).
127.0.0.1
+1 -9
View File
@@ -1,9 +1 @@
# Real-IP whitelist for proxy-header matching — exempt from HAProxy rate limits. 172.116.197.166 1
# Format: "<IP> 1" (one per line). Referenced by templates/hap_listener.tpl:
# acl is_whitelisted var(txn.real_ip),map_ip(/etc/haproxy/trusted_ips.map,0) -m int gt 0
#
# Add trusted real IPs below. Do NOT commit real/personal IPs to this repo —
# it is mirrored publicly. Keep real entries in an untracked local copy, or add
# them directly on the server (the file lives in the /etc/haproxy named volume
# and persists across container recreates).
127.0.0.1 1