Author SHA1 Message Date
shadowdaoandClaude Opus 5 b072786192 fix(domain-removal): stop deleting certificates other live sites are served from
DELETE /api/domain ran, unconditionally, for any ssl_enabled row:

    os.remove(ssl_cert_path)
    certbot delete --cert-name <domain> --non-interactive

`domains.domain` is UNIQUE. `domains.ssl_cert_path` is not, and nothing
anywhere guarded against two rows naming the same file. Sharing is not an
edge case -- it is the normal shape of the table, because
request_ssl_bundle() deliberately creates it: one SAN certificate is issued
as `--cert-name <primary>`, published once to
/etc/haproxy/certs/<primary>.pem, and then EVERY included name's row is
pointed at that same path ("Mark every name in the bundle as ssl_enabled,
all pointing at the same combined .pem").

Measured read-only on the live SQLite in the haproxy-manager containers on
2026-08-23:

  * whp01: 157 domain rows, 150 ssl_enabled, 71 distinct cert paths.
    39 of those paths are referenced by MORE THAN ONE row, covering 118 of
    the 150 SSL-enabled rows. Worst cases: brain-jar.com.pem and
    arclightcourt.com.pem with 10 domains each, hackerpublicradio.org.pem
    with 6, anhonesthost.com.pem with 5.
  * whp02: 33 rows, 31 ssl_enabled, 15 distinct paths, 11 shared across 27
    rows -- including threeworldsoneheart.org.pem, referenced by the apex,
    its www, and mail.threeworldsoneheart.org. A production cleanup of that
    mail.* row was stopped short precisely because removing it would have
    unlinked the PEM the serving site is using.

So removing one domain unlinked a file up to nine other configured domains
were being served from. HAProxy binds the crt directory
(`bind ... ssl crt /etc/haproxy/certs`), so the loss is not noticed until
the next reload or restart, at which point the listener refuses to come up
or those names fall back to the wrong certificate.

`certbot delete` is the worse half. It destroys the lineage's archive, live
symlinks and renewal config; recovery is a fresh, rate-limited ACME order.
The old code passed `--cert-name <domain>`, which is also simply the wrong
lineage for a SAN member: 81 of whp01's 150 SSL-enabled rows have a cert
path whose basename is not their own domain, so for those the call was a
silent no-op -- while for a bundle PRIMARY it deleted the one lineage still
renewing the certificate every other name in the bundle is served with.

The fix refcounts, after the row is deleted so the query answers "who else
still needs this":

  * lineage_name_for_cert_path() -- the lineage is the published bundle's
    basename minus .pem, the same derivation _quarantine_superseded_certs()
    already uses, not the domain being removed.
  * domains_referencing_cert_path() / domains_referencing_lineage() -- the
    remaining rows that name that file, and that lineage.
  * remove_domain() unlinks only when the list is empty, `certbot delete`s
    only when the list is empty, logs the retained names explicitly when it
    skips, and reports them as certificate_retained_for /
    lineage_retained_for in the API response.

ssl_enabled is deliberately not filtered on in the refcount: the two
mistakes are not symmetric. A stale PEM left in the crt directory costs a
few kilobytes; an unlinked live one is HTTPS down for every name it serves.
Cleanup is deferred, not cancelled -- removing the last name on a bundle
still unlinks the file and deletes the lineage.

No row on either production host has ssl_enabled=1 with an empty
ssl_cert_path, so the "no path, no attributable lineage" branch changes
nothing on the current fleet.

Tests (scripts/test-cert-write-safety.py, +9, suite now 31, all offline):
last reference -> file unlinked and lineage deleted; shared file survives
removal of a SAN member AND of the bundle primary, byte-for-byte, with the
edge still starting; shared lineage is not certbot-deleted; the production
mail.* shape; removing every name eventually cleans up; an unrelated
bundle is never collateral damage. Assertions are on os.path.exists, file
contents and the recorded certbot argv, never on which branch ran.

Mutation-tested, all five mutants killed:
  1. guard absent entirely (suite run with HAPROXY_MANAGER_DIR pointed at
     main) -> 6 failures, incl. "example.com is still configured and still
     served from this file".
  2. refcount taken before the row is deleted -> 5 failures, incl. "the
     last reference is gone - now it may be removed".
  3. certbot guard removed, file guard kept -> 3 failures, incl.
     [] != ['delete --cert-name example.com --non-interactive'].
  4. lineage taken from the domain name instead of the cert path -> 3
     failures, incl. 'delete --cert-name example.com' != 'delete
     --cert-name www.example.com'.
  5. file-unlink guard removed, certbot guard kept -> 4 failures, incl.
     "two sites are still served from this bundle".

Other suites unchanged and green: test-config-rollback, test-cert-scripts,
test-stick-table-contract, test-runtime-map-contract.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-23 04:47:26 -07:00
shadowdaoandClaude Opus 5 c8d16b6990 fix(ip-blocking): the runtime map fast path has never once run
add_ip_to_runtime_map() and remove_ip_from_runtime_map() sent
`add map #0 <ip> 1` / `del map #0 <ip>` to /tmp/haproxy-cli and returned True
whenever socat exited 0. Neither command has ever worked, on any deployment,
for the entire life of the feature -- while logging "Added IP x to runtime map"
every single time. Two independent defects:

  * NO `@1` PREFIX. /tmp/haproxy-cli is HAProxy's MASTER CLI socket; map
    commands are worker commands. Captured verbatim on whp01:

        $ echo "add map #0 192.0.2.77 1" | socat stdio /tmp/haproxy-cli
        Unknown command: 'add', but maybe one of the following ones is a better match:
          @!<pid>   : send a command to the <pid> process
          ...
        $ echo $?
        0

    socat exits 0 on the rejection, so `result.returncode == 0` was true. Same
    silence PR #7 fixed on the `show table` path.
  * `#0` IS NOT A VALID MAP ID. Ids are assigned at config-parse time and move
    on every config regeneration -- `@1 show map` on whp01 reports
    blocked_ips.map as 37 and trusted_ips.map as 10. There is no id 0.
    Hardcoding any number is wrong; the map is referenced by FILE PATH, which
    is what haproxy.cfg itself names in map_ip(/etc/haproxy/blocked_ips.map,0).

And a third silence, which is why a response-body check alone is not enough
here: `@1 add map #0 <ip> 1` returns an EMPTY body, exit 0, and adds nothing to
any map -- while `@1 del map #0 <ip>` and `@1 show map #0` both answer
`Unknown map identifier.`. On the add path the reply is byte-for-byte identical
to success. Only reading the entry back can tell them apart.

IP blocking itself was never broken: update_blocked_ips_map() rewrites
/etc/haproxy/blocked_ips.map and the callers reload HAProxy, which re-reads it.
That path is untouched and stays authoritative. What was broken is the
no-reload fast path, plus every report that it had worked.

  * haproxy_manager.py: both functions send `@1 add|del map
    /etc/haproxy/blocked_ips.map <ip> [1]` and READ THE ENTRY BACK with
    `get map` before returning True. runtime_map_lookup()/runtime_map_keys()
    are the read-back primitives. `sync_blocked_ips` loses `clear map #0`
    (which the master socket rejected just as loudly and just as invisibly) and
    verifies the whole set with one `show map` instead of counting commands
    that did not visibly complain; it answers 207 + `runtime_map_synced: false`
    when the runtime map does not match the database.
  * haproxy_cli() grows `expect_empty=True` for MUTATING commands: HAProxy
    answers those with nothing on success, so an empty body is the success and
    ANY non-empty body is a rejection. That is stricter than the marker list on
    purpose -- markers only recognise rejections someone has already seen, and
    it catches `'add map' expects three parameters ...`, which matches nothing.
    HaproxyCliError carries `.responses` so `del map` answering `Key not found.`
    (the requested end state) is told apart from a real failure without regex.
  * The four callers capture the boolean instead of discarding it and report
    `runtime_map_updated` / `runtime_map_failures` in the API response and the
    operation log. A runtime failure degrades to "enforced on the reload that
    already happens two lines later" -- never to an unblocked IP, never to a
    500.
  * scripts/test-runtime-map-contract.py (offline, 26 tests) asserts the bytes
    on the wire (`@1` first, map by path, value `1`), classifies every captured
    response, and scans the repo's Python string literals and shell/template
    code lines for `#<id>` map references -- comments may describe the old
    form, code may not use it. Verified to fail on each defect reintroduced
    separately: no `@1` (3 failures), `#0` (4), no read-back (2), trust-the-
    reply (1).
  * The `#0` form is also corrected in IP_BLOCKING_API.md, MIGRATION_GUIDE.md
    and the comment in templates/hap_listener.tpl -- where every copy of it
    additionally omitted the `1`, which `-m int gt 0` needs to match.

The only template change is a comment; `haproxy -c` on the live rendered config
with it applied is clean (HAProxy 3.0.11, warnings unchanged).

Verified on whp01 against the running container (docker cp + SIGHUP, no
recreate). Before: both functions returned True and logged success while
`@1 get map` answered `found=no` and entry_cnt stayed at 263. After: the fixed
add lands with value "1" and the remove takes it out again; the old command
form is now classified as a failure; a `#0` map reference returns False via the
read-back. End to end through the API, `runtime_map_updated: true`, and
/api/blocked-ips/sync -- which used to be a no-op reporting a full sync --
reports 264/264 verified present.

The runtime path was isolated from the reload that normally follows it: with
NO map-file write and NO reload (same haproxy worker pid throughout), adding
100.123.171.78 (whp01's own netbird overlay address -- not a customer IP, not
in the is_local ranges) to the runtime map alone flipped a live site from
HTTP 200 to 403, and removing it flipped it back to 200. That is the fast path
working for the first time. All test IPs were removed afterwards: 0 rows in
blocked_ips, 0 lines in the map file, entry_cnt back to 263. Six customer
sites, the panel /health and `haproxy -c` are byte-identical to the baseline
taken before the change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 11:15:01 -07:00
jknapp e33167159d Merge pull request 'fix(security-stats): stop reporting counters the stick tables never stored' (#7) from fix/stick-table-field-contract into main
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m20s
Reviewed-on: #7
2026-08-22 17:57:57 +00:00
shadowdaoandClaude Opus 5 b6a62e7f9f fix(security-stats): stop reporting counters the stick tables never stored
/api/security/stats and scripts/show-tarpit-ips.sh reported "Scan Count",
"offense count" and BLOCKED/TARPITTED status parsed from gpc0/gpc1. No stick
table in this repo has ever stored a general-purpose counter -- the `web` table
stores conn_cur, conn_rate(10s), http_req_rate(10s), http_err_rate(30s), and
the two brute-force tables store http_req_rate(60s). Every one of those figures
was fabricated, and an operator was making decisions on them.

Three independent silences kept it alive:

  * `int(parts[3])` on a positional split hit `exp=368842`, raised ValueError,
    and the loop `continue`d -- so the endpoint always answered
    `active_threats: 0` with an empty list. Live on whp01 it also reported
    parts[0], the `0x...:` allocation pointer, as the source IP.
  * The command was sent to /tmp/haproxy-cli WITHOUT the `@1` worker prefix.
    That is the MASTER CLI socket, which answers "Unknown command: 'show' ..."
    -- and socat still exits 0, so the `returncode != 0` guard never fired.
    `total_tracked_ips` was the line count of that help text (8) while the real
    table held 388 entries.
  * The shell consumers wrote `gpc0=${gpc0:-0}`, rendering a field that does
    not exist as a confident zero.

Report what the tables actually store, rather than adding gpc counters to make
the old semantics real. Adding them would mean editing hap_listener.tpl -- the
one change here with a silent-total-outage failure mode -- to rebuild
enforcement history that the edge access log (shipped 2026.08.8, on the host at
/var/log/haproxy.log) already records per request, with status codes,
termination states and request references the stick table could never hold.

  * haproxy_manager.py: STICK_TABLE_FIELD_CONTRACT names what each table
    stores. haproxy_cli() sends worker commands with `@1`, falls back to the
    bare form for a plain stats socket, and inspects the RESPONSE BODY because
    socat's exit status is worthless here. parse_stick_table_entry() reads
    name=value / name(window_ms)=value pairs by NAME, never by position.
    read_stick_table() RAISES -- naming the field -- when a row is missing a
    contract field, instead of defaulting it to 0.
  * /api/security/stats returns the four real counters with their windows, the
    true `used:` count, and no invented threat_level/blocked/offense_count.
    Fewer numbers, all of them real.
  * scripts/show-edge-ip-rates.sh replaces the fabricated report; the four
    expected fields are declared once as EXPECTED_FIELDS and drive the parser.
    show-tarpit-ips.sh becomes a shim that explains why its numbers are gone
    and points at where tarpit events actually live.
  * monitor-attacks.sh loses fourteen fabricated "threat" categories and a
    composite threat score, all permanently zero; its access-log section now
    says the log is on the host instead of silently printing nothing.
  * haproxy_tarpit_config.txt -- the never-shipped design sketch these counters
    were copied from -- gets a NOT IMPLEMENTED banner.
  * scripts/test-stick-table-contract.py (offline, 21 tests) holds the
    templates' `store` clauses, STICK_TABLE_FIELD_CONTRACT and every consumer
    to each other, and asserts each loud-failure path against the real captured
    responses. Template and consumers can no longer drift apart quietly.

No template is touched, so haproxy.cfg is unchanged.

Verified on whp01: total_tracked_ips now tracks `used:` exactly (511 vs the
table's 511, was 8 vs 388), and per-IP values match `show table web key <ip>`
field for field. haproxy PIDs unmoved, `haproxy -c` warnings unchanged, five
customer sites HTTP 200.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 10:48:49 -07:00
Claude b2f835a88c fix(logging): capture the full User-Agent, drop per-request SPOE log noise
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m19s
Two defects caught by watching the real production access log on whp01 in the
minutes after 2026.08.7 made access logging work for the first time.

1. User-Agent was being truncated to its tail.

   `http-request capture req.hdr(User-Agent)` treats the header as a
   comma-separated list and returns only the LAST element. Real User-Agent
   strings contain commas, so

     Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36
     (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36

   logged as

     ua=like Gecko) Chrome/131.0.0.0 Safari/537.36

   losing the platform half -- exactly the half needed to tell a spoofed
   crawler from a real browser, which is one of the main reasons the field was
   added. Switched to req.fhdr(), which returns the full unsplit header value.

2. SPOE was writing one log line per inspected request.

   `log global` inside the spoe-agent block emitted

     SPOE: [coraza] <GROUP:coraza-req> sid=537 st=0 0/0/0/0/0 32/32 0/0 0/467

   for every single request. Measured on whp01: 618 SPOE lines against 669 real
   access lines -- ~48% of the log volume, roughly doubling the edge's log
   footprint (~400 MB/day extra) to record `st=0` over and over.

   It carries nothing incident response needs. The WAF verdict is already in
   the access line (status 403 plus the id= UUID, which joins to
   /var/log/coraza/audit.log for the rule_id), and per-transaction WAF detail
   is written by the SPOA itself to /var/log/coraza/spoa.log. Agent-level
   failures still surface through `option set-on-error error` ->
   var(txn.coraza.error) and the fail-open path in hap_listener.tpl.

Verified: scripts/validate-rendered-config.py passes `haproxy -c` on both the
"default" and "full" scenarios against the real 3.0.11 binary; wp-admin gate,
trusted-proxy gate and xmlrpc rate-limit suites all still pass.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 10:17:07 -07:00
Claude 711c670319 fix(logging): stop silently discarding every access log line
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m46s
haproxy.cfg's global section has had `log 127.0.0.1 local2` since day one.
That is the CONTAINER's own loopback: nothing has ever listened on udp/514 in
the container netns and there is no /dev/log in the image. Every access log
line -- ~1.5M/day across ~60 customer sites -- was written to a socket with no
receiver and dropped. Nothing errored, nothing warned, and `haproxy -c` was
perfectly happy, so this survived unnoticed.

The cost only shows up during an incident. Per-IP 429s, tarpits, wp-admin gate
redirects, WAF 403s and `silent-drop`s left no record anywhere, so the edge
could not be asked what it had actually rejected -- only aggregate stick-table
counters survived. That blind spot applies to every WHP host.

Changes:

* hap_header.tpl: point `log` at {{ syslog_target }} (default 172.18.0.1:514,
  the client-net bridge gateway) with `len 2048 format rfc5424 local2 info`.
  WHP's setup-haproxy-syslog.sh installs the matching rsyslog receiver on the
  host, in a dedicated ruleset ending in stop() so 1.5M lines/day cannot flood
  /var/log/messages or the Graylog forwarder, bound to the bridge IP rather
  than 0.0.0.0.

* haproxy_manager.py: render that target from HAPROXY_SYSLOG_TARGET so
  standalone/home deployments on a different bridge subnet can retarget it.

* hap_listener.tpl: add a frontend-scoped `log-format`. `option httplog` is
  not sufficient for incident response -- it omits %ID entirely (verified
  against 3.0.11), and its %ci is the Cloudflare edge rather than the visitor
  for CF-fronted sites. The new format keeps the first 16 fields byte-identical
  to the httplog default (so existing parsers still work) and appends
  cip=<real client, from var(txn.real_ip)>, id=<uuid>, host=, ua=, sni=, hv=.
  Adds a User-Agent capture in slot 1 to feed it.

* hap_header.tpl: correct the comment claiming `option httplog` includes %ID.
  It does not, which made the documented support-correlation workflow
  (X-Request-Reference -> access log -> coraza audit.log -> rule_id) look
  supported when it could never have worked.

Deliberately NOT using `log stdout format raw local0`: it is incompatible with
the `daemon` keyword, and incompatible SILENTLY. Verified on the pinned 3.0.11
binary -- with `daemon` set, a `log stdout` config serves traffic normally and
emits zero log lines, while `haproxy -c` returns 0 with no error and no
warning, so scripts/validate-rendered-config.py could not catch it either.
Making it work would mean dropping `daemon`, which breaks the three
synchronous `subprocess.run(['haproxy', '-W', ...], check=True)` launch sites
in haproxy_manager.py -- the exact code path whose failure mode is "container
Up, ports 80/443 never bound, every site down, /health still 200".

UDP was chosen so a dead listener degrades to dropped log lines rather than a
stalled request path.

Verified: scripts/validate-rendered-config.py passes `haproxy -c` on both the
"default" and "full" scenarios against the real 3.0.11 binary; and a live
haproxy running WITH `daemon` (as production does) was confirmed to emit real
lines carrying the true client IP from CF-Connecting-IP:

  <150>1 2026-08-22T17:05:21+00:00 - haproxy 109 - - 127.0.0.1:51194
  [22/Aug/2026:17:05:21.217] t t/<NOSRV> 0/-1/-1/-1/0 200 73 - - LR--
  1/1/0/0/0 0/0 {cf-site.example|Mozilla/5.0 RealVisitor} "GET /checkout/
  HTTP/1.1" cip=203.0.113.77 id=dfe94fa9-8d95-4126-81e1-821578f22872
  host=cf-site.example ua=Mozilla/5.0 RealVisitor sni=- hv=1

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-22 10:11:30 -07:00
jknapp 67837f59cb Merge pull request 'fix(haproxy): silence the ACL pattern warning on wp_admin_asset' (#6) from fix/haproxy-acl-pattern-warning into main
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 3m28s
Reviewed-on: #6
2026-08-17 19:59:17 +00:00
shadowdao e7d08c3b30 chore: release 2026.08.6 2026-08-17 12:54:22 -07:00
shadowdao 465253c640 fix(haproxy): silence the ACL pattern warning on wp_admin_asset
HAProxy warns on any pattern whose first character is "(", because it
cannot distinguish an intended regex from a fetch-argument list with a
stray space:

  parsing acl 'wp_admin_asset' : matching 'path_reg' for pattern
  '(^|/)wp-admin/...' is likely a mistake and probably not what you want.

"--" is HAProxy's documented end-of-flags marker and is the remedy the
warning itself names. Cosmetic to matching, but not to operations: left
unsilenced it fires on every config load and every reload on every host,
which trains operators to skim past warnings and gives a real one
somewhere to hide.

Matching semantics are unchanged, verified rather than assumed. Both
forms were run side by side as two frontends under real HAProxy
3.0.11-1+deb13u3 and gave identical verdicts on all 8 vectors:

  /wp-admin/css/login.min.css        MATCH   / MATCH
  /wp-admin/js/user-profile.min.js   MATCH   / MATCH
  /wp-admin/images/x.png             MATCH   / MATCH
  /wp-admin/css/sub/deep.css         MATCH   / MATCH
  /blog/wp-admin/css/a.css           MATCH   / MATCH
  /wp-admin/css/x.php                NOMATCH / NOMATCH
  /wp-admin/plugins.php              NOMATCH / NOMATCH
  /some--path/file.css               NOMATCH / NOMATCH

The last vector is the one that matters: it proves HAProxy consumed "--"
as end-of-flags rather than adopting it as the pattern. Had it done the
latter, the ACL would have matched paths containing "--" and stopped
matching css/js -- gating every login page's own stylesheets while the
page itself still returned 200.

wp_admin_path needs no "--" only because its "-i" flag already occupies
the flag slot; it is not otherwise special.

Adds a regression test asserting both the "--" and the pattern it
guards, so this cannot pass by the pattern having been changed. Verified
to fail when the "--" is removed.
2026-08-17 12:54:22 -07:00
shadowdao b931baa9a7 chore: release 2026.08.5
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 3m22s
2026-08-14 09:55:57 -07:00
shadowdao 2148d72334 Merge branch 'ci/haproxy-config-gate' 2026-08-14 09:55:57 -07:00
shadowdao 17ca731ed3 build: pin haproxy to 3.0.11-1+deb13u3
Previously unpinned, so the binary could move at Debian's timing and break an
unrelated commit's build -- or ship an edge that refuses to start. Pinning and
the haproxy -c gate compose: pinned means the version moves deliberately, and
the gate then answers whether the new binary still accepts fleet config. Also
makes the image reproducible.
2026-08-14 09:55:51 -07:00
shadowdao bcd56b8352 ci: gate the build on a real haproxy -c of the rendered config
A config change can render perfectly, pass every unit test in scripts/, and
still be rejected outright by HAProxy. That happened on 2026-08-14: an inline
`regsub((^|/)wp-admin/.*,\1wp-login.php)` in a redirect location produced
"invalid arg 2 in converter 'regsub'". Thirteen tests were green. It was only
caught because someone built an image by hand and ran `haproxy -c`.

Nothing between commit and production would have stopped it. The unit suites
assert on the TEXT of the rendered config with regexes, which says what the
template emits, never whether HAProxy accepts it. test-config-rollback.py's
"validation" stubs the haproxy binary with a shell script that rejects one
sentinel token and has never parsed a line of real syntax. And
.gitea/workflows/build-push.yaml is checkout -> build -> push, with no tests
at all.

The production consequence is not a broken deploy, it is a silent outage:
init.py refuses to start HAProxy on an invalid config while the container
still comes up, so ports 80/443 are unbound, every site on the host is down,
and /health keeps answering 200.

scripts/validate-rendered-config.py renders the config through the real
generate_config() - every template, real order, both conditional branches
({%- if suspension_enabled %} and {%- if coraza_spoe_backend %}) rendered on
in one scenario and off in the other - creates the stub files the config
loads via `-f` (a missing one is a FATAL haproxy error and would be a false
failure), then runs `haproxy -c` and gates on its EXIT CODE. Warnings are
expected on a clean config ("Can't load stats file", path_reg advisories) and
are not failures; on a real failure the full haproxy output plus the offending
config lines go to the build log.

It runs as a Dockerfile RUN rather than a CI step so it cannot be skipped, so
it protects local builds too, and - the reason that matters most - so it
validates against the EXACT haproxy binary in the image being built. The
Dockerfile installs haproxy unpinned, so that binary moves between builds;
this turns "the new haproxy rejects our config" from a silent production risk
into a build failure. Gating in CI instead would also have meant splitting
build-push-action's single build-and-push step.

The six existing unit suites run in the same step. They had never run
anywhere automated either, and they cost about five seconds.

Verified both ways: the clean build passes and the gate's output appears in
the log; reintroducing the known-bad regsub into a copy of the template fails
the build with HAProxy's own "invalid arg 2 in converter 'regsub' : missing
arguments (got 1/2)".
2026-08-14 09:41:31 -07:00
shadowdaoandClaude Opus 5 e545f3b6e0 test(haproxy): harden wp-admin gate suite against comment-collision, fix two false doc claims
Adversarial mutation audit found the wp-admin gate test suite (26 tests, all
green) did not actually test the feature: 14 of 26 assertions ran bare
str.index/assertIn/re.search over the full rendered config, so they matched
this file's own explanatory comment blocks (which quote ACL names and whole
rules) just as happily as the real rule. Deleting the entire redirect rule,
or `acl wp_admin_allowed`, or all five normalizers, left the old suite at
26/26 PASS. rule_lines() also only stripped whole-comment lines, so a
trailing " # decoy" comment on a surviving line could impersonate a deleted
one, and one ordering test used bare cfg.index() which still "finds" a
normalize-uri directive that has been fully commented out (the substring
survives after the '#').

Rewrites every rule-presence/content/ordering assertion to go through
rule_lines()/rule_positions(), now truncating each line at the first ' #'
before matching, and adds require_rule()/require_position() guards so a
missing rule raises a named AssertionError instead of IndexError or
"substring not found". Adds dedicated declared-ACL tests for wp_admin_path,
wp_admin_asset, wp_admin_allowed and wp_gate_exempt so each has its own
direct, comment-safe check. 29 tests now (was 26).

Proved via a mutation harness (copy templates to a scratch dir, mutate the
copy, run the suite via HAPROXY_MANAGER_DIR, restore): commenting out the
redirect rule, either deny rule, any of the four wp_admin_* ACLs, any one of
the five normalize-uri lines, or expose-experimental-directives now reddens
the suite -- 13/13 required mutations caught, plus the exact trailing-comment
decoy and "all five normalizers commented at once" cases from the audit.

Also corrects two doc claims the audit found factually wrong:

- hap_listener.tpl: normalize-uri's percent-to-uppercase and
  percent-decode-unreserved rewrite the WHOLE request-target, not just the
  path -- measured examples included, and the query-sort-by-name rejection
  reasoning ("every rule matches path") was a non-sequitur given that. Real
  reason to leave it off: reordering would break signed/cached URLs. Fleet
  checked: no .NET backends, no URL-in-path proxies, no known victim today.

- hap_header.tpl: dropping expose-experimental-directives does not
  crash-loop the container. do_initial_setup() swallows the `haproxy -c`
  failure and start_haproxy() returns without raising, so start-up.sh execs
  gunicorn as PID 1 anyway -- a silent total outage (ports 80/443 unbound,
  every site down) that ensure_haproxy.py retries forever without
  escalating, while GET /health keeps answering 200.

No HAProxy rule, ACL, or normalizer changed -- comments and tests only.
Verified: all 5 required suites green, and `haproxy -c` against the real
haproxy 3.0.11 (Debian package) still exits 0 with only the same pre-existing
warnings as before (wp_admin_asset path_reg advisory, stats file).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 09:38:49 -07:00
shadowdaoandClaude Opus 5 8a8d9c5fe3 fix(haproxy): normalise the URI before matching, closing five gate bypasses
The wp-admin edge gate matched the RAW request path while the backend
normalised and decoded it before resolving a file. Every gap between those
two behaviours was a bypass, and five had already been patched individually:

  //wp-admin/plugins.php            fell through ungated
  /wp-admin/css/../plugins.php      took the static-asset bypass
  /wp-admin/js/%2e%2e/plugins.php   same, percent-encoded
  /wp%2Dadmin/plugins.php           matched no wp-admin ACL at all
  /wp-admin%2Fplugins.php           encoded separator, served by OLS

Stop patching vectors and normalise once, first, so every path-based rule in
the frontend sees the same string the backend will resolve:

    percent-to-uppercase
    percent-decode-unreserved
    path-merge-slashes
    path-strip-dot
    path-strip-dotdot full

Order was determined empirically against real haproxy 3.0.11, not from the
docs: the decoders MUST precede the path walkers, or %2e%2e is decoded to ..
only after path-strip-dotdot has already run and the traversal survives. Plain
path-strip-dotdot also leaves /../../ untouched -- "full" is required.
query-sort-by-name is deliberately not enabled; it reorders query parameters
and would break anything signing or caching on the exact query string.

normalize-uri is experimental in 3.0, so global gains
expose-experimental-directives -- without it haproxy does not start at all.
The two must be added and removed together.

%2F cannot be closed by normalisation ("/" is reserved, so decoding it is
correctly refused), so it gets its own deny, scoped to paths mentioning
wp-admin so non-WordPress apps that pass encoded slashes in path parameters
keep working. Deny rather than redirect: regsub finds no "/wp-admin/" in
"/wp-admin%2F...", so a redirect would point at the request's own URL.

Gate changes:
  * wp_admin_safe_path KEPT -- merge-slashes kills its "//" vector but not
    "/\", which no normalizer touches. Its failure mode (unsafe path is not
    redirected, therefore falls through UNGATED -- the original C1) is now
    closed by an explicit deny instead of being left implicit.
  * wp_admin_asset now excludes .php, so the asset bypass cannot cover a PHP
    entrypoint even if an encoding trick ever survives normalisation.
  * wp_admin_path is case-insensitive, paired with a matching regsub flag --
    adding either alone is an infinite redirect loop.

Verified behaviourally against real haproxy 3.0.11 with raw sockets (curl
normalises client-side and hides these), run twice: once against the rendered
templates and once against the haproxy.cfg generated by a real, healthy
container. 12/12 gated, 19/19 passed through, 7/7 with no off-site Location,
plus ~30 adversarial vectors. haproxy -c exits 0 and the container reaches
healthy. Blast radius measured on a 40-URL production-shaped corpus: 4
rewritten, all RFC-equivalent (%7E->~, /./ , //); query strings and all
non-unreserved escapes byte-identical.

Full evidence:
.superpowers/sdd/2026-08-14-wpadmin-edge-gate/task-4-normalize-report.md

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 09:09:35 -07:00
shadowdaoandClaude Opus 5 18750861b4 fix(haproxy): close open redirect in wp-admin edge gate
The redirect target is built by regsub-rewriting `path`, which only
replaces the matched "/wp-admin/.*" substring -- anything before it
survives untouched. Three request forms turn that survival into an
off-site Location header: a protocol-relative "//evil/wp-admin/x.php",
a browser-normalized "/\evil/wp-admin/x.php", and an RFC 7230
absolute-form request target. Without this gate those paths simply
404 against WordPress; the gate itself is what would have exposed a
fleet-wide phishing primitive.

Adds a positive wp_admin_safe_path ACL (path_reg ^/[^/\\]) requiring a
well-formed absolute path, required alongside the existing conditions
on the redirect rule. A path that fails it is simply not redirected
and falls through to the backend -- pre-gate behavior, so no
regression. set-var is left unguarded since it only computes a
variable; the redirect is what emits the header, so guarding it is
sufficient.

Verified against real HAProxy 3.0.11: the naive two-backslash form
fails to compile (config-line word parsing collapses "\\" to one
backslash before PCRE sees it, leaving an unterminated class); four
backslashes are required in the template so PCRE receives the
intended single-backslash class member. Confirmed live, via a
differential test against the pre-fix rule, that both the // and /\
vectors previously produced off-site Location headers and now do not,
while normal root and subdirectory-install redirects are unaffected.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:19:02 -07:00
shadowdaoandClaude Opus 5 6b0b5893b6 fix(haproxy): repair wp-admin edge gate redirect + anchor allowlist
Two defects in the wp-admin edge gate (2171bed, 704be38):

1. HAProxy 3.0.11 rejects the inline regsub redirect
   (regsub((^|/)wp-admin/.*,\1wp-login.php)) with "invalid arg 2 in
   converter 'regsub': missing arguments". Verified this is a
   converter-argument-parenthesis-counting limitation -- the inner
   "(^|/)" grouping parens are misread as closing the outer regsub()
   call, and neither quoting nor backslash-escaping the parens helps.
   Since HTTP paths always start with "/", the group is unnecessary:
   compute the login URL in its own set-var, matching the literal
   substring "/wp-admin/" (no group, no backreference) and replacing
   it with the literal "/wp-login.php" -- regsub only replaces the
   matched substring, so a subdirectory-install prefix survives
   untouched.

2. wp_admin_allowed used a bare path_end suffix match
   (/admin-ajax.php etc), so /wp-admin/evil/admin-ajax.php matched
   both wp_admin_path and the allowlist and sailed through the gate
   ungated. Anchored each entry to /wp-admin/<file>.

Verified against real HAProxy 3.0.11-1+deb13u3: haproxy -c exit 0,
and live curl against the real generated config's literal lines
confirms root-install and subdirectory-install redirects, the
anchored-allowlist fix, cookie exemption, and non-wp-admin passthrough
all behave correctly.

Extends scripts/test-wpadmin-gate.py with regression tests for the
anchored allowlist and the set-var ordering/no-inline-regsub guard.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:02:00 -07:00
shadowdao 704be38882 feat(haproxy): gate unauthenticated wp-admin requests at the edge
Redirect /wp-admin/* to the site's login page when no wordpress_logged_in_
cookie is present, so unauthenticated requests never boot PHP. Identity-based
rather than rate-based, so it is unaffected by how widely an attack is
distributed. Allowlists the paths that legitimately serve unauthenticated
visitors, including the css/js the login page itself loads.
2026-08-14 07:45:52 -07:00
shadowdao 2171bedb20 feat(haproxy): ship per-site exempt list for the wp-admin edge gate 2026-08-14 07:43:20 -07:00
shadowdao 992bf49138 chore: release 2026.08.4
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m8s
2026-08-13 14:52:59 -07:00
shadowdao 491f54928a Merge branch 'feat/xmlrpc-rate-limit' 2026-08-13 14:52:59 -07:00
shadowdaoandClaude Opus 5 ecc1184533 feat(haproxy): rate-limit POST /xmlrpc.php floods per client IP
Mirrors the existing wp-login.php brute-force protection. Generic frontend
limits trigger at 300-500 req/s (sized for media-heavy pageloads), but
observed xmlrpc floods run at just a few req/s for hours -- well under that
ceiling while still pinning PHP-FPM workers and driving 503s fleet-wide
(1,011 in one day on a single site).

Adds a dedicated stick-table (xmlrpc_bruteforce, sc2) rather than reusing
wp_bruteforce: sharing a counter would let wp-login and xmlrpc traffic from
the same IP inflate each other's rate. Tarpits at 60 req/min/IP (double
wp-login's 30, since xmlrpc is machine-to-machine and legitimately bursts --
Jetpack sync, mobile app, remote publishing). Honors the same whitelist as
every other rule in the file and does not block the endpoint outright.

Only safe to key on var(txn.real_ip) because of the trusted-proxy header
gate shipped earlier today (2026.08.3) -- before that, per-IP tracking was
trivially evaded via a spoofed X-Forwarded-For.

Adds scripts/test-xmlrpc-rate-limit.py (stdlib unittest, no pytest in this
repo) pinning the tracking rule, the tarpit threshold, the path_end ACL, and
the whitelist exclusions. Existing trusted-proxy-gate, config-rollback, and
cert-write-safety regression suites all still pass unmodified.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 14:48:25 -07:00
shadowdao 53422c35e8 Merge branch 'fix/seed-lists-past-volume'
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m3s
2026-08-13 13:50:55 -07:00
shadowdaoandClaude Opus 5 f21ade06d9 fix(haproxy): stop the trusted-proxy volume from shadowing Cloudflare/proxy lists
/etc/haproxy is a named volume in deployed containers, so the baked-in
cloudflare_ips.list and trusted_proxies.list COPYed there in the prior
task never actually reached hosts with a pre-existing volume -- the
start-up.sh guard then found them "missing" and created them empty.
With both lists empty, the from_trusted_proxy ACL in hap_listener.tpl
matched nothing, so CF-Connecting-IP / X-Real-IP / X-Forwarded-For got
stripped from every peer, including Cloudflare's own edge. Confirmed
live: image shipped 34/13 lines, running container had 0/0.

Fix: stage both files under /haproxy/defaults (outside the volume) and
apply their ownership rule in start-up.sh instead of a blind
"create if missing":
  - cloudflare_ips.list is shipped data -- always refresh it from the
    baked default so Cloudflare range updates reach existing hosts.
  - trusted_proxies.list is operator data -- seed it from the baked
    default only when missing, and never overwrite what an operator
    added on the server.
Both branches fall back to creating an empty file if the baked default
is somehow absent, since a missing "-f" target is a fatal HAProxy
config error.

Verified against a volume pre-populated to shadow the image (mimicking
a real host): cloudflare_ips.list repopulates with all 15 IPv4 + 7
IPv6 ranges even after being truncated and restarted; a distinctive
operator entry appended to trusted_proxies.list survives a restart
untouched; haproxy -c still validates cleanly.

Release-worthy fix for a defect from the just-released 2026.08.2 build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:48:34 -07:00
22 changed files with 4055 additions and 412 deletions
+47
View File
@@ -7,8 +7,55 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
### Testing
- **API Testing**: `./scripts/test-api.sh` - Tests all API endpoints with optional authentication
- **Certificate Request Testing**: `./scripts/test-certificate-request.sh` - Tests certificate generation endpoints
- **Stick-table contract**: `python3 scripts/test-stick-table-contract.py` - offline; holds the templates' `store` clauses, `STICK_TABLE_FIELD_CONTRACT`, and every consumer to each other. Run it after touching any `stick-table` line.
- **Runtime-map contract**: `python3 scripts/test-runtime-map-contract.py` - offline; asserts the runtime map commands are `@1`-prefixed, reference the map by FILE PATH (never `#<id>`), carry the value `1`, and that every captured rejection is classified as a failure. Run it after touching any `add map`/`del map`/`clear map` path.
- **Certificate destruction safety**: `python3 scripts/test-cert-write-safety.py` - offline; asserts a live `.pem` is never truncated, removed, or its certbot lineage deleted while any configured domain still references it (one bundle serves many names, so `ssl_cert_path` is routinely shared). Run it after touching any `os.remove`/`certbot delete`/PEM-write path.
- **Manual Testing**: Run `curl` commands against `http://localhost:8000` endpoints as shown in README.md
### Reading stick tables (and why it is easy to get silently wrong)
`/tmp/haproxy-cli` is HAProxy's **master** CLI socket. Worker commands
(`show table`, `show map`, `add map`, ...) need an `@1` prefix. Without it
HAProxy answers `Unknown command: 'show', ...` **and socat still exits 0** — so
an exit-status check passes and the help text gets parsed as data. Always use
`haproxy_cli(cmd, worker=True)` in Python, which inspects the response body.
Stick-table entries are `name=value` / `name(window_ms)=value` pairs, not fixed
columns; the first token is an allocation pointer (`0x...:`), not the key. Parse
by NAME, and treat a missing field as an ERROR — never default it to `0`. The
`web` table stores only `conn_cur`, `conn_rate`, `http_req_rate`,
`http_err_rate`; it holds **no history and no counter of past blocks**. What was
actually denied/tarpitted is in the edge access log on the **host** at
`/var/log/haproxy.log` (shipped 2026.08.8), not in any stick table.
This is written down because `/api/security/stats` and `show-tarpit-ips.sh`
reported "Scan Count"/"BLOCKED" figures parsed from `gpc0`/`gpc1` — fields no
stick table has ever stored — for their entire existence. See the header of
`haproxy_tarpit_config.txt` and the contract test.
### Changing a runtime map (`add map` / `del map`)
Same socket, two more ways to fail silently — and both were live in
`add_ip_to_runtime_map()`/`remove_ip_from_runtime_map()` for their whole
existence:
* **Reference the map by FILE PATH, never `#<id>`.** Ids are assigned at
config-parse time and move on every config regeneration (on whp01
`blocked_ips.map` is 37, `trusted_ips.map` is 10 — there is no id 0). Use
`add map /etc/haproxy/blocked_ips.map <ip> 1`.
* **A mutation answers NOTHING on success**, so an empty body is the only
success — any output at all is a rejection. Worse, `@1 add map #0 <ip> 1`
*also* answers nothing and adds nothing, so the body cannot prove an add
worked. **Read it back** with `@1 get map <path> <key>`.
* Entries must carry the value `1`; haproxy.cfg matches with
`map_ip(...,0) -m int gt 0`, so a valueless entry does not block.
In Python use `haproxy_cli(cmd, worker=True, expect_empty=True)` for mutations
and `runtime_map_lookup()` / `runtime_map_keys()` to verify. The runtime map is
only a fast path: `/etc/haproxy/blocked_ips.map` is authoritative and HAProxy
re-reads it on reload, so a failed runtime command must degrade to
"enforced on reload" and be reported, never swallowed.
### Running the Application
- **Docker Build**: `docker build -t haproxy-manager .`
- **Local Development**: `python haproxy_manager.py` (requires HAProxy, certbot, and dependencies installed)
+57 -4
View File
@@ -23,7 +23,23 @@ LABEL org.opencontainers.image.title="haproxy-manager-base" \
org.opencontainers.image.version="${VERSION}" \
org.opencontainers.image.licenses="MIT"
RUN apt update -y && apt dist-upgrade -y && apt install socat haproxy cron certbot curl jq net-tools -y && apt clean && rm -rf /var/lib/apt/lists/*
# haproxy is PINNED. It was previously unpinned, so the binary could move under
# us at Debian's timing — an upstream release that rejected our config would have
# broken an unrelated commit's build, or worse, shipped an edge that refuses to
# start (see the `haproxy -c` gate below for why that matters: a config HAProxy
# rejects leaves the container Up with 80/443 unbound and /health still 200).
#
# Pinning does NOT make the gate redundant, and the gate does NOT make pinning
# unnecessary — they compose. Pinned means the version moves deliberately; the
# gate then answers immediately whether the new binary still accepts fleet config.
# It also makes the image reproducible, which it previously was not.
#
# To move it: bump the version here, rebuild, and let the gate verify. If Debian
# security-updates the package (e.g. -1+deb13u4) the build FAILS until this pin is
# updated — that failure is the point, not a bug. Check availability with:
# apt-cache policy haproxy
ARG HAPROXY_VERSION=3.0.11-1+deb13u3
RUN apt update -y && apt dist-upgrade -y && apt install socat "haproxy=${HAPROXY_VERSION}" cron certbot curl jq net-tools -y && apt-mark hold haproxy && apt clean && rm -rf /var/lib/apt/lists/*
WORKDIR /haproxy
COPY ./templates /haproxy/templates
COPY requirements.txt /haproxy/
@@ -31,15 +47,52 @@ COPY haproxy_manager.py /haproxy/
COPY scripts /haproxy/scripts
COPY trusted_ips.list /etc/haproxy/trusted_ips.list
COPY trusted_ips.map /etc/haproxy/trusted_ips.map
COPY cloudflare_ips.list /etc/haproxy/cloudflare_ips.list
COPY trusted_proxies.list /etc/haproxy/trusted_proxies.list
# /etc/haproxy is a named volume in deployed containers, so baked-in files
# under that path get shadowed by the volume on existing deployments.
# under that path get shadowed by the volume on existing deployments. The
# trusted_ips.* pair above predates that discovery and is handled by the
# older start-up.sh guard (out of scope here). cloudflare_ips.list and
# trusted_proxies.list are staged under /haproxy/defaults instead, so
# start-up.sh can always read the image's baked copy regardless of what the
# volume shadows /etc/haproxy with.
COPY cloudflare_ips.list /haproxy/defaults/cloudflare_ips.list
COPY trusted_proxies.list /haproxy/defaults/trusted_proxies.list
COPY wpadmin_gate_exempt.list /haproxy/defaults/wpadmin_gate_exempt.list
# Place errorfiles outside the volumed path; the HAProxy config references
# them by absolute path.
COPY errors /haproxy/errors
RUN chmod +x /haproxy/scripts/*
RUN pip install -r requirements.txt
# ---------------------------------------------------------------------------
# Build gate: no image ships unless the real haproxy binary accepts the config
# this image's templates actually produce.
#
# On 2026-08-14 a template change rendered fine, passed all 13 unit tests, and
# was rejected by HAProxy ("invalid arg 2 in converter 'regsub'"). It was only
# caught because someone built an image by hand and ran `haproxy -c`. Nothing
# in the build or in CI would have stopped it: .gitea/workflows/build-push.yaml
# is checkout -> build -> push, and test-config-rollback.py's "haproxy" is a
# shell stub that only rejects a sentinel token. In production an invalid
# haproxy.cfg means init.py refuses to start HAProxy while the container stays
# Up - ports 80/443 unbound, every site on the host down, /health still 200.
#
# This lives in the Dockerfile rather than in the workflow deliberately:
# * it cannot be skipped, and it protects local `docker build` too;
# * no workflow restructuring (build-push-action builds and pushes in one
# step, so gating in CI would mean splitting build from push);
# * it validates against the EXACT haproxy binary in this image. Line 26
# installs haproxy unpinned, so that binary moves between builds - this
# turns "the new haproxy rejects our config" from a silent production
# risk into a build failure.
#
# The unit suites run here too. They had never run anywhere automated either,
# and they cost a few seconds.
RUN python3 /haproxy/scripts/test-wpadmin-gate.py \
&& python3 /haproxy/scripts/test-trusted-proxy-gate.py \
&& python3 /haproxy/scripts/test-xmlrpc-rate-limit.py \
&& python3 /haproxy/scripts/test-config-rollback.py \
&& python3 /haproxy/scripts/test-cert-write-safety.py \
&& python3 /haproxy/scripts/test-cert-scripts.py \
&& python3 /haproxy/scripts/validate-rendered-config.py
# Create log directories
RUN mkdir -p /var/log && touch /var/log/haproxy-manager.log /var/log/haproxy-manager-errors.log
RUN chmod 755 /var/log/haproxy-manager.log /var/log/haproxy-manager-errors.log
+30 -5
View File
@@ -508,20 +508,45 @@ curl -X POST http://localhost:8000/api/blocked-ips/sync \
For advanced users, you can interact directly with HAProxy's runtime API:
Three things about these commands are easy to get wrong, and each one fails
**silently** (socat exits 0 either way — the rejection, if any, is only in the
response body):
* `/tmp/haproxy-cli` is HAProxy's **master** CLI socket. Map commands are
worker commands and need the `@1` prefix. Without it the reply is
`Unknown command: 'add', ...`.
* Reference the map by its **file path**, never by `#<id>`. Ids are assigned at
config-parse time and move on every config regeneration (on a live edge,
`blocked_ips.map` is id 37, `trusted_ips.map` is 10 — there is no id 0).
Worse, `@1 add map #0 <ip> 1` returns an **empty** reply and adds nothing.
* Entries must carry the value `1`. `haproxy.cfg` matches with
`map_ip(...,0) -m int gt 0`, so a valueless entry does not block. (`add map`
with no value is rejected: `'add map' expects three parameters ...`.)
```bash
MAP=/etc/haproxy/blocked_ips.map
# Add IP to runtime (immediate effect)
echo "add map #0 192.168.1.100" | socat stdio /var/run/haproxy.sock
echo "@1 add map $MAP 192.168.1.100 1" | socat stdio /tmp/haproxy-cli
# Remove IP from runtime
echo "del map #0 192.168.1.100" | socat stdio /var/run/haproxy.sock
echo "@1 del map $MAP 192.168.1.100" | socat stdio /tmp/haproxy-cli
# Confirm what actually happened (do not trust the exit status)
echo "@1 get map $MAP 192.168.1.100" | socat stdio /tmp/haproxy-cli
# Clear all blocked IPs from runtime
echo "clear map #0" | socat stdio /var/run/haproxy.sock
echo "@1 clear map $MAP" | socat stdio /tmp/haproxy-cli
# Show all runtime map entries
echo "show map #0" | socat stdio /var/run/haproxy.sock
# Show all runtime map entries, and the map ids currently in use
echo "@1 show map $MAP" | socat stdio /tmp/haproxy-cli
echo "@1 show map" | socat stdio /tmp/haproxy-cli
```
The runtime map is a **fast path only**. `/etc/haproxy/blocked_ips.map` is
authoritative: HAProxy re-reads it on reload, so a failed runtime command
delays a block until the next reload rather than losing it.
## Migration from ACL Method
If you're upgrading from the old ACL-based method:
+10 -2
View File
@@ -50,14 +50,22 @@ http-request deny status 403 if { src -f /etc/haproxy/blocked_ips.map }
- **Graceful error handling**
### 2. Runtime IP Management
Map commands go to a **worker** (`@1`), reference the map by **file path**
(ids move between config regenerations, and `#0` silently adds nothing), and
carry the value `1` that `map_ip(...,0) -m int gt 0` matches on:
```bash
# Add IP without reload (immediate effect)
echo "add map #0 192.168.1.100" | socat stdio /var/run/haproxy.sock
echo "@1 add map /etc/haproxy/blocked_ips.map 192.168.1.100 1" | socat stdio /tmp/haproxy-cli
# Remove IP without reload
echo "del map #0 192.168.1.100" | socat stdio /var/run/haproxy.sock
echo "@1 del map /etc/haproxy/blocked_ips.map 192.168.1.100" | socat stdio /tmp/haproxy-cli
```
socat exits 0 even when HAProxy rejects the command, so read the response body
(or read the entry back with `@1 get map ...`) rather than the exit status.
See IP_BLOCKING_API.md for the full set.
### 3. New API Endpoints
#### Safe Config Reload
+1 -1
View File
@@ -1 +1 @@
2026.08.2
2026.08.11
+674 -148
View File
@@ -1624,6 +1624,61 @@ def request_certificates():
else:
return jsonify(response), 500 # All failed
def lineage_name_for_cert_path(cert_path):
"""The certbot lineage name implied by a published bundle path.
Every writer in this module publishes to ``{SSL_CERTS_DIR}/<name>.pem`` and
issues that lineage with ``--cert-name <name>`` (see request_ssl() and
request_ssl_bundle()); _quarantine_superseded_certs() derives the lineage
from the filename the same way. So the basename minus ``.pem`` IS the
lineage, and it is NOT necessarily the domain being removed: a bundle
issued as ``--cert-name example.com`` also serves www.example.com and any
other SAN, all of whose DB rows point at ``/etc/haproxy/certs/example.com.pem``.
"""
if not cert_path:
return None
base = os.path.basename(cert_path)
return base[:-len('.pem')] if base.endswith('.pem') else base
def domains_referencing_cert_path(cursor, cert_path):
"""Domains (still in the table) whose ssl_cert_path is exactly `cert_path`.
Call this AFTER the row being removed has been deleted, so the answer is
"who else still needs this file".
ssl_enabled is deliberately NOT filtered on. HAProxy binds the whole crt
directory, so a file is load-bearing for any row that names it; and the
asymmetry of the two mistakes is total - keeping a stale PEM costs nothing,
unlinking a live one is HTTPS down for every name it serves.
"""
if not cert_path:
return []
cursor.execute(
'SELECT domain FROM domains WHERE ssl_cert_path = ? ORDER BY domain',
(cert_path,))
return [row[0] for row in cursor.fetchall()]
def domains_referencing_lineage(cursor, lineage):
"""Domains (still in the table) served by certbot lineage `lineage`.
Same shape as domains_referencing_cert_path(), one level further back:
`certbot delete --cert-name X` destroys the archive, live symlinks and
renewal config for X. If any remaining domain is served by a bundle
published from that lineage, deleting it means the next renewal silently
stops happening for all of them and the material cannot be recovered
without a fresh, rate-limited ACME order.
"""
if not lineage:
return []
cursor.execute(
"SELECT domain, ssl_cert_path FROM domains "
"WHERE ssl_cert_path IS NOT NULL AND ssl_cert_path != ''")
return sorted({domain for domain, path in cursor.fetchall()
if lineage_name_for_cert_path(path) == lineage})
@app.route('/api/domain', methods=['DELETE'])
@require_api_key
def remove_domain():
@@ -1662,33 +1717,74 @@ def remove_domain():
# Delete domain
cursor.execute('DELETE FROM domains WHERE id = ?', (domain_id,))
# Refcount the certificate BEFORE anything is unlinked or deleted,
# with the row above already gone so the query answers "who else
# still needs this". One .pem serves many names: request_ssl_bundle()
# issues a single SAN cert and points EVERY included domain's row at
# the same /etc/haproxy/certs/<primary>.pem. Removing one of those
# names used to os.remove() that file and `certbot delete` its
# lineage unconditionally - taking HTTPS down for every other name
# in the bundle, and destroying the only recoverable copy with it.
cert_path_users = domains_referencing_cert_path(cursor, ssl_cert_path)
lineage = lineage_name_for_cert_path(ssl_cert_path) if ssl_cert_path else None
# The lineage to delete is the one this domain's bundle came from,
# not the domain's own name - for a SAN member those differ.
lineage_users = domains_referencing_lineage(cursor, lineage)
cert_retained_for = []
lineage_retained_for = []
# Delete SSL certificate from HAProxy certs directory
if ssl_enabled and ssl_cert_path:
try:
os.remove(ssl_cert_path)
logger.info(f"Removed HAProxy certificate file: {ssl_cert_path}")
except OSError as e:
logger.warning(f"Failed to remove certificate file {ssl_cert_path}: {e}")
if cert_path_users:
cert_retained_for = cert_path_users
logger.info(
"Kept HAProxy certificate file %s while removing %s: still "
"referenced by %d other domain(s): %s",
ssl_cert_path, domain, len(cert_path_users),
', '.join(cert_path_users))
else:
try:
os.remove(ssl_cert_path)
logger.info(f"Removed HAProxy certificate file: {ssl_cert_path}")
except OSError as e:
logger.warning(f"Failed to remove certificate file {ssl_cert_path}: {e}")
# Remove certificate from certbot
if ssl_enabled:
try:
result = subprocess.run(
['certbot', 'delete', '--cert-name', domain, '--non-interactive'],
capture_output=True, text=True
)
if result.returncode == 0:
logger.info(f"Removed Let's Encrypt certificate for {domain}")
else:
logger.warning(f"Failed to remove Let's Encrypt certificate for {domain}: {result.stderr}")
except Exception as e:
logger.warning(f"Error removing Let's Encrypt certificate for {domain}: {e}")
if not lineage:
logger.info(
"Skipping certbot delete for %s: no certificate path on the "
"removed row, so no lineage can be attributed to it", domain)
elif lineage_users:
lineage_retained_for = lineage_users
logger.info(
"Kept Let's Encrypt lineage %s while removing %s: still "
"serving %d other domain(s): %s",
lineage, domain, len(lineage_users), ', '.join(lineage_users))
else:
try:
result = subprocess.run(
['certbot', 'delete', '--cert-name', lineage, '--non-interactive'],
capture_output=True, text=True
)
if result.returncode == 0:
logger.info(f"Removed Let's Encrypt certificate {lineage} for {domain}")
else:
logger.warning(f"Failed to remove Let's Encrypt certificate {lineage} for {domain}: {result.stderr}")
except Exception as e:
logger.warning(f"Error removing Let's Encrypt certificate {lineage} for {domain}: {e}")
# Regenerate HAProxy config
generate_config()
log_operation('remove_domain', True, f'Domain {domain} removed successfully')
return jsonify({'status': 'success', 'message': 'Domain configuration removed'})
return jsonify({
'status': 'success',
'message': 'Domain configuration removed',
'certificate_retained_for': cert_retained_for,
'lineage_retained_for': lineage_retained_for,
})
except Exception as e:
log_operation('remove_domain', False, str(e))
@@ -1735,8 +1831,10 @@ def add_blocked_ip():
log_operation('add_blocked_ip', False, f'Failed to update map file for {ip_address}')
return jsonify({'status': 'error', 'message': 'Failed to update blocked IPs map file'}), 500
# Add to runtime map for immediate effect
add_ip_to_runtime_map(ip_address)
# Add to runtime map for immediate effect. The map FILE above is what
# actually enforces the block once HAProxy re-reads it; this is only
# the fast path, so a False here is reported, not fatal.
runtime_ok = add_ip_to_runtime_map(ip_address)
# Reload HAProxy to ensure consistency
try:
@@ -1753,8 +1851,18 @@ def add_blocked_ip():
except Exception as e:
logger.warning(f"Error reloading HAProxy after blocking IP {ip_address}: {e}")
log_operation('add_blocked_ip', True, f'IP {ip_address} blocked successfully')
return jsonify({'status': 'success', 'blocked_ip_id': blocked_ip_id, 'message': f'IP {ip_address} has been blocked'})
log_operation('add_blocked_ip', True,
f'IP {ip_address} blocked successfully '
f'(runtime map fast path: {"ok" if runtime_ok else "FAILED, enforced on reload"})')
return jsonify({
'status': 'success',
'blocked_ip_id': blocked_ip_id,
'message': f'IP {ip_address} has been blocked',
# False = the block is enforced from the map file at reload rather
# than instantly. The caller can tell the difference; it used to be
# reported as instant unconditionally.
'runtime_map_updated': runtime_ok
})
except sqlite3.IntegrityError:
log_operation('add_blocked_ip', False, f'IP {ip_address} is already blocked')
return jsonify({'status': 'error', 'message': 'IP address is already blocked'}), 409
@@ -1790,8 +1898,9 @@ def remove_blocked_ip():
log_operation('remove_blocked_ip', False, f'Failed to update map file for {ip_address}')
return jsonify({'status': 'error', 'message': 'Failed to update blocked IPs map file'}), 500
# Remove from runtime map for immediate effect
remove_ip_from_runtime_map(ip_address)
# Remove from runtime map for immediate effect. As with blocking, the
# map file is authoritative and the reload below picks it up.
runtime_ok = remove_ip_from_runtime_map(ip_address)
# Reload HAProxy to ensure consistency
try:
@@ -1808,8 +1917,14 @@ def remove_blocked_ip():
except Exception as e:
logger.warning(f"Error reloading HAProxy after unblocking IP {ip_address}: {e}")
log_operation('remove_blocked_ip', True, f'IP {ip_address} unblocked successfully')
return jsonify({'status': 'success', 'message': f'IP {ip_address} has been unblocked'})
log_operation('remove_blocked_ip', True,
f'IP {ip_address} unblocked successfully '
f'(runtime map fast path: {"ok" if runtime_ok else "FAILED, applied on reload"})')
return jsonify({
'status': 'success',
'message': f'IP {ip_address} has been unblocked',
'runtime_map_updated': runtime_ok
})
except Exception as e:
log_operation('remove_blocked_ip', False, str(e))
return jsonify({'status': 'error', 'message': str(e)}), 500
@@ -1843,103 +1958,375 @@ def sync_blocked_ips():
cursor.execute('SELECT ip_address FROM blocked_ips ORDER BY ip_address')
blocked_ips = [row[0] for row in cursor.fetchall()]
# Try to clear all entries from runtime map (might fail if empty, that's ok)
# Clear the runtime map and re-add every blocked IP. The map is
# referenced by FILE PATH: `clear map #0` (what this used to send, with
# no `@1` either) was answered by the master socket with "Unknown
# command: 'clear'" and socat exited 0, so this whole block was a no-op
# that reported a full sync.
try:
if os.path.exists(HAPROXY_SOCKET_PATH):
socket_path = HAPROXY_SOCKET_PATH
else:
socket_path = '/tmp/haproxy-cli'
subprocess.run(f'echo "clear map #0" | socat stdio {socket_path}',
shell=True, capture_output=True)
except:
pass # Clear might fail if map is empty
# Add all IPs to runtime map
success_count = 0
haproxy_cli('clear map %s' % BLOCKED_IPS_MAP_PATH,
worker=True, expect_empty=True)
except HaproxyCliError as e:
log_operation('sync_blocked_ips', False, f'Failed to clear runtime map: {e}')
logger.warning("Failed to clear runtime map: %s. The map file is "
"still authoritative and correct.", e)
return jsonify({
'status': 'error',
'message': f'Failed to clear runtime map: {e}',
'map_file_updated': True,
'runtime_map_synced': False,
'total_ips': len(blocked_ips)
}), 500
# verify=False here: one `show map` read-back below costs a single
# round trip instead of one per IP, and answers the same question for
# the whole set.
accepted = 0
for ip in blocked_ips:
if add_ip_to_runtime_map(ip):
success_count += 1
log_operation('sync_blocked_ips', True, f'Synced {success_count}/{len(blocked_ips)} IPs to runtime map')
if add_ip_to_runtime_map(ip, verify=False):
accepted += 1
# Ground truth, not a count of commands that did not visibly complain.
try:
present = runtime_map_keys(BLOCKED_IPS_MAP_PATH)
except HaproxyCliError as e:
log_operation('sync_blocked_ips', False, f'Could not read back runtime map: {e}')
return jsonify({
'status': 'error',
'message': f'Could not read back the runtime map to verify the sync: {e}',
'map_file_updated': True,
'runtime_map_synced': False,
'total_ips': len(blocked_ips)
}), 500
missing = [ip for ip in blocked_ips if ip not in present]
synced = len(blocked_ips) - len(missing)
ok = not missing
if missing:
logger.warning(
"Runtime map sync incomplete: %d/%d IPs are not in %s (first "
"few: %s). They remain blocked via the map file on reload.",
len(missing), len(blocked_ips), BLOCKED_IPS_MAP_PATH, missing[:5])
log_operation('sync_blocked_ips', ok,
f'Verified {synced}/{len(blocked_ips)} IPs present in the runtime map')
return jsonify({
'status': 'success',
'message': f'Synced {success_count}/{len(blocked_ips)} IPs to runtime map',
'status': 'success' if ok else 'partial',
'message': f'Verified {synced}/{len(blocked_ips)} IPs present in the runtime map',
'total_ips': len(blocked_ips),
'synced_ips': success_count
})
'synced_ips': synced,
'accepted_commands': accepted,
'missing_ips': missing[:50],
'runtime_map_synced': ok
}), (200 if ok else 207)
except Exception as e:
log_operation('sync_blocked_ips', False, str(e))
return jsonify({'status': 'error', 'message': str(e)}), 500
# ---------------------------------------------------------------------------
# HAProxy runtime API (stick tables)
#
# WHY THIS SECTION IS SO DEFENSIVE
# --------------------------------
# The previous /api/security/stats read `gpc0` and `gpc1` out of `show table
# web` and reported them as "scan count" / "offense count" / "blocked". No
# stick table in this repo has EVER stored a general-purpose counter — see
# STICK_TABLE_FIELD_CONTRACT below and the `store` clauses in
# templates/hap_listener.tpl and templates/hap_security_tables.tpl. Three
# separate silences let that survive:
#
# 1. `int(parts[3])` on a positional split raised ValueError on `exp=368842`
# and the loop just `continue`d, so every row was skipped and the endpoint
# always answered `active_threats: 0` with an empty list. An operator
# reading that saw "no threats" and could not tell it apart from "the
# parser is broken".
# 2. The command was sent WITHOUT a worker prefix. /tmp/haproxy-cli is the
# MASTER socket; `show table web` there is answered with "Unknown command:
# 'show', but maybe one of the following ones is a better match: ..." --
# and socat still exits 0, so `result.returncode != 0` never fired. The
# reported `total_tracked_ips` was literally the number of lines in that
# help text minus one (8), while the real table held 388 entries.
# 3. The shell consumers defaulted every missing field to 0 (`${gpc0:-0}`),
# so a field that does not exist rendered as a confident zero.
#
# Rules for anything added here, all three aimed at the same failure mode:
# * Read the CONTRACT, not positions. Stick-table output is `name=value` /
# `name(window)=value` pairs whose order and presence follow the template's
# `store` clause. Positional indexing silently reads the wrong column the
# moment that clause changes.
# * NEVER default a missing field to a number. A field the table does not
# store must surface as an error naming the field, not as 0.
# * NEVER trust socat's exit status. HAProxy reports command errors in the
# response BODY and the socket still closes cleanly. Use haproxy_cli().
#
# scripts/test-stick-table-contract.py holds STICK_TABLE_FIELD_CONTRACT, the
# rendered templates and the shell consumers to each other, and fails if any
# one of them drifts.
# ---------------------------------------------------------------------------
# What each stick table ACTUALLY stores, per its `store` clause. The single
# source of truth for every consumer in this repo. Keep in sync with the
# templates -- the contract test enforces that, in both directions.
STICK_TABLE_FIELD_CONTRACT = {
'web': ('conn_cur', 'conn_rate', 'http_req_rate', 'http_err_rate'),
'wp_bruteforce': ('http_req_rate',),
'xmlrpc_bruteforce': ('http_req_rate',),
}
# Metadata every stick-table entry carries regardless of the `store` clause.
STICK_TABLE_ENTRY_META = ('key', 'use', 'exp', 'shard')
# HAProxy answers a rejected runtime command in the response body and the
# socket still closes 0. These are the prefixes it uses.
_HAPROXY_CLI_ERROR_MARKERS = (
'Unknown command',
'No such table',
# `add map #0 ...` / `get map /etc/haproxy/nope.map ...` -- captured
# verbatim from HAProxy 3.0.11 on the live edge.
'Unknown map identifier',
'Key not found',
'Permission denied',
"Can't find the specified process",
'unknown process',
'Missing ',
)
_STICK_TABLE_TOKEN_RE = re.compile(r'^([a-z_][a-z0-9_]*)(?:\(([^)]*)\))?=(.*)$')
_STICK_TABLE_HEADER_RE = re.compile(
r'^#\s*table:\s*(?P<name>[^,]+),\s*type:\s*(?P<type>[^,]+),\s*'
r'size:\s*(?P<size>\d+),\s*used:\s*(?P<used>\d+)')
class HaproxyCliError(RuntimeError):
"""A runtime-API command was rejected, timed out, or answered nothing.
Exists so a rejected command cannot be mistaken for an empty result. That
distinction is the whole point of this module's stick-table code.
`.responses` holds the raw, stripped response body of every attempt, so a
caller can tell one rejection apart from another (`del map` answering
"Key not found." is a no-op, not a failure) without regex-matching the
formatted message.
"""
responses = ()
def _haproxy_socket_path():
return HAPROXY_SOCKET_PATH if os.path.exists(HAPROXY_SOCKET_PATH) else '/tmp/haproxy-cli'
def _cli_response_is_error(text):
if text is None:
return True
head = text.lstrip()
return any(head.startswith(marker) for marker in _HAPROXY_CLI_ERROR_MARKERS)
def _cli_send(command, socket_path, timeout):
proc = subprocess.run(
['socat', 'stdio', socket_path],
input=command + '\n', capture_output=True, text=True, timeout=timeout)
if proc.returncode != 0:
raise HaproxyCliError(
'socat failed talking to %s (exit %d): %s'
% (socket_path, proc.returncode, (proc.stderr or '').strip()))
return proc.stdout
def haproxy_cli(command, worker=False, timeout=None, expect_empty=False):
"""Send one runtime-API command and return its response, or raise.
`worker=True` marks a command that only the WORKER answers (show table,
show map, add map, ...). On this deployment the socket is HAProxy's MASTER
CLI, where those need an `@1` prefix; on a plain stats socket they must NOT
have one. Rather than guessing from configuration that can change under us,
try the prefixed form and fall back -- and raise if BOTH are rejected,
instead of returning HAProxy's help text as if it were data.
`expect_empty=True` is for MUTATING commands (add/del/clear map, set map).
HAProxy answers those with nothing at all on success, so the rule inverts:
an empty body is the success, and ANY non-empty body is a rejection. That
is deliberately stricter than matching _HAPROXY_CLI_ERROR_MARKERS -- the
marker list can only ever recognise the rejections someone has already
seen, and a mutation that prints anything has not done what was asked. Two
real examples this catches that the marker list did not:
`'add map' expects three parameters ...` and `Unknown map identifier.`
"""
socket_path = _haproxy_socket_path()
timeout = timeout if timeout is not None else DEFAULT_SUBPROCESS_TIMEOUT
attempts = (['@1 ' + command, command] if worker else [command])
failures = []
bodies = []
for attempt in attempts:
try:
out = _cli_send(attempt, socket_path, timeout)
except subprocess.TimeoutExpired:
raise HaproxyCliError('timed out after %ss running %r on %s'
% (timeout, attempt, socket_path))
body = out.strip()
if expect_empty:
if not body:
return out
elif body and not _cli_response_is_error(out):
return out
bodies.append(body)
failures.append('%r -> %r' % (attempt, body[:200] or '<empty response>'))
error = HaproxyCliError(
'HAProxy rejected %r on %s (socat exited 0 -- the rejection is in the '
'response body, which is exactly why this is checked): %s'
% (command, socket_path, '; '.join(failures)))
error.responses = tuple(bodies)
raise error
def parse_stick_table_entry(line):
"""{field: {'value': str, 'window_ms': int|None}} for one `show table` row.
Parses `name=value` / `name(window)=value` pairs by NAME. The leading
`0x...:` allocation pointer is skipped -- reading it as the key is how the
old code came to report memory addresses as IP addresses.
"""
fields = {}
for token in line.split():
m = _STICK_TABLE_TOKEN_RE.match(token)
if not m:
continue # the 0x...: pointer, or anything else unnamed
name, window, value = m.group(1), m.group(2), m.group(3)
fields[name] = {
'value': value,
'window_ms': int(window) if window and window.isdigit() else None,
}
return fields
def read_stick_table(table):
"""(header dict, [(raw line, parsed fields)]) for a stick table.
Raises HaproxyCliError if the response is not a stick-table dump, or if any
row is missing a field the contract says the table stores. A field the
table does not carry is an ERROR here, never a zero.
"""
expected = STICK_TABLE_FIELD_CONTRACT.get(table)
if expected is None:
raise HaproxyCliError(
'no field contract for stick table %r; add it to '
'STICK_TABLE_FIELD_CONTRACT (and to the templates) first' % table)
raw = haproxy_cli('show table %s' % table, worker=True)
lines = raw.strip().split('\n')
header = _STICK_TABLE_HEADER_RE.match(lines[0]) if lines else None
if not header:
raise HaproxyCliError(
'response to `show table %s` is not a stick-table dump; first line '
'was %r' % (table, lines[0][:200] if lines else ''))
entries = []
for line in lines[1:]:
if not line.strip() or line.lstrip().startswith('#'):
continue
fields = parse_stick_table_entry(line)
if 'key' not in fields:
raise HaproxyCliError(
'stick-table row for %r has no key= field: %r' % (table, line[:200]))
missing = [f for f in expected if f not in fields]
if missing:
raise HaproxyCliError(
'stick table %r no longer stores %s -- STICK_TABLE_FIELD_CONTRACT '
'and the `store` clause in templates/hap_listener.tpl have drifted '
'apart. Present: %s. Offending row: %r'
% (table, ', '.join(missing),
', '.join(sorted(fields)), line[:200]))
entries.append((line, fields))
return {
'name': header.group('name'),
'type': header.group('type'),
'size': int(header.group('size')),
'used': int(header.group('used')),
}, entries
@app.route('/api/security/stats', methods=['GET'])
@require_api_key
def get_security_stats():
"""Get current security statistics from HAProxy stick table"""
"""Per-source connection and request rates from the `web` stick table.
Reports ONLY what the table stores: conn_cur, conn_rate, http_req_rate and
http_err_rate, each with the window HAProxy is actually counting over. It
deliberately does NOT classify a "threat level" or report a "blocked" flag:
the thresholds live in templates/hap_listener.tpl and a copy here would be
a second source of truth free to drift, which is the class of bug this
endpoint used to be. Sources at or over a limit are visible from the rates
themselves, and the enforcement that actually happened -- 429s, tarpits
(termination state PT), WAF denials -- is in the edge access log on the
HOST at /var/log/haproxy.log, which records per-request outcomes the stick
table never held.
Query params:
limit max sources returned (default 50)
min_req_rate only sources at or above this http_req_rate (default 1,
i.e. sources with current activity; pass 0 for all)
"""
try:
if os.path.exists(HAPROXY_SOCKET_PATH):
socket_path = HAPROXY_SOCKET_PATH
else:
socket_path = '/tmp/haproxy-cli'
limit = max(1, min(int(request.args.get('limit', 50)), 1000))
min_req_rate = max(0, int(request.args.get('min_req_rate', 1)))
except ValueError:
return jsonify({'status': 'error',
'message': 'limit and min_req_rate must be integers'}), 400
# Get stick table data
cmd = f'echo "show table web" | socat stdio {socket_path}'
result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
if result.returncode != 0:
return jsonify({'status': 'error', 'message': 'Failed to get stick table data'}), 500
# Parse stick table output
lines = result.stdout.strip().split('\n')
threats = []
for line in lines[1:]: # Skip header
parts = line.split()
if len(parts) >= 8:
ip = parts[0]
try:
gpc0 = int(parts[3]) if len(parts) > 3 else 0
gpc1 = int(parts[4]) if len(parts) > 4 else 0
req_rate = int(parts[5]) if len(parts) > 5 else 0
err_rate = int(parts[6]) if len(parts) > 6 else 0
conn_rate = int(parts[7]) if len(parts) > 7 else 0
# Only include IPs with significant activity
if gpc0 > 0 or gpc1 > 0 or req_rate > 30 or err_rate > 5 or conn_rate > 10:
threat_level = 'low'
if gpc1 > 2:
threat_level = 'critical'
elif gpc0 > 0 or err_rate > 10:
threat_level = 'high'
elif req_rate > 40 or conn_rate > 15:
threat_level = 'medium'
threats.append({
'ip': ip,
'blocked': gpc0 > 0,
'repeat_offender': gpc1 > 2,
'offense_count': gpc1,
'request_rate': req_rate,
'error_rate': err_rate,
'connection_rate': conn_rate,
'threat_level': threat_level
})
except (ValueError, IndexError):
continue
# Sort by threat level
threats.sort(key=lambda x: (x['offense_count'], x['error_rate'], x['request_rate']), reverse=True)
return jsonify({
'status': 'success',
'total_tracked_ips': len(lines) - 1,
'active_threats': len(threats),
'threats': threats[:50] # Limit to top 50
})
try:
header, entries = read_stick_table('web')
except HaproxyCliError as e:
log_operation('get_security_stats', False, str(e))
return jsonify({'status': 'error', 'message': str(e)}), 502
except Exception as e:
log_operation('get_security_stats', False, str(e))
return jsonify({'status': 'error', 'message': str(e)}), 500
counters = STICK_TABLE_FIELD_CONTRACT['web']
windows = {}
sources = []
active = 0
for line, fields in entries:
row = {'ip': fields['key']['value']}
for name in counters:
# A non-numeric counter means HAProxy's output format changed under
# us. Say so; do not coerce it to 0 and report it as a measurement.
try:
row[name] = int(fields[name]['value'])
except ValueError:
msg = ('stick table web reported a non-numeric %s=%r; the '
'`show table` output format has changed. Row: %r'
% (name, fields[name]['value'], line[:200]))
log_operation('get_security_stats', False, msg)
return jsonify({'status': 'error', 'message': msg}), 502
if fields[name]['window_ms'] is not None:
windows.setdefault(name, fields[name]['window_ms'])
if any(row[name] > 0 for name in counters):
active += 1
if row['http_req_rate'] >= min_req_rate:
sources.append(row)
sources.sort(key=lambda r: (r['http_req_rate'], r['http_err_rate'],
r['conn_rate'], r['conn_cur']), reverse=True)
return jsonify({
'status': 'success',
'table': header['name'],
'table_size': header['size'],
'total_tracked_ips': header['used'],
'sources_with_activity': active,
'counters': list(counters),
'counter_windows_ms': windows,
'returned': len(sources[:limit]),
'min_req_rate': min_req_rate,
'sources': sources[:limit],
'note': ('Current per-source rates only. The stick table stores no '
'history and no counter of past blocks; enforcement events '
'are in the edge access log on the host at /var/log/haproxy.log.'),
})
@app.route('/api/security/temporary-block', methods=['POST'])
@require_api_key
def temporary_block():
@@ -1977,13 +2364,16 @@ def temporary_block():
if not update_blocked_ips_map():
return jsonify({'status': 'error', 'message': 'Failed to update map file'}), 500
add_ip_to_runtime_map(ip_address)
runtime_ok = add_ip_to_runtime_map(ip_address)
log_operation('temporary_block', True, f'Temporarily blocked {ip_address} for {duration_minutes} minutes')
log_operation('temporary_block', True,
f'Temporarily blocked {ip_address} for {duration_minutes} minutes '
f'(runtime map fast path: {"ok" if runtime_ok else "FAILED, enforced on reload"})')
return jsonify({
'status': 'success',
'message': f'IP {ip_address} temporarily blocked for {duration_minutes} minutes',
'expires_at': expiry_time.isoformat()
'expires_at': expiry_time.isoformat(),
'runtime_map_updated': runtime_ok
})
except Exception as e:
log_operation('temporary_block', False, str(e))
@@ -1996,6 +2386,10 @@ def clear_expired_blocks():
try:
current_time = datetime.now()
expired_ips = []
# IPs the runtime fast path could not drop. They still come out of the
# map file below, so they unblock on the next reload -- but silently
# reporting them as cleared is what this whole change is about.
runtime_failures = []
with sqlite3.connect(DB_FILE) as conn:
cursor = conn.cursor()
@@ -2016,17 +2410,21 @@ def clear_expired_blocks():
# Remove expired IPs
for ip in expired_ips:
cursor.execute('DELETE FROM blocked_ips WHERE ip_address = ?', (ip,))
remove_ip_from_runtime_map(ip)
if not remove_ip_from_runtime_map(ip):
runtime_failures.append(ip)
# Update map file if any IPs were removed
if expired_ips:
update_blocked_ips_map()
log_operation('clear_expired_blocks', True, f'Cleared {len(expired_ips)} expired IP blocks')
log_operation('clear_expired_blocks', not runtime_failures,
f'Cleared {len(expired_ips)} expired IP blocks '
f'({len(runtime_failures)} not removed from the runtime map)')
return jsonify({
'status': 'success',
'status': 'success' if not runtime_failures else 'partial',
'message': f'Cleared {len(expired_ips)} expired IP blocks',
'cleared_ips': expired_ips
'cleared_ips': expired_ips,
'runtime_map_failures': runtime_failures
})
except Exception as e:
log_operation('clear_expired_blocks', False, str(e))
@@ -2351,9 +2749,20 @@ def generate_config():
except Exception as e:
logger.error(f"Failed to create {suspended_list_path}: {e}")
# Access-log destination for the `log` line in the global section.
# Default 172.18.0.1:514 is the docker bridge gateway for WHP's
# `client-net`, i.e. the host, where rsyslog's imudp listener is bound
# by setup-haproxy-logrotate.sh. Overridable so this image stays usable
# on standalone/home deployments with a different bridge subnet or a
# remote log collector -- set HAPROXY_SYSLOG_TARGET to `<ip>:<port>`.
# UDP, so an absent listener drops log lines and never affects request
# handling.
syslog_target = os.environ.get('HAPROXY_SYSLOG_TARGET', '172.18.0.1:514').strip()
# Add Haproxy Default Headers
default_headers = template_env.get_template('hap_header.tpl').render(
cluster_secret = get_or_create_cluster_secret(),
syslog_target = syslog_target,
)
config_parts.append(default_headers)
@@ -3054,50 +3463,167 @@ def update_blocked_ips_map(promote_backup=True):
logger.error(f"Failed to update blocked IPs map: {e}")
return False
def add_ip_to_runtime_map(ip_address):
"""Add IP to HAProxy runtime map without reload"""
# ---------------------------------------------------------------------------
# Runtime map fast path (blocked IPs)
#
# WHAT WAS WRONG, AND WHY NOBODY NOTICED FOR ITS ENTIRE EXISTENCE
# ---------------------------------------------------------------
# add_ip_to_runtime_map()/remove_ip_from_runtime_map() sent
# `add map #0 <ip> 1` / `del map #0 <ip>` to /tmp/haproxy-cli and returned True
# whenever socat exited 0. Two independent defects, three silences:
#
# 1. NO `@1` PREFIX. /tmp/haproxy-cli is HAProxy's MASTER CLI socket. Map
# commands are worker commands. The master answers
# `Unknown command: 'add', but maybe one of the following ones is a better
# match: ...` -- and socat still exits 0, so `result.returncode == 0` was
# true and the function logged "Added IP x to runtime map".
# 2. `#0` IS NOT A VALID MAP ID. Ids are assigned at config-parse time and
# move whenever the config is regenerated; on whp01 blocked_ips.map is
# id 37 and trusted_ips.map is 10. There is no id 0. Hardcoding ANY number
# is wrong -- reference the map by its FILE PATH, which is stable because
# it is what haproxy.cfg names in `map_ip(/etc/haproxy/blocked_ips.map,0)`.
# 3. `add map #0 ...` fails SILENTLY EVEN WITH `@1`. Captured on HAProxy
# 3.0.11: `@1 add map #0 192.0.2.88 1` returns an EMPTY body, exit 0, and
# adds nothing to any map -- while `@1 del map #0 <ip>` and
# `@1 show map #0` both answer `Unknown map identifier.`. So a
# response-body check alone cannot catch defect 2 on the add path. That is
# why every mutation here is READ BACK with `get map` instead of trusting
# either the exit status or the (empty) reply.
#
# The blocking itself never depended on this: update_blocked_ips_map() rewrites
# /etc/haproxy/blocked_ips.map and the callers reload HAProxy, which re-reads
# the file. The FILE IS AUTHORITATIVE; this is only the no-reload fast path.
# Every function below therefore returns a bool the caller can report, and
# never raises into a request handler -- a runtime-map failure must degrade to
# "enforced on reload", not to "not blocked" and not to a 500.
#
# scripts/test-runtime-map-contract.py holds the command strings and the
# classification of every captured response to these rules.
# ---------------------------------------------------------------------------
# The value every blocked_ips.map entry must carry. haproxy.cfg matches with
# `map_ip(/etc/haproxy/blocked_ips.map,0) -m int gt 0`, so a keyed entry with
# no value evaluates to 0 and is NOT blocked. Runtime map and file must agree.
BLOCKED_IPS_MAP_VALUE = '1'
_GET_MAP_FOUND_RE = re.compile(r'\bfound=(yes|no)\b')
_GET_MAP_VALUE_RE = re.compile(r'\bvalue="([^"]*)"')
def runtime_map_lookup(map_path, key):
"""(found, value) for one key in a runtime map, or raise HaproxyCliError.
Reads back what `add map`/`del map` actually did. `get map` answers
`type=ip, case=sensitive, found=yes, idx=tree, key="1.2.3.4", value="1",
type="str"` or `type=ip, case=sensitive, found=no`; an unusable map
reference answers `Unknown map identifier.`, which haproxy_cli() rejects.
"""
out = haproxy_cli('get map %s %s' % (map_path, key), worker=True)
match = _GET_MAP_FOUND_RE.search(out)
if not match:
raise HaproxyCliError(
'unparseable `get map %s %s` response (no found=yes/no): %r'
% (map_path, key, out.strip()[:200]))
if match.group(1) == 'no':
return (False, None)
value = _GET_MAP_VALUE_RE.search(out)
return (True, value.group(1) if value else None)
def runtime_map_keys(map_path):
"""The set of keys currently in a runtime map, or raise HaproxyCliError.
`show map <file>` emits `<0x-pointer> <key> <value>` per line. An empty map
legitimately emits nothing, which is why this does not go through the
non-empty check.
"""
try:
if os.path.exists(HAPROXY_SOCKET_PATH):
socket_path = HAPROXY_SOCKET_PATH
else:
socket_path = '/tmp/haproxy-cli'
out = haproxy_cli('show map %s' % map_path, worker=True)
except HaproxyCliError as e:
# An empty map is a real, distinguishable state -- not a rejection.
if e.responses and all(body == '' for body in e.responses):
return set()
raise
keys = set()
for line in out.splitlines():
parts = line.split()
if len(parts) >= 2 and parts[0].startswith('0x'):
keys.add(parts[1])
return keys
# Add to runtime map (map file ID 0 for blocked IPs)
# Format: add map #<id> <key> <value>
# For IP blocking, value is always "1"
cmd = f'echo "add map #0 {ip_address} 1" | socat stdio {socket_path}'
result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
if result.returncode == 0:
logger.info(f"Added IP {ip_address} to runtime map")
return True
else:
logger.warning(f"Failed to add IP to runtime map: {result.stderr}")
return False
def add_ip_to_runtime_map(ip_address, verify=True):
"""Add IP to the running HAProxy's blocked map without a reload.
Returns True only when the entry is verifiably present with the value the
config matches on. False means the FILE + reload path is what will enforce
this block -- which it does regardless; see the section comment above.
"""
try:
haproxy_cli(
'add map %s %s %s' % (BLOCKED_IPS_MAP_PATH, ip_address, BLOCKED_IPS_MAP_VALUE),
worker=True, expect_empty=True)
if verify:
found, value = runtime_map_lookup(BLOCKED_IPS_MAP_PATH, ip_address)
if not found:
raise HaproxyCliError(
'`add map` was accepted but %s is not in %s afterwards -- '
'the command did nothing (this is exactly how `#0` failed)'
% (ip_address, BLOCKED_IPS_MAP_PATH))
if value != BLOCKED_IPS_MAP_VALUE:
raise HaproxyCliError(
'%s is in %s with value %r, not %r -- haproxy.cfg matches '
'with `-m int gt 0`, so this entry does NOT block'
% (ip_address, BLOCKED_IPS_MAP_PATH, value, BLOCKED_IPS_MAP_VALUE))
logger.info(f"Added IP {ip_address} to runtime map (verified={verify})")
return True
except HaproxyCliError as e:
logger.warning(
"Runtime map fast path FAILED for %s: %s. The block is NOT lost -- "
"%s was rewritten and HAProxy re-reads it on reload -- but it does "
"not take effect until that reload completes.",
ip_address, e, BLOCKED_IPS_MAP_PATH)
return False
except Exception as e:
logger.error(f"Error adding IP to runtime map: {e}")
logger.error(f"Error adding IP {ip_address} to runtime map: {e}")
return False
def remove_ip_from_runtime_map(ip_address):
"""Remove IP from HAProxy runtime map without reload"""
def remove_ip_from_runtime_map(ip_address, verify=True):
"""Remove IP from the running HAProxy's blocked map without a reload.
Returns True only when the key is verifiably gone. `Key not found.` means
the runtime map never had it, which is the requested end state, so that is
a success -- but it is logged, because it also means the runtime map and
the file had drifted apart.
"""
try:
if os.path.exists(HAPROXY_SOCKET_PATH):
socket_path = HAPROXY_SOCKET_PATH
else:
socket_path = '/tmp/haproxy-cli'
# Remove from runtime map (map file ID 0 for blocked IPs)
cmd = f'echo "del map #0 {ip_address}" | socat stdio {socket_path}'
result = subprocess.run(cmd, shell=True, capture_output=True, text=True)
if result.returncode == 0:
logger.info(f"Removed IP {ip_address} from runtime map")
return True
else:
logger.warning(f"Failed to remove IP from runtime map: {result.stderr}")
return False
try:
haproxy_cli('del map %s %s' % (BLOCKED_IPS_MAP_PATH, ip_address),
worker=True, expect_empty=True)
except HaproxyCliError as e:
if not any(body.startswith('Key not found') for body in e.responses):
raise
logger.info(
"Runtime map had no entry for %s to remove (`Key not found.`); "
"the file and the runtime map had drifted", ip_address)
if verify:
found, _ = runtime_map_lookup(BLOCKED_IPS_MAP_PATH, ip_address)
if found:
raise HaproxyCliError(
'`del map` was accepted but %s is STILL in %s'
% (ip_address, BLOCKED_IPS_MAP_PATH))
logger.info(f"Removed IP {ip_address} from runtime map (verified={verify})")
return True
except HaproxyCliError as e:
logger.warning(
"Runtime map fast path FAILED for %s: %s. The unblock is NOT lost -- "
"%s was rewritten and HAProxy re-reads it on reload -- but the IP "
"stays blocked until that reload completes.",
ip_address, e, BLOCKED_IPS_MAP_PATH)
return False
except Exception as e:
logger.error(f"Error removing IP from runtime map: {e}")
logger.error(f"Error removing IP {ip_address} from runtime map: {e}")
return False
def start_haproxy():
+34
View File
@@ -1,3 +1,37 @@
# =============================================================================
# NOT IMPLEMENTED. THIS FILE IS A DESIGN SKETCH THAT WAS NEVER SHIPPED.
# =============================================================================
#
# Nothing here is deployed, has ever been deployed, or is rendered into
# haproxy.cfg. The real edge config is generated from templates/*.tpl. Compare:
#
# THIS FILE proposes: store gpc0,gpc1,gpc2,http_err_rate(30s),...
# plus sc-inc-gpc0/1/2 scan-escalation rules
# templates/hap_listener.tpl ACTUALLY has:
# store conn_cur,conn_rate(10s),http_req_rate(10s),http_err_rate(30s)
#
# There is no gpc0, no gpc1, no gpc2, and no scan-escalation state anywhere on
# the edge, and never has been.
#
# WHY THE BANNER. This file was mistaken for the shipped config. Two consumers
# -- /api/security/stats in haproxy_manager.py and scripts/show-tarpit-ips.sh --
# were written to parse gpc0/gpc1 out of `show table web`, defaulting missing
# fields to 0. The result was a "Scan Count" column and BLOCKED/TARPITTED
# statuses that were pure fabrication, presented to an operator as fact, for
# the entire life of both tools. Fixed 2026-08-22; scripts/test-stick-table-
# contract.py now fails if any consumer's field expectations and the templates'
# `store` clauses ever drift apart again.
#
# IF YOU WANT TO REVIVE ANY OF THIS: the counters must be added to a template
# `store` clause and to STICK_TABLE_FIELD_CONTRACT in haproxy_manager.py first.
# Adding them to a consumer alone produces confident zeros, not data. Note also
# that sc0/sc1/sc2 are all in use and HAProxy's tune.stick-counters defaults to
# 3, and that since 2026.08.8 the edge has real per-request access logging on
# the host at /var/log/haproxy.log -- which records what was actually denied,
# tarpitted and rate-limited, with request references. That log is a better
# source for most of what this sketch was reaching for.
# =============================================================================
global
daemon
log stdout local0 info
+130 -118
View File
@@ -1,136 +1,148 @@
#!/bin/bash
#!/usr/bin/env bash
#
# monitor-attacks.sh — HAProxy edge activity monitor.
#
# Two sections, both fed from real data:
# 1. Current per-IP rates, from the `web` stick table (delegated to
# show-edge-ip-rates.sh — there is exactly one stick-table parser).
# 2. Recent enforcement events, from the HAProxy access log.
#
# HISTORY / WHY THIS IS SHORTER THAN IT USED TO BE
# The previous version printed a "Threat Intelligence Dashboard" with
# fourteen categories (auth_fail, authz_fail, scanner, sql_inj, traversal,
# wp_brute, admin_scan, shell_att, repeat_off, manual_bl, auto_bl,
# glitch_rate, ...) and a composite "threat score", all parsed out of
# gpc(0), gpc(1), gpc(3), gpc(12), gpc(13) and glitch_rate(300s). NONE of
# those fields exist: the `web` table stores only conn_cur, conn_rate,
# http_req_rate and http_err_rate. Every category was permanently 0 and the
# whole dashboard printed nothing while implying it was watching. All of it
# has been deleted rather than "fixed" — there was no data source to fix it
# against.
#
# Usage: monitor-attacks.sh [live]
# Env: LOG_FILE=<path> access log to read (default /var/log/haproxy.log)
# LOG_LINES=<n> how many trailing log lines to scan (default 500)
# Real-time attack monitoring for HAProxy
# Shows blocked requests and suspicious activity
set -uo pipefail
LOG_FILE="/var/log/haproxy.log"
SOCKET="/tmp/haproxy-cli"
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
LOG_FILE="${LOG_FILE:-/var/log/haproxy.log}"
LOG_LINES="${LOG_LINES:-500}"
echo "==================================================="
echo "HAProxy Security Monitor - Real-time Attack Detection"
echo "==================================================="
echo ""
# Function to show current threats with HAProxy 3.0.11 metrics
show_threats() {
echo "HAProxy 3.0.11 Threat Intelligence Dashboard:"
echo "show table web" | socat stdio "$SOCKET" 2>/dev/null | \
awk 'NR>1 {
# Parse the stick table output for array-based GPC values
ip = $1
# Look for GPC array values in the data
auth_fail = 0
authz_fail = 0
rate_viol = 0
scanner = 0
sql_inj = 0
traversal = 0
wp_brute = 0
admin_scan = 0
shell_att = 0
repeat_off = 0
manual_bl = 0
auto_bl = 0
glitch_rate = 0
threat_score = 0
# Extract relevant metrics (simplified parsing)
if ($0 ~ /gpc\(0\)=([0-9]+)/) {
match($0, /gpc\(0\)=([0-9]+)/, arr); auth_fail = arr[1]
}
if ($0 ~ /gpc\(1\)=([0-9]+)/) {
match($0, /gpc\(1\)=([0-9]+)/, arr); authz_fail = arr[1]
}
if ($0 ~ /gpc\(3\)=([0-9]+)/) {
match($0, /gpc\(3\)=([0-9]+)/, arr); scanner = arr[1]
}
if ($0 ~ /gpc\(12\)=([0-9]+)/) {
match($0, /gpc\(12\)=([0-9]+)/, arr); repeat_off = arr[1]
}
if ($0 ~ /gpc\(13\)=([0-9]+)/) {
match($0, /gpc\(13\)=([0-9]+)/, arr); manual_bl = arr[1]
}
if ($0 ~ /glitch_rate\(300s\)=([0-9]+)/) {
match($0, /glitch_rate\(300s\)=([0-9]+)/, arr); glitch_rate = arr[1]
}
# Calculate composite threat score (simplified)
threat_score = auth_fail*10 + authz_fail*8 + scanner*12 + repeat_off*25 + manual_bl*100
# Only show IPs with significant threat indicators
if (auth_fail > 0 || authz_fail > 0 || scanner > 0 || repeat_off > 0 || manual_bl > 0 || glitch_rate > 0) {
threat_level = "LOW"
if (threat_score >= 100) threat_level = "CRITICAL"
else if (threat_score >= 50) threat_level = "HIGH"
else if (threat_score >= 20) threat_level = "MEDIUM"
printf "%-15s [%8s] Score:%-3d Auth:%-2d Authz:%-2d Scanner:%-1d Repeat:%-1d Glitch:%-2d\n",
ip, threat_level, threat_score, auth_fail, authz_fail, scanner, repeat_off, glitch_rate
}
}' | head -15
echo ""
echo "Top HTTP/2 Protocol Violators:"
echo "show table web" | socat stdio "$SOCKET" 2>/dev/null | \
awk 'NR>1 && $0 ~ /glitch/ {
if ($0 ~ /glitch_rate\(300s\)=([0-9]+)/) {
match($0, /glitch_rate\(300s\)=([0-9]+)/, arr)
if (arr[1] > 2) {
printf "%-15s glitch_rate:%-3s\n", $1, arr[1]
}
}
}' | head -5
echo "---------------------------------------------------"
# --- Section 1: current rates (real stick-table data) -----------------------
show_rates() {
"$SCRIPT_DIR/show-edge-ip-rates.sh" "$@"
}
# Function to show recent blocks
# --- Section 2: recent enforcement events (real access-log data) ------------
show_recent_blocks() {
echo "Recent Blocked Requests:"
tail -100 "$LOG_FILE" 2>/dev/null | \
grep -E "(bot_scanner|scan_admin|scan_shells|sql_injection|directory_traversal|rate_abuse|tarpit|denied|403)" | \
tail -10 | \
awk '{
if (match($0, /[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+:[0-9]+/)) {
ip = substr($0, RSTART, RLENGTH)
gsub(/:.*/, "", ip)
reason = ""
if ($0 ~ /bot_scanner/) reason = "BOT_SCANNER"
else if ($0 ~ /scan_admin/) reason = "ADMIN_SCAN"
else if ($0 ~ /scan_shells/) reason = "SHELL_SCAN"
else if ($0 ~ /sql_injection/) reason = "SQL_INJECTION"
else if ($0 ~ /directory_traversal/) reason = "DIR_TRAVERSAL"
else if ($0 ~ /rate_abuse/) reason = "RATE_ABUSE"
else if ($0 ~ /tarpit/) reason = "TARPIT"
else if ($0 ~ /denied/) reason = "DENIED"
else if ($0 ~ /403/) reason = "BLOCKED"
printf "[%s] %-15s %s\n", strftime("%H:%M:%S"), ip, reason
echo "Recent enforcement events (last $LOG_LINES log lines):"
echo
if [ ! -f "$LOG_FILE" ] || [ ! -r "$LOG_FILE" ]; then
cat <<MSG
Access log not readable at: $LOG_FILE
This is expected INSIDE the haproxy-manager container: HAProxy logs to
syslog on the DOCKER HOST, and the file lives on the host, not in here.
Read it from the host instead:
grep -aE ' (PT|PR)--' /var/log/haproxy.log | tail -20 # tarpit / deny
grep -aE ' (429|403) ' /var/log/haproxy.log | tail -20 # rate-limited / blocked
grep -a 'cip=<IP>' /var/log/haproxy.log | tail -50 # one client IP
grep -a 'id=<uuid>' /var/log/haproxy.log # one request reference
# (the UUID on the block page)
tail -f /var/log/haproxy.log | grep -aE ' (PT|PR)--' # live
Or point this script at a copy: LOG_FILE=/path/to/haproxy.log $0
MSG
echo
return 0
fi
printf "%-15s %-16s %-4s %-5s %-28s %s\n" "TIME" "CLIENT IP" "CODE" "TERM" "HOST" "REQUEST / REQUEST-ID"
printf "%s\n" "-----------------------------------------------------------------------------------------------------"
local found
found=$(tail -n "$LOG_LINES" "$LOG_FILE" 2>/dev/null | awk '
{
status = ""; term = ""; cip = ""; host = ""; id = ""; ts = ""; req = ""
# %tr is bracketed: [22/Aug/2026:10:11:12.345] -> keep HH:MM:SS
# No {n} interval expressions here: not every awk in a slim Debian
# image supports them. Spelled out instead.
if (match($0, /\[[0-9][0-9]\/[A-Za-z][A-Za-z][A-Za-z]\/[0-9][0-9][0-9][0-9]:[0-9][0-9]:[0-9][0-9]:[0-9][0-9]/)) {
ts = substr($0, RSTART + 13, 8)
}
# Anchor on the %TR/%Tw/%Tc/%Tr/%Ta timers block: %ST follows it, and
# the termination state (%tsc) is 4 fields further on (%B %CC %CS %tsc).
for (i = 1; i <= NF; i++) {
if ($i ~ /^[+-]?[0-9]+\/[+-]?[0-9]+\/[+-]?[0-9]+\/[+-]?[0-9]+\/[+-]?[0-9]+$/) {
status = $(i + 1)
term = $(i + 5)
break
}
}'
echo ""
}
# Only enforcement outcomes: tarpit (PT--), deny (PR--), 403, 429.
if (!(term ~ /^PT/ || term ~ /^PR/ || status == "403" || status == "429")) next
for (i = 1; i <= NF; i++) {
if (substr($i, 1, 4) == "cip=") cip = substr($i, 5)
if (substr($i, 1, 5) == "host=") host = substr($i, 6)
if (substr($i, 1, 3) == "id=") id = substr($i, 4)
}
if (match($0, /"[A-Z]+ [^"]*"/)) {
req = substr($0, RSTART + 1, RLENGTH - 2)
if (length(req) > 42) req = substr(req, 1, 41) "..."
}
if (cip == "") cip = "-"
if (host == "") host = "-"
if (term == "") term = "-"
if (status == "") status = "-"
printf "%-15s %-16s %-4s %-5s %-28s %s\n", ts, cip, status, term, host, req
if (id != "" && id != "-") printf "%-15s %s\n", "", " id=" id
n++
}
END { if (n == 0) print "(no tarpit/deny/403/429 events in the scanned window)" }
')
printf '%s\n' "$found"
echo
echo "TERM = HAProxy termination state: PT-- tarpit, PR-- deny (incl. WAF/rate limit)."
echo "id= = request reference; it is printed on the block page and is how a"
echo " customer support ticket correlates to an exact request here."
}
# Monitor mode selection
if [ "$1" == "live" ]; then
echo "Live monitoring mode - Press Ctrl+C to exit"
echo ""
banner() {
echo "==================================================="
echo "HAProxy Edge Monitor - $(date '+%Y-%m-%d %H:%M:%S')"
echo "==================================================="
echo
}
if [ "${1:-}" = "live" ]; then
echo "Live monitoring mode - Press Ctrl+C to exit"
while true; do
clear
echo "==================================================="
echo "HAProxy Security Monitor - $(date '+%Y-%m-%d %H:%M:%S')"
echo "==================================================="
echo ""
show_threats
echo ""
banner
show_rates || true
echo
show_recent_blocks
sleep 5
done
else
# Single run mode
show_threats
echo ""
banner
rc=0
show_rates || rc=$?
echo
show_recent_blocks
echo ""
echo "Tip: Run with 'live' parameter for continuous monitoring"
echo
echo "Tip: run with 'live' for a refreshing view."
echo "Usage: $0 [live]"
fi
# Propagate a stick-table read failure: if the rates section could not be
# produced, this run did NOT report what it claims to report.
exit "$rc"
fi
+314
View File
@@ -0,0 +1,314 @@
#!/usr/bin/env bash
#
# show-edge-ip-rates.sh — real, current per-IP rate counters from the HAProxy
# `web` stick table.
#
# WHAT THIS CAN TELL YOU
# The `web` stick table (templates/hap_listener.tpl) stores exactly four
# counters per client IP:
# conn_cur, conn_rate(10s), http_req_rate(10s), http_err_rate(30s)
# Those are INSTANTANEOUS values — the current concurrency and the current
# sliding-window rates. This script prints them, and nothing else.
#
# WHAT THIS CANNOT TELL YOU
# * Who has been tarpitted, denied, or rate-limited. The stick table stores
# NO history and NO counter of past enforcement actions. It has no gpc0 /
# gpc1 / gpc(N) / gpc_rate / glitch_rate columns at all — any tool that
# claims to read them from this table is fabricating numbers.
# * Anything about an IP that has gone quiet: entries expire after 10m.
#
# Real enforcement events live in the HAProxy ACCESS LOG, which is on the
# DOCKER HOST at /var/log/haproxy.log (it does NOT exist inside this
# container). The log-format carries the HAProxy termination state plus
# cip= (real client IP), host=, ua= and id= (the request UUID shown on the
# block page, which correlates with customer support tickets).
#
# Ready to run ON THE HOST:
# # last 20 tarpitted (PT--) or denied (PR--) requests
# grep -aE ' (PT|PR)--' /var/log/haproxy.log | tail -20
# # everything HAProxy answered 429/403 to, newest last
# grep -aE ' (429|403) ' /var/log/haproxy.log | tail -20
# # everything for one client IP
# grep -a 'cip=203.0.113.7' /var/log/haproxy.log | tail -50
# # look up one request reference from a support ticket
# grep -a 'id=<uuid-from-the-block-page>' /var/log/haproxy.log
#
# USAGE
# show-edge-ip-rates.sh [-a|--all]
# -a, --all also show rows whose counters are all zero (off by default:
# a table with hundreds of idle entries is pure noise)
#
# ENVIRONMENT
# SHOW_ALL=1 same as --all
# HAPROXY_SOCKET=<path> override the CLI socket (default /tmp/haproxy-cli)
# HAPROXY_TABLE_DUMP=<file>
# parse a previously captured `show table web` dump
# from a file instead of talking to the socket.
# Supported seam for offline analysis of a captured
# support bundle, and for testing this parser.
#
# NOTE ON THE SOCKET
# /tmp/haproxy-cli is HAProxy's MASTER CLI socket, so worker commands need an
# `@1` prefix. Without it HAProxy answers "Unknown command: 'show' ..." AND
# socat still exits 0 — so exit status is worthless here and this script
# inspects the RESPONSE BODY instead.
set -euo pipefail
SOCKET="${HAPROXY_SOCKET:-/tmp/haproxy-cli}"
TABLE="web"
# Fields this script expects the `web` stick table to store. Keep on ONE line
# in this exact NAME=(a b c d) shape — the contract test greps for it, and the
# parser below is driven entirely by it.
EXPECTED_FIELDS=(conn_cur conn_rate http_req_rate http_err_rate)
# Which of EXPECTED_FIELDS to sort on (descending). Falls back to the first
# field if this name is not in the list.
SORT_FIELD="http_req_rate"
SHOW_ALL="${SHOW_ALL:-0}"
while [ $# -gt 0 ]; do
case "$1" in
-a|--all) SHOW_ALL=1 ;;
-h|--help) sed -n '2,60p' "$0" | sed 's/^# \{0,1\}//'; exit 0 ;;
*) echo "Unknown argument: $1" >&2; echo "Usage: $0 [-a|--all]" >&2; exit 2 ;;
esac
shift
done
die() { echo "ERROR: $*" >&2; exit 1; }
# First non-blank line of a blob.
#
# Deliberately NOT `printf ... | sed -n '/./{p;q;}'`. sed quits after the first
# match and closes the pipe; on a real 550-entry table dump printf is still
# writing and takes SIGPIPE, so under `set -o pipefail` the whole command
# substitution returns 141 and `set -e` kills the script -- silently, with no
# output at all. That is the same class of failure this script exists to stop
# hiding, so it does not get to happen here. A plain read loop has no pipeline
# and no early close.
first_nonblank() {
local line
while IFS= read -r line || [ -n "$line" ]; do
case "$line" in
*[![:space:]]*) printf '%s\n' "$line"; return 0 ;;
esac
done <<EOF
$1
EOF
return 0
}
# Return 0 if the CLI response body is a rejection rather than table data.
# Checked on the body because socat's exit status is 0 either way.
body_is_rejected() {
local first
first=$(first_nonblank "$1")
case "$first" in
"Unknown command"*|"No such table"*|"Permission denied"*) return 0 ;;
*) return 1 ;;
esac
}
send_cmd() {
printf '%s\n' "$1" | socat stdio "$SOCKET" 2>/dev/null
}
# ---------------------------------------------------------------- fetch data
BODY=""
SOURCE=""
if [ -n "${HAPROXY_TABLE_DUMP:-}" ]; then
[ -r "$HAPROXY_TABLE_DUMP" ] || die "HAPROXY_TABLE_DUMP is set but '$HAPROXY_TABLE_DUMP' is not readable."
BODY=$(cat "$HAPROXY_TABLE_DUMP")
SOURCE="file $HAPROXY_TABLE_DUMP"
else
[ -S "$SOCKET" ] || die "HAProxy CLI socket not found at $SOCKET (is HAProxy running, and are you inside the haproxy-manager container?)"
command -v socat >/dev/null 2>&1 || die "socat is not installed; cannot talk to $SOCKET"
# Master socket form first, then the plain stats-socket form.
BODY=$(send_cmd "@1 show table $TABLE" || true)
SOURCE="socket $SOCKET (@1 show table $TABLE)"
if [ -z "${BODY//[[:space:]]/}" ] || body_is_rejected "$BODY"; then
FALLBACK=$(send_cmd "show table $TABLE" || true)
if [ -n "${FALLBACK//[[:space:]]/}" ] && ! body_is_rejected "$FALLBACK"; then
BODY="$FALLBACK"
SOURCE="socket $SOCKET (show table $TABLE)"
else
echo "ERROR: HAProxy rejected BOTH '@1 show table $TABLE' and 'show table $TABLE'." >&2
echo " @1 response : $(first_nonblank "$BODY")" >&2
echo " bare response: $(first_nonblank "$FALLBACK")" >&2
echo " Check the socket is HAProxy's CLI and that the table '$TABLE' exists" >&2
echo " (a config reload without the frontend would drop it)." >&2
exit 1
fi
fi
fi
# ------------------------------------------------------------- header checks
HEADER=$(first_nonblank "$BODY")
case "$HEADER" in
"# table: $TABLE,"*) : ;;
*)
echo "ERROR: unexpected first line from '$SOURCE'." >&2
echo " expected it to start with: # table: $TABLE," >&2
echo " got : $HEADER" >&2
exit 1
;;
esac
TBL_SIZE=$(printf '%s\n' "$HEADER" | sed -n 's/.*size:\([0-9]*\).*/\1/p')
TBL_USED=$(printf '%s\n' "$HEADER" | sed -n 's/.*used:\([0-9]*\).*/\1/p')
[ -n "$TBL_SIZE" ] || TBL_SIZE="?"
[ -n "$TBL_USED" ] || TBL_USED="?"
# ------------------------------------------------------------------- parsing
# awk emits:
# W \t <field>:<window-seconds-or-dash> ... (one line, from first row)
# R \t <sortkey> \t <ip> \t <value per EXPECTED_FIELDS in order>
# and exits 1 after reporting any row missing an expected field.
PARSED=""
if ! PARSED=$(printf '%s\n' "$BODY" | awk -v fieldlist="${EXPECTED_FIELDS[*]}" -v sortfield="$SORT_FIELD" '
BEGIN {
nf = split(fieldlist, F, " ")
sortidx = 1
for (i = 1; i <= nf; i++) if (F[i] == sortfield) sortidx = i
wprinted = 0
}
/^#/ { next }
!/key=/ { next }
{
split("", val, " "); split("", win, " ")
for (i = 1; i <= NF; i++) {
tok = $i
p = index(tok, "=")
if (p == 0) continue
lhs = substr(tok, 1, p - 1)
rhs = substr(tok, p + 1)
b = index(lhs, "(")
if (b > 0) {
nm = substr(lhs, 1, b - 1)
win[nm] = substr(lhs, b + 1, length(lhs) - b - 1)
} else {
nm = lhs
win[nm] = ""
}
val[nm] = rhs
}
missing = ""
for (i = 1; i <= nf; i++) if (!(F[i] in val)) missing = missing (missing == "" ? "" : ", ") F[i]
if (missing != "") {
printf "ERROR: stick table row is missing expected field(s): %s\n", missing > "/dev/stderr"
printf " offending row: %s\n", $0 > "/dev/stderr"
printf " this script expects the web table to store: %s\n", fieldlist > "/dev/stderr"
print " Those expectations and the templates/hap_listener.tpl `store` clause have DRIFTED." > "/dev/stderr"
print " Fix one or the other; refusing to print 0 for a counter HAProxy never reported." > "/dev/stderr"
exit 1
}
if (!("key" in val)) {
printf "ERROR: stick table row has no key= field: %s\n", $0 > "/dev/stderr"
exit 1
}
if (!wprinted) {
line = "W"
for (i = 1; i <= nf; i++) {
w = win[F[i]]
if (w ~ /^[0-9]+$/) w = sprintf("%g", w / 1000); else w = "-"
line = line "\t" F[i] ":" w
}
print line
wprinted = 1
}
nonzero = 0
row = ""
for (i = 1; i <= nf; i++) {
v = val[F[i]]
if (v + 0 != 0) nonzero = 1
row = row "\t" v
}
printf "R\t%s\t%s\t%d%s\n", val[F[sortidx]] + 0, val["key"], nonzero, row
}
'); then
exit 1
fi
# ------------------------------------------------------------------ printing
WINSPEC=$(printf '%s\n' "$PARSED" | sed -n 's/^W\t//p' || true)
echo "==================================================================="
echo " HAProxy edge IP rates — table '$TABLE' (current values only)"
echo "==================================================================="
echo "Source : $SOURCE"
echo "Tracked : ${TBL_USED} of ${TBL_SIZE} slots in use"
if [ "$SHOW_ALL" = "1" ]; then
echo "Filter : showing ALL tracked IPs"
else
echo "Filter : showing only IPs with a non-zero counter (use --all for every row)"
fi
echo
# Column headers, with each counter's window rendered in SECONDS (HAProxy
# reports the window in milliseconds, e.g. conn_rate(10000) = 10s).
HDR=$(printf "%-18s" "IP Address")
i=0
for f in "${EXPECTED_FIELDS[@]}"; do
w=$(printf '%s\n' "$WINSPEC" | tr '\t' '\n' | sed -n "s/^${f}://p")
if [ -n "$w" ] && [ "$w" != "-" ]; then
label="${f}/${w}s"
else
label="$f"
fi
HDR="$HDR $(printf '%18s' "$label")"
i=$((i + 1))
done
echo "$HDR"
printf '%s\n' "$HDR" | sed 's/./-/g'
ROWS=$(printf '%s\n' "$PARSED" | sed -n 's/^R\t//p' || true)
shown=0
if [ -n "$ROWS" ]; then
while IFS=$'\t' read -r sortkey ip nonzero rest; do
[ -n "${ip:-}" ] || continue
if [ "$SHOW_ALL" != "1" ] && [ "$nonzero" = "0" ]; then
continue
fi
line=$(printf "%-18s" "$ip")
oldifs="$IFS"; IFS=$'\t'
# shellcheck disable=SC2086
set -- $rest
IFS="$oldifs"
for v in "$@"; do
line="$line $(printf '%18s' "$v")"
done
echo "$line"
shown=$((shown + 1))
done < <(printf '%s\n' "$ROWS" | sort -t"$(printf '\t')" -k1,1nr)
fi
if [ "$shown" -eq 0 ]; then
if [ "$SHOW_ALL" = "1" ]; then
echo "(no IPs currently tracked)"
else
echo "(no IP currently has a non-zero counter — re-run with --all to list idle entries)"
fi
fi
echo
echo "==================================================================="
echo "These are CURRENT values. The table keeps no history and no record of"
echo "past tarpits/denials. For actual enforcement events, read the access"
echo "log ON THE DOCKER HOST (it does not exist in this container):"
echo " grep -aE ' (PT|PR)--' /var/log/haproxy.log | tail -20 # tarpit / deny"
echo " grep -aE ' (429|403) ' /var/log/haproxy.log | tail -20 # rate-limit / block"
echo " grep -a 'cip=<IP>' /var/log/haproxy.log | tail -50 # one client"
echo " grep -a 'id=<uuid>' /var/log/haproxy.log # one request reference"
echo
echo "Operator actions (via the MASTER CLI socket — the @1 prefix is required):"
echo " printf '@1 show table $TABLE key <IP>\\n' | socat stdio $SOCKET"
echo " printf '@1 set table $TABLE key <IP> data.http_req_rate 0\\n' | socat stdio $SOCKET"
echo " printf '@1 clear table $TABLE key <IP>\\n' | socat stdio $SOCKET # drop one entry"
echo " printf '@1 clear table $TABLE\\n' | socat stdio $SOCKET # drop ALL entries"
echo "==================================================================="
+30 -118
View File
@@ -1,123 +1,35 @@
#!/bin/bash
# Script to display IPs that have been tarpitted by HAProxy 3.0
# Uses HAProxy stats socket to query stick-table data
#!/usr/bin/env bash
#
# Usage in Docker container:
# docker exec -it haproxy-manager /haproxy/scripts/show-tarpit-ips.sh
# DEPRECATED SHIM — kept so existing docs/runbooks/muscle memory keep working.
#
# This script used to print a "Tarpitted IPs Report" with a "Scan Count" and a
# BLOCKED / SILENT-DROP / TARPIT status per IP, all derived from gpc0 and gpc1
# stick-table columns. Those columns DO NOT EXIST: the `web` table
# (templates/hap_listener.tpl) stores only conn_cur, conn_rate, http_req_rate
# and http_err_rate. The old parser defaulted every missing field to 0, so the
# whole report was fabricated — every IP showed "Scan Count 0 / Normal"
# regardless of what it was actually doing.
#
# The stick table also keeps NO history, so nothing in it can identify who was
# tarpitted. Real enforcement events live in the access log ON THE DOCKER HOST
# at /var/log/haproxy.log (it does not exist inside this container):
# grep -aE ' (PT|PR)--' /var/log/haproxy.log | tail -20 # tarpit / deny
# grep -aE ' (429|403) ' /var/log/haproxy.log | tail -20 # rate-limit / block
#
# What IS knowable from the stick table — the current per-IP rates — is printed
# by show-edge-ip-rates.sh, which this shim now runs.
SOCKET="/tmp/haproxy-cli"
set -euo pipefail
# Check if socket exists
if [ ! -S "$SOCKET" ]; then
echo "Error: HAProxy socket not found at $SOCKET"
echo "Make sure HAProxy is running with stats socket enabled"
exit 1
fi
cat >&2 <<'NOTE'
NOTE: show-tarpit-ips.sh is deprecated and cannot report tarpits.
The HAProxy stick table stores no history and no gpc0/gpc1 counters, so
the old "Scan Count"/"BLOCKED" columns were fabricated numbers.
Actual tarpit/deny events are in /var/log/haproxy.log ON THE HOST:
grep -aE ' (PT|PR)--' /var/log/haproxy.log | tail -20
Running show-edge-ip-rates.sh instead (current rates, real values):
echo "==================================================================="
echo " HAProxy Tarpitted IPs Report "
echo "==================================================================="
echo
echo "Showing IPs tracked in the stick-table with scan detection counters:"
echo "(gpc0 = total scan attempts, gpc1 = escalation level)"
echo
NOTE
# In HAProxy 3.0, we need to use the proper process prefix
# The web frontend table is in the worker process, not master
# First check which process has the table
# Note: grep for actual worker line, not the header
PROCESS_ID=$(echo "show proc" | socat stdio "$SOCKET" 2>/dev/null | grep -E '^[0-9]+.*worker' | awk '{print $1}' | head -1)
if [ -z "$PROCESS_ID" ]; then
echo "Error: Could not find HAProxy worker process"
echo "Try: echo 'show proc' | socat stdio $SOCKET"
exit 1
fi
# Show stick-table entries from the web frontend using the worker process
# Use printf to avoid bash history expansion issues with !
printf "@!%s show table web\n" "${PROCESS_ID}" | socat stdio "$SOCKET" 2>/dev/null | {
# Skip the header line
read header
# Check if we got an error or empty response
if echo "$header" | grep -q "No such table"; then
echo "Error: Table 'web' not found. HAProxy may need to be reloaded."
exit 1
fi
has_data=false
echo "IP Address | Scan Count | Level | HTTP Err Rate | Status"
echo "---------------------|------------|-------|---------------|------------------"
# Process each line
while IFS= read -r line; do
# Skip empty lines and comments
if [ -z "$line" ] || echo "$line" | grep -q "^#"; then
continue
fi
# HAProxy 3.0 format: 0x... key=<ip> use=... exp=... gpc0=... gpc1=... http_err_rate(10s)=...
if echo "$line" | grep -q "key="; then
has_data=true
# Extract IP and counters
ip=$(echo "$line" | grep -o 'key=[^ ]*' | cut -d'=' -f2)
gpc0=$(echo "$line" | grep -o 'gpc0=[0-9]*' | cut -d'=' -f2)
gpc1=$(echo "$line" | grep -o 'gpc1=[0-9]*' | cut -d'=' -f2)
err_rate=$(echo "$line" | grep -o 'http_err_rate([^)]*=[0-9]*' | grep -o '[0-9]*$')
# Set defaults if values are empty
gpc0=${gpc0:-0}
gpc1=${gpc1:-0}
err_rate=${err_rate:-0}
# Determine status based on scan count and escalation
status=""
if [ "$gpc0" -ge 100 ]; then
status="BLOCKED (429)"
elif [ "$gpc0" -ge 60 ]; then
status="SILENT-DROP"
elif [ "$gpc0" -ge 40 ]; then
if [ "$gpc1" -ge 2 ]; then
status="SILENT-DROP (repeat)"
else
status="TARPIT 10s"
fi
elif [ "$gpc0" -ge 25 ]; then
status="TARPIT 10s"
else
status="Normal"
fi
# Format output
printf "%-20s | %10s | %5s | %13s | %s\n" "$ip" "$gpc0" "$gpc1" "$err_rate/10s" "$status"
fi
done
if [ "$has_data" = false ]; then
echo "(No IPs currently tracked - table is empty)"
fi
}
echo
echo "==================================================================="
echo "Legend:"
echo " - Scan Count 25-39: Low scanner → TARPIT 10s delay"
echo " - Scan Count 40-59: Medium scanner → TARPIT 10s (1st), SILENT-DROP (repeat)"
echo " - Scan Count 60-99: High scanner → SILENT-DROP (immediate disconnect)"
echo " - Scan Count 100+: Critical scanner → BLOCKED (429 response)"
echo " - Burst (5+ in 10s): → TARPIT 10s (1st), SILENT-DROP (repeat)"
echo "==================================================================="
echo "Note: Only counts suspicious scripts/configs, NOT missing images/fonts/CSS"
echo "Note: IPs are tracked for 1 hour since last activity"
echo
echo "To clear a specific IP from the table:"
echo " printf '@!${PROCESS_ID} del table web key <IP>\\n' | socat stdio $SOCKET"
echo
echo "To clear all entries:"
echo " printf '@!${PROCESS_ID} clear table web\\n' | socat stdio $SOCKET"
echo
echo "Debug: Worker PID is ${PROCESS_ID}"
echo
SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
exec "$SCRIPT_DIR/show-edge-ip-rates.sh" "$@"
+33 -2
View File
@@ -21,8 +21,39 @@ set -eo pipefail
mkdir -p /etc/haproxy
[ -f /etc/haproxy/trusted_ips.list ] || : > /etc/haproxy/trusted_ips.list
[ -f /etc/haproxy/trusted_ips.map ] || : > /etc/haproxy/trusted_ips.map
[ -f /etc/haproxy/cloudflare_ips.list ] || : > /etc/haproxy/cloudflare_ips.list
[ -f /etc/haproxy/trusted_proxies.list ] || : > /etc/haproxy/trusted_proxies.list
# cloudflare_ips.list is SHIPPED DATA: it must always match what this image
# bakes in (/haproxy/defaults), so a Cloudflare range refresh actually reaches
# existing hosts instead of being permanently shadowed by the volume.
# Overwrite it from the baked copy on every start.
#
# trusted_proxies.list and wpadmin_gate_exempt.list are OPERATOR DATA:
# operators add entries directly on the server and those must survive
# restarts/recreates. Seed each from the baked copy only when it's missing;
# never overwrite an existing one.
#
# All branches fall back to an empty file if the baked default is somehow
# absent, because "acl ... -f <missing file>" is a fatal HAProxy config
# error -- the list files must exist unconditionally by the time HAProxy starts.
if [ -f /haproxy/defaults/cloudflare_ips.list ]; then
cp /haproxy/defaults/cloudflare_ips.list /etc/haproxy/cloudflare_ips.list
else
[ -f /etc/haproxy/cloudflare_ips.list ] || : > /etc/haproxy/cloudflare_ips.list
fi
if [ ! -f /etc/haproxy/trusted_proxies.list ]; then
if [ -f /haproxy/defaults/trusted_proxies.list ]; then
cp /haproxy/defaults/trusted_proxies.list /etc/haproxy/trusted_proxies.list
else
: > /etc/haproxy/trusted_proxies.list
fi
fi
if [ ! -f /etc/haproxy/wpadmin_gate_exempt.list ]; then
if [ -f /haproxy/defaults/wpadmin_gate_exempt.list ]; then
cp /haproxy/defaults/wpadmin_gate_exempt.list /etc/haproxy/wpadmin_gate_exempt.list
else
: > /etc/haproxy/wpadmin_gate_exempt.list
fi
fi
cron &
+180
View File
@@ -27,6 +27,8 @@ These tests pin the invariants:
* nothing is published that is not a complete, validated cert+key pair;
* no old certificate file is removed and no lineage deleted until the
replacement is validated, in place, and actually loaded by HAProxy;
* removing ONE domain never unlinks a .pem, or deletes a certbot lineage,
that other still-configured domains are being served from;
* only final .pem files ever exist in the crt directory.
Running
@@ -894,6 +896,184 @@ class TestBundleValidation(CertPublishTestCase):
'a private key must never be group/world writable')
class TestSharedCertificateSurvivesDomainRemoval(CertPublishTestCase):
"""Bug 5: DELETE /api/domain unlinked a PEM other live sites were served from.
`domains.domain` is UNIQUE; `domains.ssl_cert_path` is not, and nothing
ever made it so. request_ssl_bundle() issues one SAN certificate and points
every included name's row at the same /etc/haproxy/certs/<primary>.pem, so
sharing is not an edge case - it is the normal shape of the table. Measured
on production the day this was written: 39 of 71 distinct cert paths on one
host were referenced by more than one domain row (118 of 150 SSL-enabled
rows), and one shared path was
/etc/haproxy/certs/threeworldsoneheart.org.pem, referenced by the live
apex, its www, and a mail.* alias.
remove_domain() did, unconditionally:
os.remove(ssl_cert_path)
certbot delete --cert-name <domain>
Removing the mail.* alias would therefore have deleted the PEM the apex was
serving on, and (had the alias been the bundle primary) the lineage behind
it - HTTPS down for every other name in the bundle, with no local copy and
only a fresh, rate-limited ACME order to recover from.
These tests assert file-system and certbot-invocation outcomes, not which
branch was taken.
"""
BUNDLE = '/etc/haproxy/certs' # documentation only; SSL_CERTS_DIR is stubbed
def _cert_for(self, primary):
"""Publish a bundle for `primary` and return its path."""
return self.publish_live_bundle(primary)
def remove(self, domain):
return self.client.delete('/api/domain', json={'domain': domain})
# -- last reference: the cleanup must still happen --------------------
def test_last_reference_removal_unlinks_the_pem(self):
cert = self._cert_for('example.com')
self.add_domain('example.com', 'be_example', ssl_cert_path=cert)
resp = self.remove('example.com')
self.assertEqual(200, resp.status_code, resp.data)
self.assertFalse(
os.path.exists(cert),
'nothing else referenced this bundle - it must be cleaned up, or '
'the crt directory accumulates certs for domains that are gone')
def test_last_reference_removal_deletes_the_lineage(self):
cert = self._cert_for('example.com')
self.add_domain('example.com', 'be_example', ssl_cert_path=cert)
self.remove('example.com')
self.assertEqual(
['delete --cert-name example.com --non-interactive'],
self.certbot_deletes(),
'the last name on a lineage went away - the lineage should go too')
def test_lineage_deleted_is_the_bundles_not_the_domains_own_name(self):
"""Removing a SAN member must target the lineage that issued the file."""
cert = self._cert_for('example.com')
self.add_domain('www.example.com', 'be_www', ssl_cert_path=cert)
self.remove('www.example.com')
self.assertEqual(
['delete --cert-name example.com --non-interactive'],
self.certbot_deletes(),
'the lineage is named after the bundle primary (--cert-name), not '
'after whichever SAN happened to be removed last')
# -- shared reference: nothing may be destroyed -----------------------
def test_shared_pem_survives_removal_of_one_name(self):
cert = self._cert_for('example.com')
before = self.read(cert)
self.add_domain('example.com', 'be_apex', ssl_cert_path=cert)
self.add_domain('www.example.com', 'be_www', ssl_cert_path=cert)
resp = self.remove('www.example.com')
self.assertEqual(200, resp.status_code, resp.data)
self.assertTrue(
os.path.exists(cert),
'example.com is still configured and still served from this file')
self.assertEqual(before, self.read(cert),
'the surviving bundle must be byte-for-byte intact')
self.assertTrue(self.edge_would_start(),
'the edge must still load the crt directory')
def test_shared_pem_survives_removal_of_the_bundle_primary(self):
"""The worst shape: the name being removed IS the lineage/file name."""
cert = self._cert_for('example.com')
before = self.read(cert)
self.add_domain('example.com', 'be_apex', ssl_cert_path=cert)
self.add_domain('www.example.com', 'be_www', ssl_cert_path=cert)
self.remove('example.com')
self.assertTrue(
os.path.exists(cert),
'www.example.com is still configured and is served from this exact '
'file - removing the apex must not unlink it')
self.assertEqual(before, self.read(cert))
self.assertTrue(self.edge_would_start())
def test_shared_lineage_is_not_certbot_deleted(self):
cert = self._cert_for('example.com')
self.add_domain('example.com', 'be_apex', ssl_cert_path=cert)
self.add_domain('www.example.com', 'be_www', ssl_cert_path=cert)
self.remove('example.com')
self.assertEqual(
[], self.certbot_deletes(),
'certbot delete destroys archive, live and renewal config for a '
'lineage that is still renewing the certificate www.example.com '
'is served with')
def test_production_shape_mail_alias_removal(self):
"""The exact row that stopped a production cleanup.
mail.threeworldsoneheart.org carried
ssl_cert_path=/etc/haproxy/certs/threeworldsoneheart.org.pem - the live
PEM of a different, serving site.
"""
cert = self._cert_for('threeworldsoneheart.org')
before = self.read(cert)
self.add_domain('threeworldsoneheart.org', 'be_apex', ssl_cert_path=cert)
self.add_domain('www.threeworldsoneheart.org', 'be_www', ssl_cert_path=cert)
self.add_domain('mail.threeworldsoneheart.org', 'be_mail', ssl_cert_path=cert)
self.remove('mail.threeworldsoneheart.org')
self.assertTrue(os.path.exists(cert),
'two sites are still served from this bundle')
self.assertEqual(before, self.read(cert))
self.assertEqual([], self.certbot_deletes())
self.assertTrue(self.edge_would_start())
def test_removing_every_name_eventually_cleans_up(self):
"""The guard defers cleanup, it does not cancel it."""
cert = self._cert_for('example.com')
self.add_domain('example.com', 'be_apex', ssl_cert_path=cert)
self.add_domain('www.example.com', 'be_www', ssl_cert_path=cert)
self.remove('www.example.com')
self.assertTrue(os.path.exists(cert), 'apex still needs it')
self.remove('example.com')
self.assertFalse(os.path.exists(cert),
'the last reference is gone - now it may be removed')
self.assertEqual(
['delete --cert-name example.com --non-interactive'],
self.certbot_deletes())
def test_unrelated_domains_certificate_is_untouched(self):
"""Sharing is by exact path; different bundles stay independent."""
mine = self._cert_for('example.com')
theirs = self._cert_for('other.example')
theirs_before = self.read(theirs)
self.add_domain('example.com', 'be_mine', ssl_cert_path=mine)
self.add_domain('other.example', 'be_theirs', ssl_cert_path=theirs)
self.remove('example.com')
self.assertFalse(os.path.exists(mine))
self.assertTrue(os.path.exists(theirs),
'a different bundle must not be collateral damage')
self.assertEqual(theirs_before, self.read(theirs))
self.assertEqual(
['delete --cert-name example.com --non-interactive'],
self.certbot_deletes())
@unittest.skipIf(TESTING_FOREIGN_TREE,
'HAPROXY_MANAGER_DIR points at another tree')
class TestPublisherApiIsPresent(unittest.TestCase):
+397
View File
@@ -0,0 +1,397 @@
#!/usr/bin/env python3
"""Contract test: the runtime-map fast path (blocked IPs).
Why this file exists
--------------------
`add_ip_to_runtime_map()` and `remove_ip_from_runtime_map()` spent their whole
existence sending
add map #0 <ip> 1
del map #0 <ip>
to `/tmp/haproxy-cli` and returning True whenever socat exited 0. Neither
command has ever worked. Two independent defects:
* **No `@1` prefix.** `/tmp/haproxy-cli` is HAProxy's MASTER CLI socket; map
commands are worker commands. The master answers `Unknown command: 'add',
but maybe one of the following ones is a better match: ...` -- and **socat
still exits 0**, so `result.returncode == 0` was true and the function
logged "Added IP x to runtime map".
* **`#0` is not a valid map id.** Ids are assigned at config-parse time and
move on every config regeneration; on the live edge `blocked_ips.map` is
id 37 and `trusted_ips.map` is 10. There is no id 0. Any hardcoded number
is wrong -- the map must be referenced by its FILE PATH, which is what
haproxy.cfg itself names in `map_ip(/etc/haproxy/blocked_ips.map,0)`.
And a third silence that makes a response-body check alone insufficient:
`@1 add map #0 <ip> 1` returns an **empty body**, exit 0, and adds nothing
anywhere -- while `@1 del map #0 <ip>` answers `Unknown map identifier.`. The
add path can therefore only be trusted after reading the entry back.
IP blocking still worked, because `update_blocked_ips_map()` rewrites
`/etc/haproxy/blocked_ips.map` and the callers reload HAProxy, which re-reads
it. The FILE is authoritative; this is the no-reload fast path, and it has
never once run while reporting that it did.
What it enforces
----------------
1. The command strings actually sent: `@1` prefix first, map referenced by
PATH and never by `#<id>`, and the value `1` that
`map_ip(...,0) -m int gt 0` requires.
2. Every captured rejection is classified as FAILURE (returns False), not
success -- including the two that carry no error text at all.
3. Success is only reported when the entry reads back in the state asked
for. Not the exit status, not an empty reply.
4. No source in this repo builds a map command with a `#<id>` reference.
Comments may describe the old form; code may not use it.
Runs fully offline: `_cli_send()` is replaced, so no socket, no socat, no
HAProxy, no network.
Running
-------
python3 scripts/test-runtime-map-contract.py
"""
import ast
import io
import os
import re
import sys
import glob
import logging
import tempfile
import unittest
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda filename, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(filename)), *a, **kw)
)
import haproxy_manager # noqa: E402
logging.getLogger().setLevel(logging.CRITICAL)
MAP = haproxy_manager.BLOCKED_IPS_MAP_PATH
IP = '192.0.2.77'
# ---------------------------------------------------------------------------
# Responses captured VERBATIM from the haproxy-manager container on the live
# edge (HAProxy 3.0.11, 2026-08-22). socat exited 0 for every single one of
# them, which is the entire reason none of this is decided on exit status.
# ---------------------------------------------------------------------------
# What the MASTER socket answers to an unprefixed `add map ...` -- i.e. the
# reply the old code read as success.
MASTER_REJECTS_ADD = """\
Unknown command: 'add', but maybe one of the following ones is a better match:
@!<pid> : send a command to the <pid> process
@master : send a command to the master process
hard-reload : achieve a hard-reload (-st) of haproxy
reload : achieve a soft-reload (-sf) of haproxy
user : lower the level of the current CLI session to user
help [<command>] : list matching or all commands
prompt [timed] : toggle interactive mode with prompt
quit : disconnect
"""
MASTER_REJECTS_DEL = MASTER_REJECTS_ADD.replace("'add'", "'del'")
# `@1 del map #0 <ip>` / `@1 get map /etc/haproxy/nope.map <ip>`.
UNKNOWN_MAP_IDENTIFIER = 'Unknown map identifier. Please use #<id> or <file>.\n'
# `@1 add map <map> <ip>` with the value omitted (as the old docs showed).
ADD_MAP_MISSING_VALUE = (
"'add map' expects three parameters (map identifier, key and value) or one "
"parameter (map identifier) and a payload\n")
# `@1 del map <map> <ip>` for a key the runtime map does not hold.
KEY_NOT_FOUND = 'Key not found.\n'
# A successful mutation. THIS IS THE WHOLE PROBLEM: it is byte-for-byte what
# `@1 add map #0 <ip> 1` also returns while adding nothing at all.
MUTATION_OK = ''
GET_FOUND = ('type=ip, case=sensitive, found=yes, idx=tree, key="%s", '
'value="1", type="str"\n' % IP)
GET_NOT_FOUND = 'type=ip, case=sensitive, found=no\n'
# `@1 get map #0 <ip>` -- "found", but with no value. haproxy.cfg matches with
# `-m int gt 0`, so an entry like this does NOT block.
GET_FOUND_NO_VALUE = ('type=ip, case=sensitive, found=yes, idx=tree, key="%s", '
'value=none\n' % IP)
SHOW_MAP_SAMPLE = (
'0x7f6e5a788700 101.36.109.130 1\n'
'0x7f6e5a788780 101.47.140.218 1\n'
'0x7f6e5a1fd500 %s 1\n' % IP)
class FakeSocket(object):
"""Replaces _cli_send(). Records every command, answers from a script.
`script` is a list of (substring, response) pairs, consulted in order; the
first whose substring appears in the command wins. Anything unmatched is
an explicit test bug, not a silent default.
"""
def __init__(self, script):
self.script = script
self.sent = []
def __call__(self, command, socket_path, timeout):
self.sent.append(command)
for needle, response in self.script:
if needle in command:
return response
raise AssertionError('test script has no response for %r' % command)
def with_socket(script):
"""Install a FakeSocket for the duration of a `with` block."""
class _Ctx(object):
def __enter__(self):
self.fake = FakeSocket(script)
self._real = haproxy_manager._cli_send
haproxy_manager._cli_send = self.fake
return self.fake
def __exit__(self, *exc):
haproxy_manager._cli_send = self._real
return False
return _Ctx()
# Every command the happy path needs, in the order the code issues them.
HAPPY_ADD = [('get map', GET_FOUND), ('add map', MUTATION_OK)]
HAPPY_DEL = [('get map', GET_NOT_FOUND), ('del map', MUTATION_OK)]
class CommandsAreWellFormed(unittest.TestCase):
"""Guard 1: the exact bytes on the wire.
Both original defects are visible here and nowhere else -- a `#0` map
reference and a missing `@1` are perfectly ordinary-looking Python.
"""
def test_add_sends_worker_prefixed_path_referenced_command(self):
with with_socket(HAPPY_ADD) as fake:
self.assertTrue(haproxy_manager.add_ip_to_runtime_map(IP))
self.assertEqual(fake.sent[0], '@1 add map %s %s 1' % (MAP, IP))
def test_del_sends_worker_prefixed_path_referenced_command(self):
with with_socket(HAPPY_DEL) as fake:
self.assertTrue(haproxy_manager.remove_ip_from_runtime_map(IP))
self.assertEqual(fake.sent[0], '@1 del map %s %s' % (MAP, IP))
def test_every_command_is_tried_on_the_worker_first(self):
"""`/tmp/haproxy-cli` is the MASTER socket; bare map commands 404."""
for run in (lambda: haproxy_manager.add_ip_to_runtime_map(IP),
lambda: haproxy_manager.remove_ip_from_runtime_map(IP)):
with with_socket(HAPPY_ADD + HAPPY_DEL) as fake:
run()
for command in fake.sent:
self.assertTrue(command.startswith('@1 '),
'%r is missing the @1 worker prefix' % command)
def test_no_command_references_a_map_by_id(self):
"""Map ids move on every config regeneration. Path, always."""
script = HAPPY_ADD + HAPPY_DEL + [('show map', SHOW_MAP_SAMPLE)]
with with_socket(script) as fake:
haproxy_manager.add_ip_to_runtime_map(IP)
haproxy_manager.remove_ip_from_runtime_map(IP)
haproxy_manager.runtime_map_keys(MAP)
for command in fake.sent:
self.assertNotRegex(
command, r'\bmap\s+#',
'%r references a map by id; ids are not stable' % command)
self.assertIn(MAP, command,
'%r does not name the map file' % command)
def test_add_carries_the_value_the_config_matches_on(self):
"""`map_ip(...,0) -m int gt 0`: a valueless entry does not block."""
self.assertEqual(haproxy_manager.BLOCKED_IPS_MAP_VALUE, '1')
with with_socket(HAPPY_ADD) as fake:
haproxy_manager.add_ip_to_runtime_map(IP)
self.assertTrue(fake.sent[0].endswith(' %s 1' % IP),
'%r has no value; HAProxy rejects it' % fake.sent[0])
class RejectionIsFailure(unittest.TestCase):
"""Guard 2+3: nothing may report success unless the map really changed.
Every response below was returned by the live socket with **exit code 0**.
The old code returned True for all of them.
"""
def test_master_socket_rejection_of_add_is_failure(self):
with with_socket([('get map', GET_NOT_FOUND), ('add map', MASTER_REJECTS_ADD)]):
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
def test_master_socket_rejection_of_del_is_failure(self):
with with_socket([('get map', GET_FOUND), ('del map', MASTER_REJECTS_DEL)]):
self.assertFalse(haproxy_manager.remove_ip_from_runtime_map(IP))
def test_unknown_map_identifier_is_failure(self):
with with_socket([('get map', GET_NOT_FOUND), ('add map', UNKNOWN_MAP_IDENTIFIER)]):
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
with with_socket([('get map', GET_FOUND), ('del map', UNKNOWN_MAP_IDENTIFIER)]):
self.assertFalse(haproxy_manager.remove_ip_from_runtime_map(IP))
def test_missing_value_rejection_is_failure(self):
"""Carries no marker word at all -- caught by 'a mutation says nothing'."""
with with_socket([('get map', GET_NOT_FOUND), ('add map', ADD_MAP_MISSING_VALUE)]):
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
def test_silent_noop_add_is_failure(self):
"""The `#0` failure mode: accepted, empty reply, nothing added.
Nothing in the response distinguishes this from success. Only the
read-back does -- which is why the read-back is not optional.
"""
with with_socket([('get map', GET_NOT_FOUND), ('add map', MUTATION_OK)]):
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
def test_add_that_lands_without_a_value_is_failure(self):
with with_socket([('get map', GET_FOUND_NO_VALUE), ('add map', MUTATION_OK)]):
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
def test_del_that_leaves_the_key_behind_is_failure(self):
with with_socket([('get map', GET_FOUND), ('del map', MUTATION_OK)]):
self.assertFalse(haproxy_manager.remove_ip_from_runtime_map(IP))
def test_key_not_found_on_del_is_the_requested_end_state(self):
"""Not a failure: the runtime map already lacks the key."""
with with_socket([('get map', GET_NOT_FOUND), ('del map', KEY_NOT_FOUND)]):
self.assertTrue(haproxy_manager.remove_ip_from_runtime_map(IP))
def test_verified_success_is_reported_as_success(self):
with with_socket(HAPPY_ADD):
self.assertTrue(haproxy_manager.add_ip_to_runtime_map(IP))
with with_socket(HAPPY_DEL):
self.assertTrue(haproxy_manager.remove_ip_from_runtime_map(IP))
def test_a_runtime_failure_never_raises_into_the_request_handler(self):
"""The map file + reload still enforces the block; degrade, don't 500."""
with with_socket([('get map', GET_NOT_FOUND), ('add map', MASTER_REJECTS_ADD)]):
self.assertIs(haproxy_manager.add_ip_to_runtime_map(IP), False)
def test_mutations_are_checked_for_an_empty_body_not_a_marker_list(self):
"""_HAPROXY_CLI_ERROR_MARKERS can only know rejections already seen."""
with self.assertRaises(haproxy_manager.HaproxyCliError):
with with_socket([('add map', 'something nobody has ever seen\n')]):
haproxy_manager.haproxy_cli('add map %s x 1' % MAP,
worker=True, expect_empty=True)
def test_the_new_error_markers_are_recognised(self):
for response in (UNKNOWN_MAP_IDENTIFIER, KEY_NOT_FOUND,
MASTER_REJECTS_ADD):
self.assertTrue(haproxy_manager._cli_response_is_error(response),
'%r must be classified as an error' % response[:40])
for response in (GET_FOUND, GET_NOT_FOUND, SHOW_MAP_SAMPLE):
self.assertFalse(haproxy_manager._cli_response_is_error(response),
'%r is data, not an error' % response[:40])
def test_socat_exit_zero_carries_no_information(self):
"""FakeSocket never signals failure any other way, and neither did socat."""
with with_socket([('get map', GET_NOT_FOUND), ('add map', MASTER_REJECTS_ADD)]) as fake:
self.assertFalse(haproxy_manager.add_ip_to_runtime_map(IP))
self.assertTrue(fake.sent, 'the command was sent and "succeeded" at the '
'process level; only the body says otherwise')
class ReadBack(unittest.TestCase):
"""The read-back primitives the guarantees above rest on."""
def test_lookup_reports_found_with_value(self):
with with_socket([('get map', GET_FOUND)]):
self.assertEqual(haproxy_manager.runtime_map_lookup(MAP, IP),
(True, '1'))
def test_lookup_reports_not_found(self):
with with_socket([('get map', GET_NOT_FOUND)]):
self.assertEqual(haproxy_manager.runtime_map_lookup(MAP, IP),
(False, None))
def test_lookup_raises_on_a_rejected_reference(self):
with with_socket([('get map', UNKNOWN_MAP_IDENTIFIER)]):
with self.assertRaises(haproxy_manager.HaproxyCliError):
haproxy_manager.runtime_map_lookup(MAP, IP)
def test_keys_parses_show_map_output(self):
with with_socket([('show map', SHOW_MAP_SAMPLE)]):
self.assertEqual(
haproxy_manager.runtime_map_keys(MAP),
{'101.36.109.130', '101.47.140.218', IP})
def test_empty_map_is_not_a_rejection(self):
with with_socket([('show map', '')]):
self.assertEqual(haproxy_manager.runtime_map_keys(MAP), set())
def test_rejected_show_map_still_raises(self):
with with_socket([('show map', UNKNOWN_MAP_IDENTIFIER)]):
with self.assertRaises(haproxy_manager.HaproxyCliError):
haproxy_manager.runtime_map_keys(MAP)
MAP_BY_ID_RE = re.compile(r'\b(?:add|del|clear|show|get)\s+map\s+#')
class NoSourceBuildsAMapIdCommand(unittest.TestCase):
"""Guard 4: `map #<id>` may be described in comments, never executed.
Scanning string literals rather than raw text is deliberate -- the whole
reason this bug is documented at length in the source is so the next reader
does not reintroduce it, and a plain grep would fail on those comments.
"""
def test_no_python_string_literal_builds_a_map_id_command(self):
tree = ast.parse(io.open('haproxy_manager.py', encoding='utf-8').read())
offenders = [
node.value for node in ast.walk(tree)
if isinstance(node, ast.Constant) and isinstance(node.value, str)
and MAP_BY_ID_RE.search(node.value)
and not (ast.get_docstring(tree) == node.value)
]
# Docstrings are string literals too; exclude any literal that is a
# docstring of a module/class/function.
docstrings = set()
for node in ast.walk(tree):
if isinstance(node, (ast.Module, ast.ClassDef, ast.FunctionDef,
ast.AsyncFunctionDef)):
doc = ast.get_docstring(node, clean=False)
if doc:
docstrings.add(doc)
offenders = [o for o in offenders if o not in docstrings]
self.assertEqual(offenders, [],
'these string literals build a map command with an '
'unstable #<id> reference')
def test_no_shell_or_template_code_line_uses_a_map_id(self):
targets = (glob.glob('scripts/*.sh') + glob.glob('templates/*.tpl')
+ ['Dockerfile'])
offenders = []
for path in targets:
if not os.path.exists(path):
continue
for lineno, line in enumerate(
io.open(path, encoding='utf-8').read().splitlines(), 1):
if line.lstrip().startswith('#'):
continue # a comment describing the old form is fine
if MAP_BY_ID_RE.search(line):
offenders.append('%s:%d: %s' % (path, lineno, line.strip()))
self.assertEqual(offenders, [],
'map ids are assigned at config-parse time and move; '
'reference the map file by path')
if __name__ == '__main__':
unittest.main(verbosity=2)
+475
View File
@@ -0,0 +1,475 @@
#!/usr/bin/env python3
"""Contract test: what the stick tables STORE vs what the consumers READ.
Why this file exists
--------------------
`/api/security/stats` and `scripts/show-tarpit-ips.sh` spent their whole
existence reporting "Scan Count", "offense count" and "BLOCKED" figures parsed
out of `gpc0` and `gpc1`. **No stick table in this repo has ever stored a
general-purpose counter.** Every one of those numbers was fabricated, and an
operator was making decisions on them.
Nothing caught it, because each layer failed silently in a different way:
* `int(parts[3])` on a positional split hit `exp=368842`, raised ValueError,
and the loop `continue`d -- so the endpoint answered `active_threats: 0`
with an empty list. "No threats" and "the parser is broken" looked
identical.
* The command went to the MASTER CLI socket without the `@1` worker prefix.
HAProxy answered `Unknown command: 'show', but maybe one of the following
ones is a better match: ...` and **socat still exited 0**, so the
`returncode != 0` guard never fired. The reported `total_tracked_ips` was
the line count of that help text (8) while the real table held 388 entries.
* The shell consumers wrote `gpc0=${gpc0:-0}`, so a field that does not exist
rendered as a confident zero.
The durable fix is not "parse better" -- it is making the template and its
consumers unable to drift apart without something going red. That is this file.
What it enforces
----------------
1. `STICK_TABLE_FIELD_CONTRACT` in haproxy_manager.py equals, exactly and in
both directions, the `store` clauses in the rendered templates.
2. Every shell consumer's `EXPECTED_FIELDS=(...)` array equals the contract.
3. A captured sample of REAL `show table` output from the live edge parses to
exactly the contract's fields plus the entry metadata -- so the contract
describes reality, not just itself.
4. The loud-failure behaviour: `read_stick_table()` RAISES on a rejected
command, on a non-table response, and on a row missing a contract field.
It must never answer zeros. Guard 4 is the one that would have caught the
original bug on day one.
5. No `store` clause names a general-purpose counter, and no template tracks
one -- the state this repo is actually in, asserted rather than assumed.
Assertions about the TEMPLATES go through `rule_lines()`, which strips comments
before matching. These templates quote their own rules in prose at length; a
bare `assertIn` over the rendered text passes just as happily against a rule
that has been commented out. Same lesson, and same helper, as
scripts/test-wpadmin-gate.py.
Runs fully offline -- no HAProxy, no socket, no network.
Running
-------
python3 scripts/test-stick-table-contract.py
"""
import os
import re
import sys
import logging
import tempfile
import unittest
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda filename, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(filename)), *a, **kw)
)
import haproxy_manager # noqa: E402
# ---------------------------------------------------------------------------
# A captured sample of REAL output, so this test can assert against reality
# without a live socket.
#
# Provenance: `echo "@1 show table web" | socat stdio /tmp/haproxy-cli` inside
# the haproxy-manager container on whp01, 2026-08-22, HAProxy 3.0.11. Rows
# trimmed for length; format byte-for-byte as emitted.
#
# Note what is and is NOT here: no gpc0, no gpc1, no gpc(N), no gpc_rate, no
# glitch_rate. Note also that field windows come back in MILLISECONDS
# (`conn_rate(10000)`), not the `10s` the template writes -- a consumer that
# labels the raw number "10s" is off by a factor of 1000.
# ---------------------------------------------------------------------------
LIVE_TABLE_SAMPLE = """\
# table: web, type: ip, size:204800, used:388
0x7f6e5447e728: key=17.58.57.102 use=0 exp=368842 shard=0 conn_rate(10000)=0 conn_cur=0 http_req_rate(10000)=0 http_err_rate(30000)=0
0x7f6e541cf7c8: key=43.173.182.9 use=0 exp=171404 shard=0 conn_rate(10000)=0 conn_cur=0 http_req_rate(10000)=0 http_err_rate(30000)=0
0x7f6e5547f528: key=95.129.255.180 use=0 exp=590639 shard=0 conn_rate(10000)=1 conn_cur=0 http_req_rate(10000)=1 http_err_rate(30000)=0
0x7f6e5447e488: key=5.9.105.254 use=0 exp=556277 shard=0 conn_rate(10000)=0 conn_cur=0 http_req_rate(10000)=0 http_err_rate(30000)=1
"""
# The verbatim reply the MASTER CLI socket gives to an unprefixed worker
# command. Captured the same way. socat exits 0 on this -- which is the entire
# reason haproxy_cli() inspects the body.
MASTER_SOCKET_REJECTION = """\
Unknown command: 'show', but maybe one of the following ones is a better match:
show cli level : display the level of the current CLI session
show cli sockets : dump list of cli sockets
show proc : show processes status
show startup-logs : report logs emitted during HAProxy startup
show version : show version of the current process
help [<command>] : list matching or all commands
prompt [timed] : toggle interactive mode with prompt
quit : disconnect
"""
# Anything matching these in a `store` clause is a general-purpose counter.
GPC_PATTERN = re.compile(r'\bgpc|glitch')
# Shell consumers that must declare their field expectations as a single
# EXPECTED_FIELDS array, and the table each one reads.
SHELL_CONSUMERS = {
'scripts/show-edge-ip-rates.sh': 'web',
}
EXPECTED_FIELDS_RE = re.compile(r'^\s*EXPECTED_FIELDS=\(([^)]*)\)\s*$', re.M)
def rule_lines(cfg, needle):
"""Comment-stripped config lines containing `needle`.
A line that is entirely a comment is dropped; a line mixing config with a
trailing comment is truncated at the first ' #' before matching. Without
this, every assertion below would pass against a rule that had been
commented out but whose text survived in the surrounding prose -- and these
templates quote their own rules in prose constantly. See
scripts/test-wpadmin-gate.py, where a mutation audit proved the point.
"""
out = []
for raw in cfg.split('\n'):
stripped = raw.strip()
if stripped and not stripped.startswith('#'):
code = stripped.split(' #', 1)[0].rstrip()
if code and needle in code:
out.append(code)
return out
def render_listener():
return haproxy_manager.template_env.get_template('hap_listener.tpl').render(
crt_path='/etc/haproxy/certs',
suspension_enabled=False,
coraza_spoe_backend=None,
)
def render_security_tables():
return haproxy_manager.template_env.get_template('hap_security_tables.tpl').render()
def stick_tables_from(cfg):
"""{table name: (field, ...)} for every stick-table declared in `cfg`.
A stick-table takes the name of the frontend/backend/listen section that
declares it -- that name is what `show table <name>` wants, so the section
header is part of the contract, not incidental. Section headers and
stick-table lines are both read comment-stripped.
"""
tables = {}
section = None
for raw in cfg.split('\n'):
stripped = raw.strip()
if not stripped or stripped.startswith('#'):
continue
code = stripped.split(' #', 1)[0].rstrip()
header = re.match(r'^(frontend|backend|listen)\s+(\S+)', code)
if header:
section = header.group(2)
continue
if code.startswith('stick-table'):
store = re.search(r'\bstore\s+(\S+)', code)
if not store:
raise AssertionError(
'stick-table in section %r has no `store` clause: %r'
% (section, code))
if section is None:
raise AssertionError(
'stick-table declared outside any section: %r' % code)
# store is a comma-separated list; each item is name or name(window)
fields = tuple(item.split('(')[0]
for item in store.group(1).split(','))
tables[section] = fields
return tables
def all_template_stick_tables():
tables = {}
for cfg in (render_listener(), render_security_tables()):
for name, fields in stick_tables_from(cfg).items():
if name in tables:
raise AssertionError('stick table %r declared twice' % name)
tables[name] = fields
return tables
class StickTableContract(unittest.TestCase):
"""Guard 1 + 5: the templates and the Python contract, held together."""
def setUp(self):
self.templates = all_template_stick_tables()
self.contract = haproxy_manager.STICK_TABLE_FIELD_CONTRACT
def test_templates_actually_declare_stick_tables(self):
"""Guard the guard: an empty parse would make every other check vacuous."""
self.assertTrue(self.templates,
'parsed no stick tables out of the templates at all -- '
'stick_tables_from() is broken, not the templates')
def test_same_table_names(self):
self.assertEqual(
sorted(self.templates), sorted(self.contract),
'STICK_TABLE_FIELD_CONTRACT and the templates disagree on WHICH '
'stick tables exist. Add/remove the table in both places.')
def test_same_fields_per_table(self):
for table in sorted(self.templates):
with self.subTest(table=table):
self.assertEqual(
sorted(self.templates[table]),
sorted(self.contract.get(table, ())),
"stick table %r stores %s but STICK_TABLE_FIELD_CONTRACT "
"claims %s. Whichever is wrong, a consumer is about to read "
"a field that is never populated -- which is the bug this "
"test exists for." % (table,
list(self.templates[table]),
list(self.contract.get(table, ()))))
def test_web_table_is_the_one_the_api_reads(self):
self.assertIn('web', self.contract)
self.assertIn('web', self.templates)
def test_no_general_purpose_counters_are_stored(self):
"""The state of the world today, asserted rather than assumed.
If a gpc/glitch counter is ever genuinely added to a template, this
test is the place to update -- and updating it forces whoever does so
to also add the field to STICK_TABLE_FIELD_CONTRACT (test_same_fields_
per_table) and to the shell consumers (test_shell_consumers_match_
contract). That chain is the point: a counter cannot appear in a
consumer without existing in the table, and cannot appear in the table
without the consumers being updated.
"""
for table, fields in self.templates.items():
for field in fields:
self.assertIsNone(
GPC_PATTERN.search(field),
'stick table %r now stores %r. Update this test, '
'STICK_TABLE_FIELD_CONTRACT, and every consumer.'
% (table, field))
def test_track_sc_counters_have_a_table_each(self):
"""Every `track-scN ... table X` names a table that really exists.
A typo here is invisible to `haproxy -c` only in the sense that it is
NOT -- but it is invisible to the consumers, which would query a table
that is never written.
"""
cfg = render_listener()
for line in rule_lines(cfg, 'track-sc'):
named = re.search(r'\btable\s+(\S+)', line)
if named:
self.assertIn(
named.group(1), self.templates,
'track-sc rule references undeclared table %r: %r'
% (named.group(1), line))
def test_sc_counter_indices_fit_haproxys_limit(self):
"""sc0/sc1/sc2 are all HAProxy gives us by default.
`tune.stick-counters` defaults to 3. A `track-sc3` without raising it
is a config-time failure, and the templates' own comments assume the
limit -- so assert it rather than leaving it as folklore.
"""
cfg = render_listener() + '\n' + render_security_tables()
raised = rule_lines(cfg, 'tune.stick-counters')
limit = 3
if raised:
limit = int(re.search(r'(\d+)', raised[-1]).group(1))
for line in rule_lines(cfg, 'track-sc'):
idx = int(re.search(r'track-sc(\d+)', line).group(1))
self.assertLess(
idx, limit,
'track-sc%d exceeds tune.stick-counters (%d): %r'
% (idx, limit, line))
class ShellConsumerContract(unittest.TestCase):
"""Guard 2: the shell consumers cannot drift from the contract."""
def test_shell_consumers_match_contract(self):
for path, table in sorted(SHELL_CONSUMERS.items()):
with self.subTest(script=path):
full = os.path.join(MODULE_DIR, path)
self.assertTrue(os.path.exists(full),
'%s is missing; it is a declared consumer of '
'stick table %r' % (path, table))
with open(full) as fh:
src = fh.read()
m = EXPECTED_FIELDS_RE.search(src)
self.assertIsNotNone(
m, '%s must declare its field expectations once as a '
'single-line `EXPECTED_FIELDS=(a b c)` array so this '
'test can hold it to the template' % path)
declared = sorted(m.group(1).split())
self.assertEqual(
declared,
sorted(haproxy_manager.STICK_TABLE_FIELD_CONTRACT[table]),
'%s reads %s but stick table %r stores %s'
% (path, declared, table,
sorted(haproxy_manager.STICK_TABLE_FIELD_CONTRACT[table])))
def test_retired_script_no_longer_parses_phantom_counters(self):
"""show-tarpit-ips.sh may explain gpc0/gpc1; it may not extract them.
The shim is allowed -- encouraged -- to name the fields in prose so an
operator who runs it learns why its numbers went away. What it must not
do is go back to pulling values out of them.
"""
path = os.path.join(MODULE_DIR, 'scripts/show-tarpit-ips.sh')
if not os.path.exists(path):
self.skipTest('show-tarpit-ips.sh has been removed outright')
with open(path) as fh:
lines = fh.readlines()
for raw in lines:
stripped = raw.strip()
if not stripped or stripped.startswith('#'):
continue
code = stripped.split(' #', 1)[0]
self.assertIsNone(
re.search(r"grep -o ['\"]?gpc|gpc[0-9]*=\$|sc_get_gpc|sc-inc-gpc", code),
'show-tarpit-ips.sh is extracting a general-purpose counter '
'again: %r' % stripped)
class LiveSampleParses(unittest.TestCase):
"""Guard 3: the contract describes real HAProxy output, not just itself."""
def test_header_parses(self):
header, entries = self._read()
self.assertEqual(header['name'], 'web')
self.assertEqual(header['type'], 'ip')
self.assertEqual(header['size'], 204800)
self.assertEqual(header['used'], 388)
self.assertEqual(len(entries), 4)
def test_sample_fields_are_exactly_contract_plus_metadata(self):
expected = set(haproxy_manager.STICK_TABLE_FIELD_CONTRACT['web'])
expected |= set(haproxy_manager.STICK_TABLE_ENTRY_META)
_, entries = self._read()
for line, fields in entries:
self.assertEqual(
set(fields), expected,
'real `show table web` output carries %s, contract+metadata '
'expects %s. Row: %r'
% (sorted(fields), sorted(expected), line))
def test_key_is_the_ip_not_the_allocation_pointer(self):
"""The original bug read parts[0] -- the `0x...:` pointer -- as the IP."""
_, entries = self._read()
ips = [f['key']['value'] for _, f in entries]
self.assertIn('95.129.255.180', ips)
for ip in ips:
self.assertFalse(ip.startswith('0x'),
'parsed a memory address as an IP: %r' % ip)
def test_windows_are_milliseconds(self):
"""HAProxy reports `conn_rate(10000)` for a `conn_rate(10s)` store.
Asserted because labelling that raw 10000 as "10s" (or as seconds) is
an easy and completely silent way to be wrong by 1000x in the panel.
"""
_, entries = self._read()
_, fields = entries[0]
self.assertEqual(fields['conn_rate']['window_ms'], 10000)
self.assertEqual(fields['http_req_rate']['window_ms'], 10000)
self.assertEqual(fields['http_err_rate']['window_ms'], 30000)
self.assertIsNone(fields['conn_cur']['window_ms'],
'conn_cur is a gauge, not a rate; it has no window')
def test_values_are_the_real_ones(self):
_, entries = self._read()
by_ip = {f['key']['value']: f for _, f in entries}
self.assertEqual(by_ip['95.129.255.180']['http_req_rate']['value'], '1')
self.assertEqual(by_ip['5.9.105.254']['http_err_rate']['value'], '1')
self.assertEqual(by_ip['17.58.57.102']['http_req_rate']['value'], '0')
def _read(self):
return _read_table_from(LIVE_TABLE_SAMPLE)
class FailsLoudly(unittest.TestCase):
"""Guard 4: every way this can go wrong must raise, never return zeros.
This is the guard that would have caught the original bug immediately. Each
case below is a real response the old code accepted silently.
"""
def test_master_socket_rejection_is_not_data(self):
"""The exact reply that used to be reported as `total_tracked_ips: 8`."""
with self.assertRaises(haproxy_manager.HaproxyCliError) as ctx:
_read_table_from(MASTER_SOCKET_REJECTION)
self.assertIn('not a stick-table dump', str(ctx.exception))
def test_socat_exit_zero_does_not_mean_success(self):
"""haproxy_cli() must reject on the BODY, not the exit status.
socat returns 0 for every response above -- the rejection is only ever
visible in the text.
"""
self.assertTrue(
haproxy_manager._cli_response_is_error(MASTER_SOCKET_REJECTION))
self.assertTrue(
haproxy_manager._cli_response_is_error('No such table\n'))
self.assertTrue(
haproxy_manager._cli_response_is_error('Permission denied\n'))
self.assertFalse(
haproxy_manager._cli_response_is_error(LIVE_TABLE_SAMPLE))
def test_missing_contract_field_raises_and_names_it(self):
"""A field the table stopped storing must not silently become 0."""
degraded = LIVE_TABLE_SAMPLE.replace(' http_err_rate(30000)=0', '')
with self.assertRaises(haproxy_manager.HaproxyCliError) as ctx:
_read_table_from(degraded)
msg = str(ctx.exception)
self.assertIn('http_err_rate', msg,
'the error must name the missing field')
self.assertIn('drifted', msg,
'the error must say what actually went wrong')
def test_empty_response_raises(self):
with self.assertRaises(haproxy_manager.HaproxyCliError):
_read_table_from('')
def test_row_without_key_raises(self):
broken = LIVE_TABLE_SAMPLE.replace('key=17.58.57.102 ', '')
with self.assertRaises(haproxy_manager.HaproxyCliError) as ctx:
_read_table_from(broken)
self.assertIn('no key=', str(ctx.exception))
def test_unknown_table_raises_before_touching_the_socket(self):
"""Querying a table with no contract is a programming error, not a 0."""
with self.assertRaises(haproxy_manager.HaproxyCliError) as ctx:
haproxy_manager.read_stick_table('does_not_exist')
self.assertIn('no field contract', str(ctx.exception))
def test_empty_table_is_not_an_error(self):
"""A table with zero entries is a legitimate, distinguishable result."""
header, entries = _read_table_from(
'# table: web, type: ip, size:204800, used:0\n')
self.assertEqual(header['used'], 0)
self.assertEqual(entries, [])
def _read_table_from(response, table='web'):
"""Run read_stick_table() against a canned response instead of a socket."""
real = haproxy_manager.haproxy_cli
haproxy_manager.haproxy_cli = lambda cmd, worker=False, timeout=None: response
try:
return haproxy_manager.read_stick_table(table)
finally:
haproxy_manager.haproxy_cli = real
if __name__ == '__main__':
unittest.main(verbosity=2)
+530
View File
@@ -0,0 +1,530 @@
#!/usr/bin/env python3
"""Regression tests for the WordPress admin edge gate in hap_listener.tpl.
Why this file exists
--------------------
Unauthenticated GETs to /wp-admin/* were reaching PHP, booting WordPress just to
produce a login redirect and exhausting lsphp pools under a distributed
low-and-slow attack. The gate redirects them at the edge instead.
Two properties are easy to get wrong and invisible to `haproxy -c`:
* ORDERING. `is_whitelisted` reads var(txn.real_ip). If the rule renders before
the set-var chain, the whitelist evaluates against an unset variable.
* THE ALLOWLIST. wp-login.php loads its OWN css/js from /wp-admin/. Dropping
those entries leaves every login page on the fleet unstyled, while still
returning 200 -- a silent regression.
A THIRD property, added after an adversarial mutation audit: every assertion
here must be scoped to the CODE, not the surrounding prose. This file's own
comment blocks quote ACL names, rule fragments and even whole rules to explain
them -- which means a bare `assertIn` / `re.search` / `str.index` run over the
raw rendered config passes just as happily when the real rule has been deleted
(or merely commented out) and only its explanation survives. The audit proved
this concretely: commenting out the entire redirect rule, or the
`acl wp_admin_allowed` line, or all five normalizers, left the previous version
of this file at 26/26 PASS. See `rule_lines()` below, and use it (or one of the
guarded helpers built on it) for every assertion about whether a rule exists,
what it says, or where it sits relative to another rule. Do not add a new
`self.cfg.index(...)`, `self.assertIn(x, self.cfg)`, or `re.search(pattern,
self.cfg)` to this file -- none of them can tell code from comment.
Running
-------
python3 scripts/test-wpadmin-gate.py
"""
import os
import re
import sys
import logging
import tempfile
import unittest
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda filename, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(filename)), *a, **kw)
)
import haproxy_manager # noqa: E402
ALLOWLIST = ('/admin-ajax.php', '/admin-post.php',
'/load-styles.php', '/load-scripts.php')
EXCLUSIONS = ('!wp_admin_allowed', '!wp_admin_asset', '!has_wp_logged_in',
'!wp_gate_exempt', '!is_local', '!is_trusted_ip', '!is_whitelisted')
# The normalizer set, in the order it MUST render. Decoding has to precede the
# path walkers or "%2e%2e" is decoded to ".." only after path-strip-dotdot has
# already run, leaving the ".." unresolved -- measured against real HAProxy
# 3.0.11, both orders side by side.
NORMALIZERS = ('percent-to-uppercase',
'percent-decode-unreserved',
'path-merge-slashes',
'path-strip-dot',
'path-strip-dotdot full')
def render_listener():
return haproxy_manager.template_env.get_template('hap_listener.tpl').render(
crt_path='/etc/haproxy/certs',
suspension_enabled=False,
coraza_spoe_backend=None,
)
def render_header():
return haproxy_manager.template_env.get_template('hap_header.tpl').render(
cluster_secret=None,
)
# ---------------------------------------------------------------------------
# Comment-safe config inspection
#
# Every helper below operates on `rule_positions()`'s output, never on the raw
# rendered string. That is the one rule this whole module exists to enforce on
# itself: a mutation audit found that commenting out a real rule (prefixing it
# with '#', or -- more slyly -- deleting it and appending its own text as a
# TRAILING comment on the line above) left the previous version of these tests
# fully green, because plain `str.index` / `assertIn` / `re.search` over
# `self.cfg` cannot distinguish code from a comment that merely quotes it.
# ---------------------------------------------------------------------------
def rule_positions(cfg, needle):
"""[(comment-stripped line, char offset in cfg)] for every non-comment
line containing `needle`, in document order.
Two things a bare substring/regex search over `self.cfg` gets wrong, both
fixed here:
1. A line that is ENTIRELY a comment (starts with '#' once stripped) is
dropped. This is necessary but not sufficient -- see (2).
2. A line that MIXES real config with a trailing comment
(`live-code # note`, or the decoy `live-code # was: <the other
rule's exact text>`) is truncated at the first ' #' before matching,
so text stuffed into a trailing comment cannot masquerade as the
rule itself. A real HAProxy comment always starts at a '#' preceded
by whitespace here -- none of these templates use bare '#' as a
value character -- so this truncation does not clip real rules.
Returning (line, position) pairs together -- rather than making callers
re-derive one from the other with a second `cfg.index(line)` -- also
avoids a subtler bug: if the same comment-stripped line occurs twice
(e.g. a duplicated rule), re-deriving the position with `str.index` always
finds the FIRST copy regardless of which one you meant. Walking the file
once and recording positions as we go keeps first/last unambiguous.
"""
out = []
pos = 0
for raw in cfg.split('\n'):
stripped = raw.strip()
if stripped and not stripped.startswith('#'):
code = stripped.split(' #', 1)[0].rstrip()
if code and needle in code:
out.append((code, pos))
pos += len(raw) + 1 # +1 for the '\n' split() consumed
return out
def rule_lines(cfg, needle):
"""Comment-stripped rule lines containing `needle` (text only, no
position). See `rule_positions()` for what this guards against. Every
assertion about a RULE's presence or content must go through this (or
`rule_positions`/the guarded helpers below) -- never a bare
`needle in cfg` or `re.search(pattern, cfg)`.
"""
return [line for line, _ in rule_positions(cfg, needle)]
def require_rule(cfg, needle, what=None):
"""The single rule line containing `needle`.
Raises a plain AssertionError naming what was being looked for -- not
IndexError from an unguarded `rule_lines(...)[0]`, and not
`ValueError: substring not found` from a bare `cfg.index(...)` -- when
the rule is missing. A missing rule and a broken test harness must not
look identical in a failure report.
Raises the same way if `needle` is ambiguous (matches more than one rule
line): silently taking the first match in that case would hide the
ambiguity instead of surfacing it.
"""
label = what or needle
lines = rule_lines(cfg, needle)
if not lines:
raise AssertionError('no rule found for %r (expected: %s)' % (needle, label))
if len(lines) > 1:
raise AssertionError(
'%r matched %d rule lines, expected exactly one (%s): %r'
% (needle, len(lines), label, lines))
return lines[0]
def require_position(cfg, needle, what=None, last=False):
"""(line, char offset) for an ordering assertion, guarded the same way as
`require_rule` -- but tolerant of the needle matching multiple lines
(e.g. a multi-line set-var "chain"), since ordering checks often want the
first or last of several. Pass last=True for the last occurrence.
"""
label = what or needle
positions = rule_positions(cfg, needle)
if not positions:
raise AssertionError(
'no rule found for %r, cannot check ordering (expected: %s)' % (needle, label))
return positions[-1] if last else positions[0]
class WpAdminGate(unittest.TestCase):
def setUp(self):
self.cfg = render_listener()
def test_wp_admin_path_acl_declared(self):
line = require_rule(self.cfg, 'acl wp_admin_path', 'wp_admin_path ACL')
self.assertIn('path_reg', line)
def test_wp_admin_asset_acl_declared(self):
line = require_rule(self.cfg, 'acl wp_admin_asset', 'wp_admin_asset ACL')
self.assertIn('path_reg', line)
def test_wp_admin_allowed_acl_declared(self):
line = require_rule(self.cfg, 'acl wp_admin_allowed', 'wp_admin_allowed ACL')
self.assertIn('path_end', line)
def test_wp_gate_exempt_acl_declared(self):
"""Scoped to the ACL line itself, not `self.cfg` as a whole -- the
surrounding prose (see hap_listener.tpl's "per-site opt-out" comment)
also spells out /etc/haproxy/wpadmin_gate_exempt.list verbatim, so an
unscoped `assertIn` would still pass with the real ACL deleted.
"""
line = require_rule(self.cfg, 'acl wp_gate_exempt', 'wp_gate_exempt ACL')
self.assertIn('/etc/haproxy/wpadmin_gate_exempt.list', line)
def test_allowlist_entries_present(self):
"""wp-login.php loads its own css/js from /wp-admin/ -- see module docstring."""
line = require_rule(self.cfg, 'acl wp_admin_allowed', 'wp_admin_allowed ACL')
for entry in ALLOWLIST:
with self.subTest(entry=entry):
self.assertIn(entry, line)
def test_static_asset_dirs_allowed(self):
line = require_rule(self.cfg, 'acl wp_admin_asset', 'wp_admin_asset ACL')
self.assertRegex(line, r'path_reg.*\(css\|js\|images\)')
def test_install_php_is_NOT_allowlisted(self):
"""install.php is deliberately gated -- a takeover vector on abandoned installs."""
line = require_rule(self.cfg, 'acl wp_admin_allowed', 'wp_admin_allowed ACL')
self.assertNotIn('install.php', line)
def test_redirect_rule_has_all_exclusions(self):
rule = require_rule_by_predicate(
self.cfg, 'http-request redirect', lambda ln: 'wp_admin_path' in ln,
'wp-admin redirect rule')
for excl in EXCLUSIONS:
with self.subTest(exclusion=excl):
self.assertIn(excl, rule)
def test_rule_renders_after_real_ip_resolution(self):
"""is_whitelisted reads txn.real_ip; before the set-var chain it is unset."""
_, setvar_pos = require_position(self.cfg, 'set-var(txn.real_ip)',
'real_ip set-var chain')
_, rule_pos = require_position(self.cfg, 'wp_admin_path', 'wp_admin_path ACL/rule')
self.assertLess(setvar_pos, rule_pos,
'wp_admin_path renders before txn.real_ip is resolved')
def test_rule_renders_after_has_wp_logged_in_declared(self):
"""HAProxy resolves ACLs as it parses; use-before-declare fails."""
_, decl_pos = require_position(self.cfg, 'acl has_wp_logged_in',
'has_wp_logged_in ACL declaration')
_, rule_pos = require_position(self.cfg, 'wp_admin_path', 'wp_admin_path ACL/rule')
self.assertLess(decl_pos, rule_pos,
'wp_admin_path renders before has_wp_logged_in is declared')
def test_only_one_has_wp_logged_in_declaration(self):
self.assertEqual(len(rule_lines(self.cfg, 'acl has_wp_logged_in')), 1)
def test_allowlist_entries_are_anchored_to_wp_admin(self):
"""Bare `path_end /admin-ajax.php` also matches
/wp-admin/evil/admin-ajax.php, which ALSO matches wp_admin_path
(path_reg only requires /wp-admin/ to appear somewhere) -- an
attacker-inserted path segment would then sail through the
allowlist ungated. Entries must be anchored to sit directly under
wp-admin/. Scoped to the captured ACL line only, since the
surrounding comment block also mentions these bare filenames.
"""
line = require_rule(self.cfg, 'acl wp_admin_allowed', 'wp_admin_allowed ACL')
for entry in ALLOWLIST:
with self.subTest(entry=entry):
self.assertIn('/wp-admin' + entry, line)
self.assertNotRegex(
line, r'(?<!wp-admin)' + re.escape(entry) + r'(?!\S)',
'found a bare, unanchored allowlist entry: ' + entry)
def test_wp_login_url_setvar_renders_in_correct_order(self):
"""The inline regsub-in-`location` form is rejected by real HAProxy
3.0.11 (invalid arg 2 in converter 'regsub'), so the regsub is
computed in its own set-var line instead. That set-var must render
after the set-var(txn.real_ip) chain (it must not disturb that
load-bearing chain) and before the redirect rule that consumes it.
"""
_, last_real_ip_pos = require_position(self.cfg, 'set-var(txn.real_ip)',
'real_ip set-var chain', last=True)
_, wp_login_pos = require_position(self.cfg, 'set-var(txn.wp_login_url)',
'wp_login_url set-var')
_, redirect_pos = require_position(
self.cfg, 'http-request redirect code 302 location %[var(txn.wp_login_url)]',
'wp-admin redirect rule')
self.assertLess(last_real_ip_pos, wp_login_pos,
'wp_login_url set-var must render after the real_ip set-var chain')
self.assertLess(wp_login_pos, redirect_pos,
'wp_login_url set-var must render before the redirect rule that uses it')
def test_redirect_rule_uses_setvar_not_inline_regsub(self):
"""Guards against reintroducing the rejected inline form."""
rule = require_rule_by_predicate(
self.cfg, 'http-request redirect', lambda ln: 'wp_admin_path' in ln,
'wp-admin redirect rule')
self.assertIn('%[var(txn.wp_login_url)]', rule)
self.assertNotIn('regsub', rule)
def test_redirect_rule_requires_safe_path(self):
"""OPEN REDIRECT guard. The redirect target is built by rewriting
`path` with regsub, which only replaces the matched substring --
everything before "/wp-admin/" survives untouched in the output.
Three concrete requests turn that into an off-site `Location:`
header: "//evil.example.com/wp-admin/x.php" (protocol-relative,
browsers resolve "//host/path" to "https://host/path"),
"/\\evil.example.com/wp-admin/x.php" (browsers normalise a leading
"/\\" the same as "//"), and an RFC 7230 absolute-form request
target ("https://evil.example.com/wp-admin/x.php") which can make
HAProxy's `path` fetch return a full URI. wp_admin_safe_path
(requiring a well-formed absolute path) must be a POSITIVE
condition on the redirect rule -- scoped to the captured rule line
only, since the surrounding comment block also mentions this ACL
name and a bare substring match would pass even if the condition
were dropped from the rule itself.
"""
rule = require_rule_by_predicate(
self.cfg, 'http-request redirect', lambda ln: 'wp_admin_path' in ln,
'wp-admin redirect rule')
self.assertIn('wp_admin_safe_path', rule)
self.assertNotIn('!wp_admin_safe_path', rule,
'wp_admin_safe_path must be a positive condition, not negated')
def test_wp_admin_safe_path_acl_declared(self):
line = require_rule(self.cfg, 'acl wp_admin_safe_path', 'wp_admin_safe_path ACL')
self.assertIn('path_reg', line)
def test_unsafe_wp_admin_path_is_denied_not_passed_through(self):
"""wp_admin_safe_path being a POSITIVE condition on the redirect means
a path that fails it is simply not redirected -- which used to mean it
fell through to the backend UNGATED, i.e. exactly the PHP-booting
request the gate exists to stop. Normalisation removes the "//"
spelling of that, but not "/\\", so the fall-through must be closed
with an explicit deny rather than left implicit.
"""
denies = rule_lines(self.cfg, '!wp_admin_safe_path')
self.assertTrue(denies,
'no rule denies a wp-admin path that fails wp_admin_safe_path')
self.assertTrue(any(d.startswith('http-request deny') for d in denies),
'the !wp_admin_safe_path rule must be a deny: %r' % denies)
def require_rule_by_predicate(cfg, needle, predicate, what):
"""Like require_rule(), but for rules identified by needle + a predicate
over the comment-stripped line (e.g. "the `http-request redirect` line
that also mentions wp_admin_path", since the frontend has more than one
`http-request redirect`). Raises a clear AssertionError, not IndexError
or a silently-empty match, if no line satisfies both.
"""
candidates = [ln for ln in rule_lines(cfg, needle) if predicate(ln)]
if not candidates:
raise AssertionError('no rule found matching %s' % what)
if len(candidates) > 1:
raise AssertionError('%s matched more than one rule line: %r' % (what, candidates))
return candidates[0]
class UriNormalisation(unittest.TestCase):
"""The gate matches the RAW path; the backend normalises and decodes it.
Every gap between those is a bypass -- five were found this way. These
tests pin the normalisation that closes the gap as a class.
NOTE: these are config-TEXT assertions. They are necessary but NOT
sufficient: the previous revision of this file passed while five live
bypasses shipped. The real evidence is the behavioural matrix run against
real haproxy 3.0.11 with raw sockets -- see
.superpowers/sdd/2026-08-14-wpadmin-edge-gate/task-4-normalize-report.md.
"""
def setUp(self):
self.cfg = render_listener()
self.header = render_header()
def test_experimental_directives_exposed_in_global(self):
"""normalize-uri is experimental in 3.0; without this HAProxy refuses
to start (`haproxy -c` exits 1, ALERT). That does NOT crash-loop the
container, though -- see hap_header.tpl's comment and
haproxy_manager.py's start_haproxy()/do_initial_setup(): the failure
is swallowed, and the container comes up with haproxy simply never
running. This test exists so that silent-outage mode is never
reintroduced by dropping this line.
"""
lines = rule_lines(self.header, 'expose-experimental-directives')
matches = [ln for ln in lines if ln == 'expose-experimental-directives']
self.assertTrue(
matches, 'expose-experimental-directives missing from the global section')
def test_all_normalizers_render(self):
for norm in NORMALIZERS:
with self.subTest(normalizer=norm):
self.assertTrue(
rule_lines(self.cfg, 'normalize-uri ' + norm),
'missing normalizer: ' + norm)
def test_normalizer_order_decode_before_path_walkers(self):
"""Reverse this order and /wp-admin/js/%2e%2e/plugins.php reaches the
ACLs as /wp-admin/js/../plugins.php -- decoded but unresolved.
"""
lines = [ln for ln in rule_lines(self.cfg, 'http-request normalize-uri')
if ln.startswith('http-request normalize-uri')]
names = [ln.split('normalize-uri ', 1)[1] for ln in lines]
self.assertEqual(
names, list(NORMALIZERS),
'normalize-uri directives are missing, reordered, or duplicated: %r' % names)
def test_normalisation_precedes_every_path_based_rule(self):
"""A normalizer placed after a path rule normalises nothing for it."""
norm_positions = []
for n in NORMALIZERS:
_, pos = require_position(self.cfg, 'http-request normalize-uri ' + n,
'normalizer: ' + n)
norm_positions.append(pos)
last_norm = max(norm_positions)
for marker in ('acl is_health_check', 'acl wp_login_path',
'acl xmlrpc_path', 'acl wp_batch_path',
'acl wp_admin_path', 'http-request set-path'):
with self.subTest(rule=marker):
_, marker_pos = require_position(self.cfg, marker, marker)
self.assertLess(last_norm, marker_pos,
marker + ' renders before URI normalisation')
def test_query_sort_by_name_is_not_enabled(self):
"""query-sort-by-name reorders query parameters, which would break
anything that signs or caches on the exact query string. This is NOT
because the enabled normalizers already leave the query alone --
percent-to-uppercase and percent-decode-unreserved rewrite the WHOLE
request-target, query string included (see hap_listener.tpl's BLAST
RADIUS comment) -- it is a deliberate line between "case-fold /
decode" (no-ops under RFC 3986) and "reorder" (not a no-op for a
signed/cached query string).
"""
self.assertFalse(rule_lines(self.cfg, 'normalize-uri query-sort-by-name'))
def test_dotdot_normalizer_uses_full(self):
"""Without "full", ".." segments that climb above the root are left in
place and /../../wp-admin/plugins.php survives -- measured.
"""
lines = rule_lines(self.cfg, 'normalize-uri path-strip-dotdot')
self.assertTrue(lines, 'missing normalizer: path-strip-dotdot')
for ln in lines:
self.assertTrue(ln.endswith('path-strip-dotdot full'), ln)
def test_encoded_separator_on_wp_admin_is_denied(self):
"""percent-decode-unreserved deliberately leaves %2F encoded ("/" is
reserved), but OpenLiteSpeed decodes it and serves the file --
/wp-admin%2Fplugins.php was measured booting PHP on the OLS tier while
matching no wp-admin ACL. Normalisation cannot close this; it needs its
own rule.
"""
denies = [ln for ln in rule_lines(self.cfg, 'path_has_encoded_sep')
if ln.startswith('http-request deny')]
self.assertTrue(denies, 'no deny rule for encoded separators')
acl_line = require_rule(self.cfg, 'acl path_has_encoded_sep', 'path_has_encoded_sep ACL')
self.assertIn('%2f', acl_line.lower())
def test_encoded_separator_deny_is_scoped_to_wp_admin(self):
"""A blanket "deny any %2F in any path" would break non-WordPress
customer apps that legitimately pass an encoded slash in a path
parameter. The deny must be conditioned on the path mentioning
wp-admin.
"""
denies = [ln for ln in rule_lines(self.cfg, 'path_has_encoded_sep')
if ln.startswith('http-request deny')]
self.assertTrue(denies, 'no deny rule for encoded separators')
for d in denies:
self.assertIn('wp_admin_word', d)
def test_encoded_separator_deny_honors_the_same_whitelist(self):
denies = [ln for ln in rule_lines(self.cfg, 'path_has_encoded_sep')
if ln.startswith('http-request deny')]
self.assertTrue(denies, 'no deny rule for encoded separators')
for d in denies:
for excl in ('!has_wp_logged_in', '!wp_gate_exempt', '!is_local',
'!is_trusted_ip', '!is_whitelisted'):
with self.subTest(rule=d, exclusion=excl):
self.assertIn(excl, d)
def test_encoded_separator_acl_matches_a_substring_not_a_prefix(self):
"""/blog%2Fwp-admin/plugins.php hides the separator BEFORE "wp-admin",
where an anchored pattern never matches, and OLS still resolves it.
"""
acl_line = require_rule(self.cfg, 'acl path_has_encoded_sep', 'path_has_encoded_sep ACL')
self.assertIn('-m sub', acl_line)
def test_wp_admin_asset_bypass_cannot_cover_a_php_entrypoint(self):
"""The asset bypass anchored its prefix but not its suffix, so
/wp-admin/css/../plugins.php took it and the backend then resolved
".." and booted plugins.php. path-strip-dotdot is the real fix; this
keeps the bypass structurally incapable of covering PHP.
"""
acl_line = require_rule(self.cfg, 'acl wp_admin_asset', 'wp_admin_asset ACL')
self.assertIn('.php', acl_line,
'wp_admin_asset must exclude .php explicitly')
def test_wp_admin_asset_pattern_is_end_of_flags_guarded(self):
"""A pattern starting with "(" makes HAProxy warn on EVERY load/reload:
parsing acl 'wp_admin_asset' : matching 'path_reg' for pattern
'(^|/)wp-admin/...' is likely a mistake ... Maybe you need to
remove the extraneous space before '('.
"--" is the end-of-flags marker HAProxy itself names as the fix. It is
cosmetic to matching but not to operations: an unsilenced warning on
every reload on every host trains people to skim past warnings, which
is how a real one gets missed. Assert the pattern is still the one we
think it is, so this can never pass by the pattern having been changed.
"""
acl_line = require_rule(self.cfg, 'acl wp_admin_asset', 'wp_admin_asset ACL')
self.assertRegex(
acl_line, r'path_reg\s+--\s+\(',
'wp_admin_asset pattern begins with "(" and MUST be preceded by "--"')
self.assertIn('(^|/)wp-admin/(css|js|images)/', acl_line)
def test_case_insensitive_acl_and_regsub_are_kept_in_sync(self):
"""A case-insensitive wp_admin_path with a case-sensitive regsub is an
INFINITE REDIRECT LOOP: regsub finds no "/wp-admin/" in
"/WP-ADMIN/plugins.php", returns `path` unchanged, and the Location
then points at the request's own URL.
"""
acl_line = require_rule(self.cfg, 'acl wp_admin_path', 'wp_admin_path ACL')
setvar_line = require_rule(self.cfg, 'set-var(txn.wp_login_url)', 'wp_login_url set-var')
acl_ci = bool(re.search(r'path_reg\s+-i\s', acl_line))
regsub_ci = bool(re.search(r'regsub\([^)]*,\s*i\)', setvar_line))
self.assertEqual(
acl_ci, regsub_ci,
'wp_admin_path case-sensitivity (%s) and regsub flags (%s) disagree'
% (acl_line, setvar_line))
if __name__ == '__main__':
unittest.main(verbosity=2)
+135
View File
@@ -0,0 +1,135 @@
#!/usr/bin/env python3
"""Regression tests for per-client-IP rate limiting on POST /xmlrpc.php.
Why this file exists
--------------------
POST /xmlrpc.php floods were unthrottled fleet-wide. The generic frontend
rate limits (hap_listener.tpl) trigger at 3000/5000 req/10s -- i.e. 300-500
req/s -- but the observed floods run at a few req/s for hours, well under
that ceiling. The existing wp_bruteforce mechanism (dedicated stick-table,
60s window, per real client IP) solves exactly this shape of problem for
POST /wp-login.php; this change adds an equivalent dedicated table/rule pair
for POST /xmlrpc.php.
This is only safe to key on var(txn.real_ip) because of the trusted-proxy
gate added earlier (release 2026.08.3, see test-trusted-proxy-gate.py) --
before that fix, a direct client could spoof any client IP via
X-Forwarded-For and evade all per-IP tracking.
These tests pin:
- a dedicated stick-table for xmlrpc tracking exists in
hap_security_tables.tpl (own sc slot / own counter -- not sharing the
wp_bruteforce counter, so a wp-login brute-force run and an xmlrpc flood
from the same IP don't inflate each other's rate)
- the tracking rule only fires on POST /xmlrpc.php (path_end, so
subdirectory WP installs are covered)
- the limiting rule tarpits over the chosen threshold
- the limiting rule honors the same whitelist as every other rule in the
file (!is_local !is_trusted_ip !is_whitelisted)
- xmlrpc is not blocked outright -- only the rate-limit ACL is present,
there's no blanket deny of the path
Running
-------
python3 scripts/test-xmlrpc-rate-limit.py
"""
import os
import re
import sys
import logging
import tempfile
import unittest
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda filename, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(filename)), *a, **kw)
)
import haproxy_manager # noqa: E402
def render_listener():
return haproxy_manager.template_env.get_template('hap_listener.tpl').render(
crt_path='/etc/haproxy/certs',
suspension_enabled=False,
coraza_spoe_backend=None,
)
def render_security_tables():
return haproxy_manager.template_env.get_template(
'hap_security_tables.tpl').render()
class XmlrpcRateLimit(unittest.TestCase):
def setUp(self):
self.listener_cfg = render_listener()
self.tables_cfg = render_security_tables()
def test_dedicated_stick_table_defined(self):
"""A dedicated table (not wp_bruteforce) tracks xmlrpc requests, so a
wp-login brute-force run and an xmlrpc flood from the same IP can't
inflate each other's rate counter."""
self.assertRegex(
self.tables_cfg,
r'backend\s+xmlrpc_bruteforce\s*\n\s*stick-table\s+type\s+ip\b.*store.*http_req_rate',
)
# Must not be the same table wp-login already uses.
self.assertNotIn('backend wp_bruteforce\n stick-table type ip size 100k expire 30m store http_req_rate(60s)\nbackend xmlrpc_bruteforce', self.tables_cfg)
def test_xmlrpc_path_acl_uses_path_end(self):
"""path_end (not path_beg) so subdirectory WP installs are covered,
matching the wp-login rule's reasoning."""
self.assertRegex(
self.listener_cfg,
r'acl\s+xmlrpc_path\s+path_end\s+/xmlrpc\.php',
)
def test_tracking_rule_only_fires_on_post_xmlrpc(self):
self.assertRegex(
self.listener_cfg,
r'http-request\s+track-sc2\s+var\(txn\.real_ip\)\s+table\s+xmlrpc_bruteforce\s+if\s+METH_POST\s+xmlrpc_path',
)
def test_limiting_rule_tarpits_over_threshold_with_whitelist(self):
pattern = (
r'http-request\s+tarpit\s+deny_status\s+429\s+if\s+METH_POST\s+xmlrpc_path\s+'
r'\{\s*sc_http_req_rate\(2\)\s+gt\s+(\d+)\s*\}\s+'
r'!is_local\s+!is_trusted_ip\s+!is_whitelisted'
)
match = re.search(pattern, self.listener_cfg)
self.assertIsNotNone(
match, 'expected a tarpit rule tracking sc2 with the full whitelist')
threshold = int(match.group(1))
self.assertGreater(threshold, 0)
def test_xmlrpc_not_blocked_outright(self):
"""The endpoint must remain functional for clients under the
threshold -- only a rate-limit ACL, no blanket deny of the path."""
self.assertNotRegex(
self.listener_cfg,
r'http-request\s+deny\s+deny_status\s+\d+\s+if\s+(?:METH_POST\s+)?xmlrpc_path\s*(?:!is_local|\n)',
)
def test_rule_order_after_wp_login_block(self):
"""Not load-bearing for correctness (mutually exclusive paths), but
keep the new block grouped with the other WordPress-specific rules
rather than scattered elsewhere in the file."""
wp_login_idx = self.listener_cfg.index('wp_login_path')
xmlrpc_idx = self.listener_cfg.index('xmlrpc_path')
self.assertLess(wp_login_idx, xmlrpc_idx)
if __name__ == '__main__':
unittest.main(verbosity=2)
+423
View File
@@ -0,0 +1,423 @@
#!/usr/bin/env python3
"""Build gate: render the real HAProxy config and hand it to the real `haproxy -c`.
Why this file exists
--------------------
On 2026-08-14 a change to hap_listener.tpl rendered perfectly, passed every
unit test in scripts/ (13 green), and was then rejected outright by HAProxy:
[ALERT] config : parsing [/etc/haproxy/haproxy.cfg:...] :
invalid arg 2 in converter 'regsub' : ... unexpected empty
replacement string
Nothing between "commit" and "running in production" would have caught it.
The unit tests assert on the *text* of the rendered config with regexes, which
tells you what the template says, never whether HAProxy will accept it. And
scripts/test-config-rollback.py stubs the `haproxy` binary with a shell script
that only rejects a literal sentinel token, so its "validation" has never
parsed a single line of real HAProxy syntax.
The failure mode this guards is not cosmetic. When haproxy.cfg is invalid,
scripts/init.py refuses to start HAProxy but the container still comes up:
ports 80/443 are unbound, every site on the host is down, and /health keeps
answering 200 because the Flask API is fine.
So: render the config through the SAME code path production uses
(haproxy_manager.generate_config(), templates and all), then run the actual
`haproxy -c` against the result and gate on its exit code.
This runs as a RUN step in the Dockerfile, which means it also validates
against the exact haproxy binary that ships in the image being built - note
that the Dockerfile installs haproxy UNPINNED, so that binary can move under
us between builds. Any syntax the new binary rejects now fails the build
instead of failing at 3am on an edge node.
What it covers
--------------
* every template generate_config() touches, assembled in the real order
* a domain with SSL + a backend, a wildcard domain, a cert-only domain with
no backend, and two template_override backends
* blocked-IP map entries (single IP and CIDR)
* BOTH sides of the two conditional blocks in hap_listener.tpl -
{%- if suspension_enabled %} and {%- if coraza_spoe_backend %} - because a
syntax error inside a conditional ships undetected otherwise. Scenario
"full" turns both on; scenario "default" leaves both off, which is the
byte-identical-to-standalone shape.
Warnings vs failures
--------------------
`haproxy -c` emits warnings on a clean config here (at minimum "Can't load
stats file" because /var/lib/haproxy/stats.dat doesn't exist at build time,
plus assorted path_reg/ACL advisories). Those are NOT failures. This gate keys
on the process EXIT CODE only, and dumps the full output when it is non-zero.
Running
-------
python3 scripts/validate-rendered-config.py
Needs the real `haproxy` binary, the application's Python dependencies, and
write access to /etc/haproxy (several templates reference files there by
absolute path - see _REAL_PATH_NOTE below). Inside the image build all three
hold. On a workstation, run it in the container instead.
"""
import logging
import os
import re
import shutil
import sqlite3
import subprocess
import sys
import tempfile
MODULE_DIR = os.path.abspath(
os.environ.get('HAPROXY_MANAGER_DIR',
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
)
os.chdir(MODULE_DIR)
sys.path.insert(0, MODULE_DIR)
# Same trick the other suites use: haproxy_manager configures logging at import
# time against /var/log/haproxy-manager.log. Redirect the handlers so this runs
# without root and without polluting the image's log files.
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-validate-logs-')
_real_file_handler = logging.FileHandler
logging.FileHandler = (
lambda filename, *a, **kw: _real_file_handler(
os.path.join(_LOG_DIR, os.path.basename(filename)), *a, **kw)
)
try:
import haproxy_manager as hm # noqa: E402
finally:
logging.FileHandler = _real_file_handler
# The application logs a lot at INFO during a render, and legitimately logs at
# ERROR about things that are only true in this harness ("no existing HAProxy
# config on disk ... ROLLBACK IS NOT AVAILABLE" - correct, there is no live
# config during a build). Silence it so the build log carries the gate's own
# verdict and haproxy's output, which is what matters.
logging.getLogger('haproxy_manager').setLevel(logging.CRITICAL)
# _REAL_PATH_NOTE
# ---------------
# Most paths haproxy_manager writes to are module-level constants and are
# redirected into a temp dir below. Two cannot be:
#
# /etc/haproxy/blocked_ips.map - hardcoded inside hap_listener.tpl's
# map_ip() converter
# /etc/haproxy/coraza-spoe.cfg - hardcoded in the `filter spoe engine`
# line, and parsed by haproxy -c
#
# Redirecting the constants without editing the templates would just make
# haproxy read a different (missing) file, so those two are left at their real
# paths. Everything this script creates under /etc/haproxy is removed again on
# exit; pre-existing files (the baked trusted_ips.*) are never touched.
ETC_HAPROXY = '/etc/haproxy'
# Files referenced with `-f` / map_ip() from the templates. A missing `-f` file
# is a FATAL haproxy error, so a gate that didn't create these would fail for
# reasons that have nothing to do with the config being tested.
STUB_FILES = {
os.path.join(ETC_HAPROXY, 'trusted_ips.list'): '# validation stub\n203.0.113.10\n',
os.path.join(ETC_HAPROXY, 'trusted_ips.map'): '# validation stub\n203.0.113.11 1\n',
os.path.join(ETC_HAPROXY, 'cloudflare_ips.list'): '# validation stub\n198.51.100.0/24\n',
os.path.join(ETC_HAPROXY, 'trusted_proxies.list'): '# validation stub\n192.0.2.0/24\n',
os.path.join(ETC_HAPROXY, 'wpadmin_gate_exempt.list'): '# validation stub\nexempt.example.test\n',
os.path.join(ETC_HAPROXY, 'suspended_domains.list'): 'suspended.example.test\n',
# `lf-file` on the Coraza deny rule; loaded at parse time. Present in the
# image (COPY errors /haproxy/errors), stubbed for anything else.
'/haproxy/errors/403-waf.html': '<html><body>blocked %[unique-id]</body></html>\n',
}
# (suspension_enabled, coraza_spoe_backend) combinations to render + validate.
SCENARIOS = (
('default', {}),
('full', {
'HAPROXY_SUSPENSION_ENABLED': 'true',
'HAPROXY_CORAZA_SPOE_BACKEND': '127.0.0.1:9000',
}),
)
def log(msg):
sys.stdout.write(f'[validate-config] {msg}\n')
sys.stdout.flush()
def fail(msg):
sys.stderr.write(f'[validate-config] FAIL: {msg}\n')
sys.stderr.flush()
raise SystemExit(1)
class CreatedFiles:
"""Tracks what we put on disk outside the temp dir so it can be removed.
Two sources: files we create explicitly, and files generate_config() itself
writes into /etc/haproxy (blocked_ips.map, coraza-spoe.cfg and their
.backup copies). The latter are caught by diffing the directory listing,
which also picks up anything a future change starts writing there.
"""
def __init__(self):
self.explicit = []
self.etc_before = self._listdir(ETC_HAPROXY)
@staticmethod
def _listdir(path):
try:
return set(os.listdir(path))
except OSError:
return set()
def ensure(self, path, content):
"""Create path with content if it does not already exist."""
if os.path.exists(path):
return
os.makedirs(os.path.dirname(path), exist_ok=True)
with open(path, 'w') as fh:
fh.write(content)
os.chmod(path, 0o644)
self.explicit.append(path)
def cleanup(self):
for path in self.explicit:
try:
os.unlink(path)
except OSError:
pass
for name in self._listdir(ETC_HAPROXY) - self.etc_before:
try:
os.unlink(os.path.join(ETC_HAPROXY, name))
except OSError:
pass
def require_haproxy_binary():
"""Fail closed. A gate that skips itself when the binary is missing is a
gate that would have let the 2026-08-14 change through."""
path = shutil.which('haproxy')
if not path:
fail('no `haproxy` binary on PATH - this gate cannot validate anything. '
'Run it inside the image (the Dockerfile installs haproxy).')
version = subprocess.run([path, '-v'], capture_output=True, text=True)
log(f'using {path}: {version.stdout.strip().splitlines()[0] if version.stdout else "unknown version"}')
def make_self_signed_cert(certs_dir):
"""HAProxy loads every file in the `bind ... ssl crt <dir>` directory at
parse time, so the directory has to hold a real, loadable bundle."""
os.makedirs(certs_dir, exist_ok=True)
cert = os.path.join(certs_dir, 'cert.tmp')
key = os.path.join(certs_dir, 'key.tmp')
subprocess.run(
['openssl', 'req', '-x509', '-newkey', 'rsa:2048', '-nodes',
'-keyout', key, '-out', cert, '-days', '1',
'-subj', '/CN=validate.example.test'],
check=True, capture_output=True,
)
bundle = os.path.join(certs_dir, 'validate.example.test.pem')
with open(bundle, 'w') as out:
for part in (cert, key):
with open(part) as fh:
out.write(fh.read())
os.unlink(cert)
os.unlink(key)
os.chmod(bundle, 0o600)
def seed_database(db_path):
"""A representative fleet: SSL + backend, wildcard, cert-only, overrides."""
hm.DB_FILE = db_path
hm.init_db()
with sqlite3.connect(db_path) as conn:
cur = conn.cursor()
def add_site(domain, backend, ssl_enabled=1, wildcard=0, override=None,
servers=(('web1', '10.0.0.10', 8080, 'check'),)):
cur.execute(
'INSERT INTO domains (domain, ssl_enabled, ssl_cert_path, '
'template_override, is_wildcard) VALUES (?, ?, ?, ?, ?)',
(domain, ssl_enabled, f'/etc/haproxy/certs/{domain}.pem',
override, wildcard),
)
domain_id = cur.lastrowid
cur.execute('INSERT INTO backends (name, domain_id, settings) '
'VALUES (?, ?, ?)', (backend, domain_id, None))
backend_id = cur.lastrowid
for name, addr, port, opts in servers:
cur.execute(
'INSERT INTO backend_servers (backend_id, server_name, '
'server_address, server_port, server_options) '
'VALUES (?, ?, ?, ?, ?)',
(backend_id, name, addr, port, opts),
)
add_site('site-one.example.test', 'site-one',
servers=(('web1', '10.0.0.10', 8080, 'check'),
('web2', '10.0.0.11', 8080, 'check backup')))
add_site('site-two.example.test', 'site-two', ssl_enabled=0)
add_site('*.wildcard.example.test', 'wildcard-site', wildcard=1)
add_site('ws.example.test', 'ws-site', override='hap_backend_websocket')
add_site('sse.example.test', 'sse-site', override='hap_backend_longlived')
# Cert/management-only domain: registered for certificates, no backend.
# generate_config() has an explicit branch for this.
cur.execute(
'INSERT INTO domains (domain, ssl_enabled, ssl_cert_path, '
'template_override, is_wildcard) VALUES (?, ?, ?, ?, ?)',
('panel.example.test', 1, '/etc/haproxy/certs/panel.pem', None, 0),
)
# Both map_ip() shapes: a single address and a CIDR.
for ip in ('203.0.113.66', '198.51.100.0/24'):
cur.execute('INSERT INTO blocked_ips (ip_address, reason, blocked_by) '
'VALUES (?, ?, ?)', (ip, 'validation fixture', 'gate'))
conn.commit()
def render(scenario_name, env_overrides, workdir):
"""Render via generate_config() and return the path to the assembled config.
generate_config() is the real production entry point: it takes the rollback
snapshot, writes the blocked-IP map, renders every template, and writes
haproxy.cfg. Only the reload is stubbed - there is no HAProxy process to
reload during a build, and validating the file is the whole point.
"""
scenario_dir = os.path.join(workdir, scenario_name)
certs_dir = os.path.join(scenario_dir, 'certs')
os.makedirs(scenario_dir)
make_self_signed_cert(certs_dir)
hm.HAPROXY_CONFIG_PATH = os.path.join(scenario_dir, 'haproxy.cfg')
hm.HAPROXY_BACKUP_PATH = os.path.join(scenario_dir, 'haproxy.cfg.backup')
hm.CLUSTER_SECRET_PATH = os.path.join(scenario_dir, 'cluster-secret')
hm.HAPROXY_SOCKET_PATH = os.path.join(scenario_dir, 'haproxy.sock')
hm.SSL_CERTS_DIR = certs_dir
seed_database(os.path.join(scenario_dir, 'haproxy_config.db'))
saved_env = {}
for key in ('HAPROXY_SUSPENSION_ENABLED', 'HAPROXY_CORAZA_SPOE_BACKEND'):
saved_env[key] = os.environ.pop(key, None)
os.environ.update(env_overrides)
real_reload = hm.reload_haproxy_safely
hm.reload_haproxy_safely = lambda *a, **kw: (True, 'reload skipped: build-time validation')
try:
hm.generate_config()
finally:
hm.reload_haproxy_safely = real_reload
for key, value in saved_env.items():
os.environ.pop(key, None)
if value is not None:
os.environ[key] = value
return hm.HAPROXY_CONFIG_PATH
def assert_scenario_branches(scenario_name, config_path, env_overrides):
"""Cheap sanity check that the conditional blocks actually rendered.
Without this, a template refactor that silently stopped emitting the
suspension or Coraza block would leave the gate passing while covering
less than it claims to.
"""
with open(config_path) as fh:
text = fh.read()
expected = {
'suspension': ('acl is_suspended_domain',
'HAPROXY_SUSPENSION_ENABLED' in env_overrides),
'coraza': ('filter spoe engine coraza',
'HAPROXY_CORAZA_SPOE_BACKEND' in env_overrides),
}
for label, (needle, should_be_present) in expected.items():
present = needle in text
if present != should_be_present:
fail(f'[{scenario_name}] {label} block {"missing" if should_be_present else "unexpectedly present"} '
f'in the rendered config (looked for {needle!r}). The gate is '
f'not covering what it thinks it is.')
def _dump_context(config_path, output):
"""Print the rendered lines HAProxy complained about.
The temp dir is deleted on the way out, so the build log has to carry the
evidence. HAProxy reports `parsing [<file>:<line>]`; show a window around
each reported line rather than dumping ~1500 lines of config.
"""
line_numbers = sorted({
int(n) for n in re.findall(
r'parsing \[' + re.escape(config_path) + r':(\d+)\]', output)
})
if not line_numbers:
return
with open(config_path) as fh:
lines = fh.read().splitlines()
sys.stderr.write('[validate-config] --- rendered config around the error ---\n')
for number in line_numbers:
start = max(1, number - 6)
end = min(len(lines), number + 4)
for index in range(start, end + 1):
marker = '>>' if index == number else ' '
sys.stderr.write(f'{marker}{index:6d}| {lines[index - 1]}\n')
sys.stderr.write('[validate-config] ---\n')
def validate(scenario_name, config_path):
result = subprocess.run(['haproxy', '-c', '-f', config_path],
capture_output=True, text=True)
if result.returncode != 0:
output = result.stdout + result.stderr
sys.stderr.write(
f'\n[validate-config] ===== {scenario_name}: haproxy REJECTED the '
f'rendered configuration (exit {result.returncode}) =====\n')
sys.stderr.write(output if output.endswith('\n') else output + '\n')
_dump_context(config_path, output)
sys.stderr.write(
'[validate-config] This is a real HAProxy parse failure. Shipping it '
'would leave the container Up with ports 80/443 unbound and every '
'site on the host down, while /health still returns 200.\n')
sys.stderr.flush()
raise SystemExit(1)
# Exit code 0 is the verdict. Warnings are expected and are NOT failures:
# "Can't load stats file" always fires at build time, and HAProxy emits
# path_reg/ACL advisories on a perfectly valid config.
noise = (result.stdout + result.stderr).strip()
log(f'{scenario_name}: haproxy -c OK (exit 0)')
if noise:
for line in noise.splitlines():
log(f' {scenario_name}: haproxy said: {line}')
def main():
require_haproxy_binary()
if not os.path.isdir(ETC_HAPROXY) or not os.access(ETC_HAPROXY, os.W_OK):
fail(f'{ETC_HAPROXY} must exist and be writable - several templates '
'reference files there by absolute path. Run this inside the image.')
created = CreatedFiles()
workdir = tempfile.mkdtemp(prefix='haproxy-validate-')
try:
for path, content in STUB_FILES.items():
created.ensure(path, content)
for scenario_name, env_overrides in SCENARIOS:
log(f'rendering scenario "{scenario_name}" '
f'({env_overrides or "no optional features"})')
config_path = render(scenario_name, env_overrides, workdir)
assert_scenario_branches(scenario_name, config_path, env_overrides)
validate(scenario_name, config_path)
finally:
created.cleanup()
shutil.rmtree(workdir, ignore_errors=True)
shutil.rmtree(_LOG_DIR, ignore_errors=True)
log('all scenarios accepted by the real haproxy binary')
return 0
if __name__ == '__main__':
raise SystemExit(main())
+16 -1
View File
@@ -37,7 +37,22 @@ spoe-agent coraza
timeout processing 100ms
use-backend coraza-spoa-backend
log global
# NO `log global` here, deliberately.
#
# `log global` in a spoe-agent emits one line PER INSPECTED REQUEST, e.g.
# SPOE: [coraza] <GROUP:coraza-req> sid=537 st=0 0/0/0/0/0 32/32 0/0 0/467
# Measured on whp01 immediately after access logging started working:
# 618 SPOE lines vs 669 real access lines -- it was ~48% of the log volume,
# i.e. it would roughly DOUBLE the edge's log footprint (~400 MB/day extra)
# to record `st=0` over and over.
#
# It carries nothing incident response needs: the WAF's verdict is already
# visible in the access log line (status 403 + the `id=` UUID, which joins
# to /var/log/coraza/audit.log for the rule_id), and per-transaction WAF
# detail is written by the SPOA itself to /var/log/coraza/spoa.log.
# Agent-level failures still surface via `option set-on-error error` ->
# var(txn.coraza.error) and the fail-open path in hap_listener.tpl.
# Per-request inspection message. No `event` directive — fires only when
# explicitly invoked from haproxy.cfg via `http-request send-spoe-group`.
+71 -11
View File
@@ -2,20 +2,46 @@
# Global settings
#---------------------------------------------------------------------
global
# to have these messages end up in /var/log/haproxy.log you will
# need to:
# ACCESS LOG DESTINATION.
#
# 1) configure syslog to accept network log events. This is done
# by adding the '-r' option to the SYSLOGD_OPTIONS in
# /etc/sysconfig/syslog
# This used to be `log 127.0.0.1 local2`, which was a silent black hole:
# 127.0.0.1 is the CONTAINER's own loopback, nothing has ever listened on
# udp/514 in the container netns, and there is no /dev/log in the image.
# Every access log line -- ~1.5M/day across the whole edge -- was written
# to a socket with no receiver and dropped. Nothing errored, nothing
# warned, and `haproxy -c` was perfectly happy. The cost only shows up
# during an incident: per-IP 429s, tarpits, wp-admin gate redirects,
# WAF 403s and `silent-drop`s left no record anywhere, so the edge could
# not be asked what it had rejected. Only aggregate stick-table counters
# survived.
#
# 2) configure local2 events to go to the /var/log/haproxy.log
# file. A line like the following can be added to
# /etc/sysconfig/syslog
# Now points at the DOCKER BRIDGE GATEWAY, where the host's rsyslog has an
# imudp listener bound (installed idempotently by WHP's
# setup-haproxy-logrotate.sh, which also writes the logrotate stanza).
# The host writes local2 to /var/log/haproxy.log and stops it there, so it
# does not also flood /var/log/messages or the Graylog forwarder.
#
# local2.* /var/log/haproxy.log
# WHY NOT `log stdout format raw local0`: it is INCOMPATIBLE with the
# `daemon` keyword below, and incompatible SILENTLY. Verified on the
# pinned 3.0.11 binary: with `daemon` set, a `log stdout` config serves
# traffic normally and emits ZERO log lines, and `haproxy -c` returns 0
# with no error and no warning -- so the CI config gate
# (scripts/validate-rendered-config.py) cannot catch it either. Making it
# work means dropping `daemon` / adding -db so haproxy stays in the
# foreground, which in turn breaks the three synchronous
# `subprocess.run(['haproxy', '-W', ...], check=True)` launch sites in
# haproxy_manager.py (they would block until the 180s timeout and then be
# killed). That is a change to the exact code path whose failure mode is
# "container Up, ports 80/443 never bound, every site down, /health still
# 200". Not worth it for a logging change.
#
log 127.0.0.1 local2
# UDP means a dead listener degrades to dropped log lines, never to a
# stalled or failing request path -- the correct failure direction for an
# edge fronting ~60 customer sites.
#
# len 2048 accommodates the enriched log-format in hap_listener.tpl
# (URL + User-Agent + UUID); the default 1024 would truncate long ones.
log {{ syslog_target }} len 2048 format rfc5424 local2 info
chroot /var/lib/haproxy
pidfile /var/run/haproxy.pid
@@ -27,6 +53,36 @@ global
# SSL and Performance
tune.ssl.default-dh-param 2048
# Required by the `http-request normalize-uri` chain at the top of the
# `web` frontend (hap_listener.tpl). normalize-uri is still flagged
# EXPERIMENTAL in HAProxy 3.0, and HAProxy REFUSES TO START without this
# opt-in -- not a warning, a fatal:
# [ALERT] config : parsing [...] : 'normalize-uri' action is
# experimental, must be allowed via a global
# 'expose-experimental-directives'
# (verified against real haproxy 3.0.11-9e587df: `haproxy -c` exits 1).
# So this line and the normalize-uri rules must be added/removed together.
#
# Dropping this one alone does NOT crash-loop the container -- the truth
# is worse: it is a SILENT TOTAL OUTAGE that nothing escalates. Container
# init (scripts/init.py -> haproxy_manager.do_initial_setup()) calls
# generate_config() (which still succeeds -- Jinja doesn't validate
# HAProxy semantics) and then start_haproxy(), which runs `haproxy -c`,
# sees it fail, logs an error, and RETURNS WITHOUT RAISING. init.py exits
# 0. scripts/start-up.sh then execs gunicorn as PID 1 regardless. Result:
# the container stays "Up", ports 80/443 are never bound, EVERY SITE ON
# THE HOST IS DOWN, and the in-container supervisor loop
# (ensure_haproxy.py, every HAPROXY_SUPERVISOR_INTERVAL seconds) retries
# the identical failing render forever without ever escalating. Worse
# still, GET /health keeps returning HTTP 200 -- health_check() only
# answers 500 on a database error; a dead haproxy just flips the JSON
# body's "haproxy_status" to "stopped" while the status code a naive
# monitor checks never changes. Do not trust /health alone to catch this.
#
# This exposes ONLY the experimental directives that are actually used --
# it does not change the behaviour of anything else in this file.
expose-experimental-directives
# HTTP/3 over QUIC. The Debian haproxy package is built against system
# OpenSSL via the compatibility shim (USE_QUIC_OPENSSL_COMPAT), which is
# not a native QUIC TLS stack. HAProxy therefore rejects `quic*@` binds
@@ -92,7 +148,11 @@ defaults
maxconn 3000
# Per-request unique reference, used:
# - in the log line (httplog includes %ID)
# - in the access log line, as the `id=` field of the custom log-format
# in hap_listener.tpl. NOTE: `option httplog` does NOT include %ID
# (verified against haproxy 3.0.11) -- this comment used to claim it
# did, which made the support workflow below look supported when it
# was not. The explicit log-format is what actually carries it.
# - echoed to clients in the X-Request-Reference response header on
# WAF blocks so a customer can quote it when opening a support ticket
# - embedded in /etc/haproxy/errors/403-waf.html so a blocked visitor
+437 -2
View File
@@ -19,9 +19,149 @@ frontend web
# response, including haproxy-generated ones (blocks, default page).
http-after-response set-header alt-svc "h3=\":443\"; ma=86400"
# Capture Host header so it appears in httplog output (in %hr field)
# Capture Host header so it appears in httplog output (in %hr field).
# ORDER IS LOAD-BEARING: this is capture slot 0, referenced by the
# access log-format below as %[capture.req.hdr(0)].
http-request capture req.hdr(Host) len 64
# Capture slot 1 = User-Agent. Incident response needs it to tell a
# scanner from a browser, and it is not in `option httplog` output.
# Any new capture MUST be appended AFTER this line, never inserted
# above it, or the slot indices in the log-format silently shift and
# the access log starts attributing the wrong string to the wrong field.
# req.fhdr(), NOT req.hdr(): req.hdr() treats the header as a comma-
# separated list and returns only the LAST element. Real User-Agent strings
# contain commas -- "Mozilla/5.0 (Windows NT 10.0; Win64; x64)
# AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"
# captured with req.hdr() logs as just "like Gecko) Chrome/131.0.0.0
# Safari/537.36", silently losing the platform half -- which is exactly the
# half you need to tell a spoofed crawler from a real browser.
# Observed in production on whp01 before this was corrected.
http-request capture req.fhdr(User-Agent) len 200
# --- Access logging -----------------------------------------------------
# Scoped to THIS frontend on purpose: it references capture slots and
# var(txn.real_ip), which only exist here. Putting it in `defaults` would
# apply it to the stats frontend and every backend too, where those
# samples are undefined.
#
# `log-format` overrides `option httplog` for this proxy (haproxy emits a
# harmless warning saying so). The first 16 fields are byte-identical to
# the 3.0 httplog default, so anything that already parses httplog keeps
# working; the `key=value` tail is additive.
#
# Why the default httplog is not enough for incident response:
# %ci is the PROXY's address for Cloudflare-fronted sites, not the
# visitor. The real client is var(txn.real_ip), resolved further
# down from CF-Connecting-IP / X-Real-IP / X-Forwarded-For and only
# honoured from trusted proxies. BOTH are logged: cip= is who to
# rate-limit or block, %ci is which edge it arrived through.
# %ID is NOT included by `option httplog` (verified against
# haproxy 3.0.11). Without it the documented support workflow
# -- X-Request-Reference -> access log -> coraza audit.log -> rule_id
# -- cannot be completed. id= is what makes that join possible.
# host=/ua= identify the vhost and client; %ST/%B/%tsc give status,
# bytes and the termination state that distinguishes a rate-limit
# deny (PR--) from a tarpit (PT--) from a normal close.
log-format "%ci:%cp [%tr] %ft %b/%s %TR/%Tw/%Tc/%Tr/%Ta %ST %B %CC %CS %tsc %ac/%fc/%bc/%sc/%rc %sq/%bq %hr %hs %{+Q}r cip=%[var(txn.real_ip)] id=%ID host=%[capture.req.hdr(0)] ua=%[capture.req.hdr(1)] sni=%[ssl_fc_sni] hv=%[fc_http_major]"
# --- URI normalisation (MUST be the first path-touching block here) ---
# Every path-based control in this frontend (the ACME health-check bypass,
# wp-login, xmlrpc, the wp-json/batch virtual patch, the wp-admin gate, the
# blocked-IP and suspension set-path rules, and everything Coraza inspects)
# matched the RAW request-target while the backend NORMALISED and DECODED
# it before resolving a file. Every gap between those two behaviours is a
# bypass, and each one had to be patched individually. Five were found in
# the wp-admin gate alone, all the same class:
#
# //wp-admin/plugins.php raw path starts "//" -> the safe-path
# guard failed -> request fell through
# ungated, backend served it
# /wp-admin/css/../plugins.php matched the css/js/images asset
# bypass; backend resolved ".." and
# booted plugins.php
# /wp-admin/js/%2e%2e/plugins.php same, with the ".." percent-encoded
# /wp%2Dadmin/plugins.php "wp-admin" spelled with %2D never
# matched any wp-admin ACL at all
# /wp-admin%2Fplugins.php encoded separator; see the dedicated
# rule in the wp-admin gate below
#
# Rather than keep bolting a counter-pattern onto each rule, normalise the
# URI once, here, so every rule below matches the SAME string the backend
# will resolve. HAProxy rewrites the request-target in place, so the
# backend receives the normalised form too.
#
# ORDER IS LOAD-BEARING and was determined empirically against real
# haproxy 3.0.11, not from the docs. The decoders must run BEFORE the path
# walkers: with the reverse order, /wp-admin/js/%2e%2e/plugins.php ends up
# as /wp-admin/js/../plugins.php -- decoded, but the ".." left unresolved,
# because path-strip-dotdot had already run by the time the "%2e%2e"
# became "..". Verified both directions side by side.
#
# percent-to-uppercase %2f -> %2F. Canonicalises the spelling of
# whatever stays encoded, so downstream rules
# need one case of each escape, not two.
# percent-decode-unreserved Decodes ONLY RFC 3986 unreserved chars
# (A-Za-z0-9-._~). This is what turns %2e%2e
# into .. and wp%2Dadmin into wp-admin.
# Reserved escapes are deliberately left
# alone -- %2F in particular, which is why
# the wp-admin gate needs its own encoded-
# slash rule (see below).
# path-merge-slashes //x -> /x. Also removes the entire class of
# "leading // defeats an anchored regex".
# path-strip-dot /a/./b -> /a/b.
# path-strip-dotdot full /a/b/../c -> /a/c. "full" additionally
# resolves ".." segments that would climb
# above the root (/../../wp-admin/x.php ->
# /wp-admin/x.php); without "full" HAProxy
# leaves those in place and the vector
# survives -- measured, both forms tested.
#
# DELIBERATELY NOT ENABLED: query-sort-by-name. Reordering query-string
# parameters would silently break anything that signs or caches on the
# exact query string (signed asset URLs, HMAC'd callbacks, CDN cache
# keys). This is NOT because the enabled normalizers already leave the
# query alone -- see BLAST RADIUS just below, they don't -- it is a
# deliberate line drawn between "case-fold / decode", which RFC 3986
# defines as no-ops, and "reorder", which is not a no-op for a caller
# treating the query as an opaque signed string.
#
# BLAST RADIUS: this block applies to EVERY request for EVERY site on
# EVERY tier, so the decoding was kept minimal on purpose -- and it is
# NOT scoped to the path. percent-to-uppercase and percent-decode-
# unreserved normalise the WHOLE request-target as HAProxy parses it --
# query string included, not just the path component the ACLs below
# match on -- and the BACKEND receives the rewritten query on the wire,
# not just an internal haproxy view of it. Measured against real HAProxy
# 3.0.11:
# /a?sig=%2babc%2fdef -> /a?sig=%2Babc%2Fdef (percent-hex upper-cased)
# /a?b=%41%42%43 -> /a?b=ABC (unreserved chars decoded)
# /a?tok=%7e%2d%5f%2e -> /a?tok=~-_. (unreserved chars decoded)
# Only parameter ORDER is preserved -- that guarantee is exactly why
# query-sort-by-name above is the one normalizer in this family left
# disabled. Per RFC 3986 both enabled rewrites are defined as the same
# URI (case in a percent-escape, and an unreserved character vs. its
# escape, carry no distinct meaning), but "the same URI" is not "the
# same bytes": an application that HMACs or otherwise signs the RAW
# query string, rather than parsing it first, could see a mutated value
# and fail to verify an otherwise-legitimate request. Checked against
# this fleet: no .NET backends (.NET's UrlEncode emits lowercase
# percent-hex, which percent-to-uppercase would rewrite) and no
# URL-in-path proxies, so there is no known victim today -- but do not
# assume "path only" from this block; that was the actual bug in an
# earlier draft of this comment.
#
# normalize-uri is EXPERIMENTAL in 3.0 and requires
# `expose-experimental-directives` in the global section
# (hap_header.tpl). Without it HAProxy does not start. Remove one and you
# must remove the other.
http-request normalize-uri percent-to-uppercase
http-request normalize-uri percent-decode-unreserved
http-request normalize-uri path-merge-slashes
http-request normalize-uri path-strip-dot
http-request normalize-uri path-strip-dotdot full
# --- Trusted-proxy gate (MUST precede real-IP resolution below) ---
# CF-Connecting-IP / X-Real-IP / X-Forwarded-For are client-supplied. Any
# peer that is not a known reverse proxy gets them stripped, so the
@@ -118,6 +258,42 @@ frontend web
acl has_login_cookie req.cook(whplc) -m found
http-request deny deny_status 403 if METH_POST wp_login_path !has_login_cookie !is_local !is_trusted_ip !is_whitelisted
# --- WordPress xmlrpc.php flood protection ---
# xmlrpc.php floods are a common, sustained abuse pattern that the generic
# limits above don't catch: those trigger at 3000/5000 req/10s (300-500
# req/s, sized for media-heavy pageloads), while observed xmlrpc floods run
# at just a few req/s for hours -- comfortably under that ceiling but still
# enough to pin PHP-FPM workers and show up as 503s for the rest of the
# site. Same shape of problem as wp-login credential stuffing, so it gets
# the same fix: track POSTs to xmlrpc.php per real client IP in a DEDICATED
# 60s table (sc2 / backend xmlrpc_bruteforce, defined in
# hap_security_tables.tpl -- kept separate from wp_bruteforce so the two
# endpoints' traffic can't inflate each other's counter, see that file for
# the reasoning) and tarpit once an IP exceeds the threshold.
#
# Threshold is 60/min (double wp-login's 30/min), not because xmlrpc abuse
# is less severe but because legitimate traffic here is machine-to-machine
# rather than a human filling out a form: Jetpack sync, the WordPress
# mobile app, and remote-publishing clients (e.g. an offline blog editor)
# can legitimately burst several xmlrpc calls in quick succession. 60/min
# (1 req/s average over the window) comfortably absorbs that burst while
# still tripping well before an hours-long few-req/s flood does real
# damage -- at 2 req/s sustained the 60s counter clears the threshold in
# under a minute.
#
# Tarpit (not deny), matching the wp-login rule: this is per-IP tracking
# of a bounded set of offenders, not the wp-login cookie challenge's
# distributed hundreds-of-thousands-of-IPs scenario where holding
# connections would exhaust HAProxy itself, so tying up the flooding IP's
# connections is the cheaper and more effective response. path_end (not
# path_beg) covers subdirectory WP installs, same reasoning as wp-login.
# Honors the same whitelist (RFC1918 / trusted_ips.list / trusted_ips.map)
# so health checks and trusted infrastructure are unaffected, and legit
# clients under the threshold are never blocked outright.
acl xmlrpc_path path_end /xmlrpc.php
http-request track-sc2 var(txn.real_ip) table xmlrpc_bruteforce if METH_POST xmlrpc_path
http-request tarpit deny_status 429 if METH_POST xmlrpc_path { sc_http_req_rate(2) gt 60 } !is_local !is_trusted_ip !is_whitelisted
# WordPress REST batch endpoint lockdown ("wp2shell": CVE-2026-63030 +
# CVE-2026-60137). Chaining a core SQL injection with REST batch-route
# confusion gives unauthenticated RCE on WP 6.9.0-6.9.4 and 7.0.0-7.0.1
@@ -147,9 +323,268 @@ frontend web
http-request deny deny_status 403 if wp_batch_route !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
http-request deny deny_status 403 if wp_batch_route_enc !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
# --- WordPress admin edge gate ---
# Measured on whp02, 2026-08-14: distributed unauthenticated GETs booting
# WordPress just to bounce back a login redirect --
# 2801 GET /wp-admin/profile.php 2781 GET /wp-admin/edit.php 2761 GET /wp-admin/plugins.php
# -- spread across many source IPs at roughly 4 req/min per IP, each hit
# burning a PHP-FPM/lsphp worker. Two sites absorbed 1,243 resulting 503s
# in nine hours as their pools saturated.
#
# IDENTITY, NOT RATE. This is the one rule in this file that tracks
# nothing and has no stick-table counter or threshold. Every rate-based
# control above (the generic limits, wp_bruteforce, xmlrpc_bruteforce) is
# per-IP, and per-IP rate is exactly what this attack is engineered to
# stay under: ~4 req/min from any single IP is indistinguishable from a
# slow human, and the source set is large enough that no threshold can be
# lowered to catch it without also catching real visitors. There is also
# no free stick-table slot left to try anyway -- sc0/sc1/sc2 are already
# used above and HAProxy's default tune.stick-counters is 3, so a fourth
# tracked counter is not an option here. DO NOT "simplify" this into a
# rate/threshold rule later: the whole point is that a threshold cannot
# see this traffic. Instead we gate on identity -- a real logged-in
# WordPress user always carries a wordpress_logged_in_* cookie (the same
# ACL the wp2shell block above already declares), and an unauthenticated
# request to a wp-admin page has no legitimate reason to boot PHP at all.
#
# 302, not 403. WordPress itself redirects an unauthenticated /wp-admin/
# request to wp-login.php, so replicating that at the edge means an admin
# whose session merely expired lands on the normal login screen instead
# of an error page -- we are not trading a bot problem for a support
# ticket. Bots get a cheap redirect they ignore.
#
# path_reg, not path_beg. A subdirectory install at /blog/wp-admin/ slips
# past a prefix match; path_reg with an optional leading "/" catches both
# root and subdirectory installs, same reasoning as the wp-login and
# xmlrpc path_end rules above.
#
# regsub rewrites the redirect target itself, so /blog/wp-admin/x.php
# redirects to /blog/wp-login.php rather than 404ing at the site root.
#
# DO NOT write regsub's regex argument with a capturing group / literal
# parentheses, e.g. regsub((^|/)wp-admin/.*,\1wp-login.php) -- neither
# inlined into the redirect's `location` nor in a standalone set-var.
# HAProxy 3.0.11's converter-argument parser counts parens to find the
# end of the regsub(...) call itself, so the *inner* "(^|/)" grouping
# parens are misread as closing the outer call early -- it does not
# matter whether the argument is quoted ("...": still fails) or the
# parens are backslash-escaped (\(...\): still fails). Every such form
# was verified against real HAProxy 3.0.11-1+deb13u3 and all produce the
# same ALERT: "invalid arg 2 in converter 'regsub' : missing arguments
# (got 1/2)". This is a converter-argument-parsing limitation, not a
# log-format/`%[...]` issue -- the identical failure reproduces in a
# plain set-var (outside any log-format string), which rules out the
# `location` value's log-format context as the cause.
#
# The fix sidesteps groups/backreferences entirely: HTTP paths always
# start with "/", so the leading "(^|/)" alternation is redundant --
# matching the literal substring "/wp-admin/" (both slashes, no group)
# is sufficient to anchor to a real path segment (a false match like
# "/somewp-admin/" doesn't contain "/wp-admin/" as a substring, since
# there's no "/" directly before "wp-admin"). No backreference is
# needed either: regsub only replaces the matched substring, so
# replacing "/wp-admin/.*" with a literal "/wp-login.php" leaves
# whatever precedes it (the subdirectory-install prefix, if any)
# untouched. Computed in its own set-var so it is a plain sample
# expression, not something baked into the redirect's log-format
# string. Behaviorally verified live against real HAProxy 3.0.11:
# /wp-admin/edit.php -> /wp-login.php and
# /blog/wp-admin/plugins.php -> /blog/wp-login.php, both with
# redirect_to preserved. See
# .superpowers/sdd/2026-08-14-wpadmin-edge-gate/task-3b-report.md.
#
# redirect_to carries the path only (%[path,url_enc]), not the query
# string -- deliberate, see design spec. An admin bounced off
# post.php?post=123&action=edit lands back on a blank post.php rather
# than that exact post. Capturing the full URI needs capture.req.uri
# (extra config, and the captured value is length-capped) for a benefit
# that only matters on session expiry, so it was not worth it.
#
# THE ALLOWLIST IS MEASURED, NOT GUESSED -- taken from actual fleet
# traffic returning 200 on /wp-admin/*. admin-ajax.php and admin-post.php
# are the standard front-end AJAX/form-handler endpoints real themes and
# plugins call while logged out. Critically, wp-login.php loads its OWN
# css/js FROM /wp-admin/ (load-styles.php, load-scripts.php, and the
# static wp_admin_asset dirs below) -- miss those and every login page on
# the fleet renders unstyled with a broken password-strength meter, while
# wp-login.php itself still returns 200, making it a silent regression
# that "looks like" the gate is working.
#
# Each allowlist entry is anchored to /wp-admin/<file>, not a bare
# filename suffix. A bare `path_end /admin-ajax.php` also matches
# /wp-admin/evil/admin-ajax.php -- which ALSO matches wp_admin_path
# (path_reg only requires /wp-admin/ to appear somewhere), so an
# attacker-inserted path segment would sail through this allowlist
# ungated and boot full WordPress, exactly the resource exhaustion this
# gate exists to stop. Anchoring still covers subdirectory installs via
# suffix matching (/blog/wp-admin/admin-ajax.php ends with
# /wp-admin/admin-ajax.php) while rejecting an inserted directory.
#
# install.php is DELIBERATELY NOT allowlisted. It is legitimately
# reachable without a cookie during a fresh install, but it is also a
# standing scanner target and a real takeover vector on a site that was
# half-installed and then abandoned. Anyone genuinely installing uses the
# per-site exempt-list opt-out below instead.
#
# Honors the same whitelist as every other rule in this frontend
# (RFC1918 / trusted_ips.list / trusted_ips.map), and a per-site opt-out
# via /etc/haproxy/wpadmin_gate_exempt.list (operator-managed, seeded
# empty by start-up.sh) for sites where a plugin legitimately serves
# unauthenticated visitors from a /wp-admin/ URL outside this allowlist.
# EVERY ACL BELOW MATCHES THE NORMALISED PATH. The normalize-uri chain at
# the top of this frontend has already merged duplicate slashes, resolved
# "." / ".." segments (including percent-encoded ones) and decoded
# unreserved escapes by the time these run, so these patterns only have to
# describe the ONE canonical spelling the backend will resolve -- they do
# not have to anticipate every encoding of it. That is the whole point of
# the normalisation block; do not "harden" these regexes by re-adding
# encoding variants, fix the normalisation instead.
#
# wp_admin_safe_path guards against an OPEN REDIRECT this gate would
# otherwise introduce. The redirect target below is built by rewriting
# `path` with regsub -- regsub only replaces the matched substring, so
# everything BEFORE the matched "/wp-admin/" survives untouched in the
# output. `path` is not guaranteed to be a clean site-relative string;
# three concrete requests turn that survival into an off-site
# `Location:` header:
# //evil.example.com/wp-admin/x.php -> //evil.example.com/wp-login.php
# (protocol-relative -- browsers resolve "//host/path" to
# "https://host/path", so this redirects off-site with no scheme
# needed). NOW NEUTRALISED UPSTREAM: path-merge-slashes rewrites this
# to /evil.example.com/wp-admin/x.php before any ACL sees it, so the
# Location becomes the same-origin /evil.example.com/wp-login.php.
# Verified live.
# /\evil.example.com/wp-admin/x.php -> /\evil.example.com/wp-login.php
# (browsers normalise a leading "/\" the same as "//"). STILL LIVE
# after normalisation -- a backslash is not a slash, so no normalizer
# touches it. This ACL is the only thing that stops it.
# https://evil.example.com/wp-admin/x.php -> https://evil.example.com/wp-login.php
# (RFC 7230 absolute-form request targets can make HAProxy's `path`
# fetch return a full URI, not just the path component). HAProxy's own
# H1 parser answers 400 on this frontend; this ACL is the backstop.
# So wp_admin_safe_path is NOT redundant with the normalisation and must
# not be deleted as such -- one of its three vectors survives normalisation
# untouched.
#
# It is used TWO ways, and the pair matters:
# - as a POSITIVE condition on the redirect, so a pathological path can
# never produce a `Location:` header at all; and
# - as an explicit deny, so such a path is not merely un-redirected.
# The deny is what closes the failure mode the positive-condition form
# introduced on its own: "not redirected" used to mean "falls through to
# the backend UNGATED", i.e. the exact PHP-booting request this gate
# exists to stop, reachable by prefixing "//" (that specific spelling is
# now normalised away, but "/\" is not). Post-normalisation the only
# paths that reach the deny are "/\..." ones, which cannot resolve to a
# real file on any tier, so nothing legitimate is denied.
# -i (case-insensitive) and the matching ",i" flag on the regsub below are
# a PAIR -- adding one without the other produces an infinite redirect
# loop, because a case-sensitive regsub finds no "/wp-admin/" in
# "/WP-ADMIN/plugins.php", returns `path` UNCHANGED, and the Location then
# points at the request's own URL. Verified live that the pair is correct:
# /WP-ADMIN/plugins.php -> /wp-login.php and /blog/WP-Admin/plugins.php ->
# /blog/wp-login.php.
#
# On this fleet's Linux backends /WP-ADMIN/plugins.php 404s without booting
# PHP, so this is hardening rather than a live-bypass fix; it matters if a
# docroot ever sits on a case-insensitive mount, where that same request
# WOULD boot PHP. The cost is that a site with a real directory literally
# named e.g. /docs/WP-Admin/ now gets gated -- the same false positive the
# lowercase pattern already has, which is what the per-site exempt list
# exists to resolve.
acl wp_admin_path path_reg -i (^|/)wp-admin/
# Four literal backslashes here is NOT a typo. HAProxy's config-line word
# parser treats backslash as its OWN escape character before the value
# ever reaches the regex engine: "\\" (two backslashes) in the config
# collapses to one literal backslash by the time PCRE compiles it, which
# leaves an unterminated character class ("[^/\]") and fails with
# "missing terminating ] for character class" -- verified against real
# HAProxy 3.0.11. Four backslashes ("\\\\") collapse to two ("\\"),
# which PCRE then reads as a single escaped-backslash class member --
# the intended "reject a literal backslash" semantics.
acl wp_admin_safe_path path_reg ^/[^/\\\\]
acl wp_admin_allowed path_end /wp-admin/admin-ajax.php /wp-admin/admin-post.php /wp-admin/load-styles.php /wp-admin/load-scripts.php
# The (?!.*\.php) lookahead is DEFENCE IN DEPTH, not the primary fix. This
# ACL grants an un-gated bypass to everything under wp-admin/css|js|images,
# and it used to anchor its prefix but not its suffix, so
# /wp-admin/css/../plugins.php took the bypass and the backend then
# resolved ".." and booted plugins.php. path-strip-dotdot now rewrites that
# to /wp-admin/plugins.php before this ACL runs, which is the real fix; the
# lookahead additionally makes the bypass structurally incapable of
# covering a PHP entrypoint even if a future encoding trick survives
# normalisation. It excludes ".php" ONLY -- no static asset contains that
# substring, so it cannot cause the silent "login page renders unstyled"
# regression that an extension allowlist would risk. Requires PCRE2, which
# both the Debian (deployed) and Alpine haproxy builds have (+PCRE2).
#
# THE "--" IS LOAD-BEARING, DO NOT DELETE IT. HAProxy warns on any pattern
# whose first character is "(", because it cannot tell an intended regex
# from a fetch-argument list someone typo'd a space into:
# parsing acl 'wp_admin_asset' : matching 'path_reg' for pattern
# '(^|/)wp-admin/...' is likely a mistake and probably not what you want.
# Maybe you need to remove the extraneous space before '('.
# "--" is HAProxy's documented end-of-flags marker and is the fix it names
# itself ("please insert '--' between the match and the pattern"). It changes
# no matching semantics -- verified live on whp02: /wp-admin/css/login.min.css
# still passes and /wp-admin/css/x.php is still gated, before and after.
# Left unsilenced this fires on EVERY config load and reload on every host,
# where it trains operators to skim past warnings and can bury a real one.
# wp_admin_path escapes the warning only because its "-i" flag happens to
# consume the flag slot first; it is not otherwise special.
acl wp_admin_asset path_reg -- (^|/)wp-admin/(css|js|images)/(?!.*\.php).*$
acl wp_gate_exempt hdr(host),lower -f /etc/haproxy/wpadmin_gate_exempt.list
# ENCODED SEPARATOR. percent-decode-unreserved deliberately does NOT decode
# %2F -- "/" is a reserved character, and decoding it in the normalizer
# would change the path's structure (it would invent new segments), which
# is precisely why HAProxy refuses to. But OpenLiteSpeed DOES decode it and
# then serves the file: /wp-admin%2Fplugins.php was measured returning 302
# from a real WordPress site on the OLS tier, i.e. full PHP boot, while
# matching none of the ACLs above. Apache returns 404 for the same request
# (AllowEncodedSlashes Off), so this is an OLS-tier defect -- and OLS is the
# tier currently saturating.
#
# DENY, not "treat it as a wp-admin path and redirect". Two reasons:
# 1. The redirect target is computed by regsub(/wp-admin/.*) which finds
# no "/wp-admin/" in "/wp-admin%2Fplugins.php", so `path` would come
# back UNCHANGED and the Location would point at the request's own
# URL -- an infinite redirect loop, not a gate.
# 2. Nothing legitimate emits it. A path segment cannot contain a literal
# "/", so %2F inside a path is always either a probe or a proxy-
# confusion attempt, and the Apache tier has been 404ing it all along,
# so no site on the fleet can depend on it.
#
# SCOPED to paths that mention wp-admin, not all paths. A blanket "deny any
# %2F in any path" would also hit REST/API-style routes on non-WordPress
# customer apps that legitimately pass an encoded slash inside a path
# parameter. Scoping keeps the blast radius inside the attack surface this
# gate owns.
#
# Matching is on the SUBSTRING, not an anchored pattern, on purpose:
# /blog%2Fwp-admin/plugins.php hides the separator BEFORE "wp-admin", where
# an anchored (^|/)wp-admin/ never matches, and OLS still resolves it to
# /blog/wp-admin/plugins.php. Substring matching catches the separator
# wherever it is. percent-to-uppercase has already folded %2f into %2F;
# the -i is belt and braces so this rule stands on its own if the
# normalizer is ever reordered.
#
# %5C (encoded backslash) is denied on the same terms. On this fleet's
# Linux backends a backslash is an ordinary filename character, so
# /wp-admin%5Cplugins.php 404s rather than booting PHP -- measured, it is
# not a live bypass today. It is included because it is the same
# encoded-separator trick against a backend that happens to treat "\" as
# one, it costs nothing, and no legitimate path contains it.
acl wp_admin_word path -i -m sub wp-admin
acl path_has_encoded_sep path -i -m sub %2f %5c
http-request deny deny_status 403 if wp_admin_word path_has_encoded_sep !has_wp_logged_in !wp_gate_exempt !is_local !is_trusted_ip !is_whitelisted
http-request deny deny_status 403 if wp_admin_path !wp_admin_safe_path !has_wp_logged_in !wp_gate_exempt !is_local !is_trusted_ip !is_whitelisted
http-request set-var(txn.wp_login_url) path,regsub(/wp-admin/.*,/wp-login.php,i) if wp_admin_path
http-request redirect code 302 location %[var(txn.wp_login_url)]?redirect_to=%[path,url_enc] if wp_admin_path wp_admin_safe_path !wp_admin_allowed !wp_admin_asset !has_wp_logged_in !wp_gate_exempt !is_local !is_trusted_ip !is_whitelisted
# IP blocking using map file (manual blocks only)
# Map file format: /etc/haproxy/blocked_ips.map contains "<ip_or_cidr> 1" per line
# Runtime updates: echo "add map #0 IP_ADDRESS 1" | socat stdio /var/run/haproxy.sock
# Runtime updates (worker command, map referenced by PATH -- "#<id>" ids
# move on every config regeneration and "#0" silently adds nothing):
# echo "@1 add map /etc/haproxy/blocked_ips.map IP_ADDRESS 1" | socat stdio /tmp/haproxy-cli
# Checks the real client IP (from headers if present, otherwise src)
# map_ip() converter supports both single IPs and CIDR ranges (e.g., 192.168.1.0/24)
acl is_blocked_ip var(txn.real_ip),map_ip(/etc/haproxy/blocked_ips.map,0) -m int gt 0
+15
View File
@@ -13,4 +13,19 @@ frontend stats
# sc0 connection/rate table so the login-attempt threshold is independent of
# the (much higher) flood thresholds.
backend wp_bruteforce
stick-table type ip size 100k expire 30m store http_req_rate(60s)
# Dedicated stick-table for POST /xmlrpc.php flood tracking.
# Tracked via track-sc2 from the `web` frontend (hap_listener.tpl); counts
# only xmlrpc POSTs per real client IP over a 60s window. This is a SEPARATE
# table/counter from wp_bruteforce (sc1) rather than a shared one: both are
# machine-to-machine WordPress endpoints an attacker could hit from the same
# IP, and sharing a counter would let one endpoint's traffic inflate the
# other's rate -- an IP credential-stuffing wp-login while also flooding
# xmlrpc would trip the wp-login threshold early on xmlrpc volume alone (or
# vice versa). track-sc1 (wp-login) and track-sc2 (xmlrpc) are each gated on
# mutually exclusive path ACLs, so at most one of them ever fires per
# request -- HAProxy's "one track-sc<N> per counter per request" limit is
# never in play here since they're different counters anyway.
backend xmlrpc_bruteforce
stick-table type ip size 100k expire 30m store http_req_rate(60s)
+16
View File
@@ -0,0 +1,16 @@
# Per-site opt-out from the WordPress admin edge gate.
#
# Hostnames listed here are EXEMPT: unauthenticated /wp-admin/* requests for
# these sites pass through to PHP instead of being redirected to wp-login.php.
# One hostname per line, lowercase. Matched against the Host header.
#
# Referenced by templates/hap_listener.tpl:
# acl wp_gate_exempt hdr(host),lower -f /etc/haproxy/wpadmin_gate_exempt.list
#
# Add a site here when a plugin legitimately serves unauthenticated visitors
# from a /wp-admin/ URL that is not in the rule's allowlist. Symptom: "my
# plugin's admin page redirects to login".
#
# Do NOT commit real customer domains — this repo is mirrored publicly. Add
# entries directly on the server; the file lives in the /etc/haproxy named
# volume and persists across container recreates.