e33167159db8a45df6b1c65dee987de563d08dc9
2
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
b2f835a88c |
fix(logging): capture the full User-Agent, drop per-request SPOE log noise
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m19s
Two defects caught by watching the real production access log on whp01 in the
minutes after 2026.08.7 made access logging work for the first time.
1. User-Agent was being truncated to its tail.
`http-request capture req.hdr(User-Agent)` treats the header as a
comma-separated list and returns only the LAST element. Real User-Agent
strings contain commas, so
Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36
(KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36
logged as
ua=like Gecko) Chrome/131.0.0.0 Safari/537.36
losing the platform half -- exactly the half needed to tell a spoofed
crawler from a real browser, which is one of the main reasons the field was
added. Switched to req.fhdr(), which returns the full unsplit header value.
2. SPOE was writing one log line per inspected request.
`log global` inside the spoe-agent block emitted
SPOE: [coraza] <GROUP:coraza-req> sid=537 st=0 0/0/0/0/0 32/32 0/0 0/467
for every single request. Measured on whp01: 618 SPOE lines against 669 real
access lines -- ~48% of the log volume, roughly doubling the edge's log
footprint (~400 MB/day extra) to record `st=0` over and over.
It carries nothing incident response needs. The WAF verdict is already in
the access line (status 403 plus the id= UUID, which joins to
/var/log/coraza/audit.log for the rule_id), and per-transaction WAF detail
is written by the SPOA itself to /var/log/coraza/spoa.log. Agent-level
failures still surface through `option set-on-error error` ->
var(txn.coraza.error) and the fail-open path in hap_listener.tpl.
Verified: scripts/validate-rendered-config.py passes `haproxy -c` on both the
"default" and "full" scenarios against the real 3.0.11 binary; wp-admin gate,
trusted-proxy gate and xmlrpc rate-limit suites all still pass.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
711c670319 |
fix(logging): stop silently discarding every access log line
HAProxy Manager Build and Push / Build-and-Push (push) Successful in 1m46s
haproxy.cfg's global section has had `log 127.0.0.1 local2` since day one.
That is the CONTAINER's own loopback: nothing has ever listened on udp/514 in
the container netns and there is no /dev/log in the image. Every access log
line -- ~1.5M/day across ~60 customer sites -- was written to a socket with no
receiver and dropped. Nothing errored, nothing warned, and `haproxy -c` was
perfectly happy, so this survived unnoticed.
The cost only shows up during an incident. Per-IP 429s, tarpits, wp-admin gate
redirects, WAF 403s and `silent-drop`s left no record anywhere, so the edge
could not be asked what it had actually rejected -- only aggregate stick-table
counters survived. That blind spot applies to every WHP host.
Changes:
* hap_header.tpl: point `log` at {{ syslog_target }} (default 172.18.0.1:514,
the client-net bridge gateway) with `len 2048 format rfc5424 local2 info`.
WHP's setup-haproxy-syslog.sh installs the matching rsyslog receiver on the
host, in a dedicated ruleset ending in stop() so 1.5M lines/day cannot flood
/var/log/messages or the Graylog forwarder, bound to the bridge IP rather
than 0.0.0.0.
* haproxy_manager.py: render that target from HAPROXY_SYSLOG_TARGET so
standalone/home deployments on a different bridge subnet can retarget it.
* hap_listener.tpl: add a frontend-scoped `log-format`. `option httplog` is
not sufficient for incident response -- it omits %ID entirely (verified
against 3.0.11), and its %ci is the Cloudflare edge rather than the visitor
for CF-fronted sites. The new format keeps the first 16 fields byte-identical
to the httplog default (so existing parsers still work) and appends
cip=<real client, from var(txn.real_ip)>, id=<uuid>, host=, ua=, sni=, hv=.
Adds a User-Agent capture in slot 1 to feed it.
* hap_header.tpl: correct the comment claiming `option httplog` includes %ID.
It does not, which made the documented support-correlation workflow
(X-Request-Reference -> access log -> coraza audit.log -> rule_id) look
supported when it could never have worked.
Deliberately NOT using `log stdout format raw local0`: it is incompatible with
the `daemon` keyword, and incompatible SILENTLY. Verified on the pinned 3.0.11
binary -- with `daemon` set, a `log stdout` config serves traffic normally and
emits zero log lines, while `haproxy -c` returns 0 with no error and no
warning, so scripts/validate-rendered-config.py could not catch it either.
Making it work would mean dropping `daemon`, which breaks the three
synchronous `subprocess.run(['haproxy', '-W', ...], check=True)` launch sites
in haproxy_manager.py -- the exact code path whose failure mode is "container
Up, ports 80/443 never bound, every site down, /health still 200".
UDP was chosen so a dead listener degrades to dropped log lines rather than a
stalled request path.
Verified: scripts/validate-rendered-config.py passes `haproxy -c` on both the
"default" and "full" scenarios against the real 3.0.11 binary; and a live
haproxy running WITH `daemon` (as production does) was confirmed to emit real
lines carrying the true client IP from CF-Connecting-IP:
<150>1 2026-08-22T17:05:21+00:00 - haproxy 109 - - 127.0.0.1:51194
[22/Aug/2026:17:05:21.217] t t/<NOSRV> 0/-1/-1/-1/0 200 73 - - LR--
1/1/0/0/0 0/0 {cf-site.example|Mozilla/5.0 RealVisitor} "GET /checkout/
HTTP/1.1" cip=203.0.113.77 id=dfe94fa9-8d95-4126-81e1-821578f22872
host=cf-site.example ua=Mozilla/5.0 RealVisitor sni=- hv=1
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|