Compare commits
54
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
9d16151120 | ||
|
|
b892438070 | ||
|
|
2a2b9739fc | ||
|
|
7732e2a2ff | ||
|
|
89c74c10cf | ||
|
|
1b557b9931 | ||
|
|
6ced2f8797 | ||
|
|
3917b6d1ae | ||
|
|
d9cc5311de | ||
|
|
f1c1954378 | ||
|
|
e190ca9f8e | ||
|
|
a204a44d42 | ||
|
|
04c98b1c1b | ||
|
|
1ff51da6f0 | ||
|
|
158ad3bde8 | ||
|
|
8b74cd5a4e | ||
|
|
eb3658b68e | ||
|
|
83fee2ff78 | ||
|
|
fedb025fb8 | ||
|
|
6448cffb91 | ||
|
|
47b9c87e1d | ||
|
|
8d04fe43fd | ||
|
|
99dfe98aaf | ||
|
|
e2290192f3 | ||
|
|
2e19513851 | ||
|
|
de221a1326 | ||
|
|
e1d479b74e | ||
|
|
131284fd0c | ||
|
|
edb02d6206 | ||
|
|
c4a9d9d7e6 | ||
|
|
657584c88e | ||
|
|
4262b24b7e | ||
|
|
ddb4c20073 | ||
|
|
69bed61697 | ||
|
|
fed3457f29 | ||
|
|
df56321167 | ||
|
|
a89cdab886 | ||
|
|
b359af15e5 | ||
|
|
9b329fa52b | ||
|
|
fd79ac2b70 | ||
|
|
e11d8c4269 | ||
|
|
c22d2cd1f4 | ||
|
|
be3bc040df | ||
|
|
873048e837 | ||
|
|
c931e54d4d | ||
|
|
3770dae20f | ||
|
|
b5bb141899 | ||
|
|
5ebc3fb8e8 | ||
|
|
df758a3fde | ||
|
|
fbb94e6dc3 | ||
|
|
58bb5b4f18 | ||
|
|
68a6f1bc27 | ||
|
|
5bdc95109e | ||
|
|
978f173814 |
@@ -0,0 +1,215 @@
|
||||
---
|
||||
name: haproxy-manager-deploy
|
||||
description: Use when shipping a haproxy-manager-base code change — editing templates, the Dockerfile, the Python manager, the coraza-spoa subdir, or static assets like errors/, then getting it onto whp01 or staging. Trigger eagerly on phrases like "deploy haproxy", "ship the haproxy change", "rebuild haproxy-manager", "update the WAF block page", "recreate haproxy-manager", or any time the next step would involve `git push` from this repo, `docker pull` on the image, or `container-manager.sh recreate`. Walks the Gitea-CI-auto-build + recreate flow, surfaces the named-volume shadowing foot-gun, and includes post-deploy verification.
|
||||
---
|
||||
|
||||
# haproxy-manager-base commit / build / deploy
|
||||
|
||||
This is procedural discipline for changes to `haproxy-manager-base`. The repository builds via Gitea Actions on push, not via a local build script (the WHP flow uses `build-release.sh`; this one doesn't — don't conflate them, see the `whp-deploy` skill in the whp repo for that one). Each step has caught a real foot-gun.
|
||||
|
||||
## The pipeline at a glance
|
||||
|
||||
```
|
||||
edit code (local)
|
||||
└─> commit + push
|
||||
└─> Gitea Actions auto-build (build-push.yaml / build-push-coraza.yaml)
|
||||
├─> publishes :latest tag to repo.anhonesthost.net
|
||||
└─> wait for image (~2-4 min)
|
||||
└─> recreate container on target server
|
||||
└─> verify
|
||||
```
|
||||
|
||||
Do not skip the verify step. The container can come up "healthy" while still serving stale config or missing a baked-in file (see Step 5).
|
||||
|
||||
---
|
||||
|
||||
## Step 0a — Resolve the target host (never hardcoded)
|
||||
|
||||
This skill deliberately does **not** bake in a server hostname — this repo is mirrored to a public remote, so a real FQDN in the skill would leak into commits. Instead, resolve the deploy target into a `DEPLOY_HOST` shell variable that every `ssh` command below uses.
|
||||
|
||||
```bash
|
||||
HOST_FILE=".claude/skills/haproxy-manager-deploy/target-host.local"
|
||||
DEPLOY_HOST="$(cat "$HOST_FILE" 2>/dev/null)"
|
||||
```
|
||||
|
||||
- **If `$DEPLOY_HOST` is non-empty**, use it — that's the user's saved target. The file is gitignored, so the real hostname never lands in a commit.
|
||||
- **If it's empty**, ask the user which server this deploy targets (e.g. production vs. staging) and what its hostname or SSH alias is. Then offer to save it so future deploys don't have to ask:
|
||||
|
||||
```bash
|
||||
echo 'the-host-they-gave.example' > "$HOST_FILE" # gitignored — safe to store the real FQDN here
|
||||
```
|
||||
|
||||
Confirm `$DEPLOY_HOST` is set before running any `ssh` step:
|
||||
|
||||
```bash
|
||||
[ -n "$DEPLOY_HOST" ] || echo "DEPLOY_HOST not set — ask the user for the target server"
|
||||
```
|
||||
|
||||
All commands below assume the variable is set in the same shell session (`ssh root@"$DEPLOY_HOST" ...`).
|
||||
|
||||
---
|
||||
|
||||
## Step 0 — Confirm before pushing
|
||||
|
||||
If the user just said "deploy" or "ship the haproxy fix", confirm what's actually changing: a template, the Python manager, the coraza-spoa subdir (separate image, separate workflow), or a static asset. Look at `git status` and `git diff` and read the diff back to the user if it's non-trivial.
|
||||
|
||||
Anything that affects the customer-facing block path (e.g. `templates/hap_listener.tpl`, `errors/403-waf.html`) is **visible to every visitor on every site**. Authorization is per-deploy, not standing.
|
||||
|
||||
---
|
||||
|
||||
## Step 1 — Know which workflow your change triggers
|
||||
|
||||
- `build-push.yaml` builds `haproxy-manager-base:latest` (the main image). Triggered by changes anywhere outside `coraza-spoa/`.
|
||||
- `build-push-coraza.yaml` builds `coraza-spoa:latest`. Triggered by changes inside `coraza-spoa/`.
|
||||
- `mirror-base-image.yaml` is a scheduled job mirroring upstream base images; unrelated to feature deploys.
|
||||
|
||||
If you've changed both subtrees in one push, both workflows fire — note that the order they finish isn't guaranteed.
|
||||
|
||||
---
|
||||
|
||||
## Step 2 — Beware the `/etc/haproxy` named volume shadow
|
||||
|
||||
If your change adds a NEW file that the running container needs (a baked-in asset, an errorfile, a new config snippet), **do not place it under `/etc/haproxy/` in the Dockerfile**. That path is a Docker named volume in deployed containers — image content only seeds the volume on first creation, so existing deployments will not see your new file even after a recreate.
|
||||
|
||||
Safe paths for baked-in assets:
|
||||
- `/haproxy/...` (the image's WORKDIR — not volumed)
|
||||
- Anywhere outside `/etc/haproxy`, `/etc/letsencrypt`
|
||||
|
||||
Reference the asset by absolute path from the haproxy config templates (e.g. `lf-file /haproxy/errors/403-waf.html`).
|
||||
|
||||
See `feedback-haproxy-named-volume` memory for the full pattern.
|
||||
|
||||
---
|
||||
|
||||
## Step 3 — Commit + push
|
||||
|
||||
Standard commit format with trailing `Co-Authored-By:` line. Match the recent commit message style (`git log --oneline -5`). Stage files explicitly by name.
|
||||
|
||||
```bash
|
||||
git push origin main
|
||||
```
|
||||
|
||||
Pushing immediately triggers the Gitea Actions build.
|
||||
|
||||
---
|
||||
|
||||
## Step 4 — Wait for the build
|
||||
|
||||
The Go build inside coraza-spoa takes ~2-3 minutes; the haproxy-manager-base build is faster (~1-2 min). Don't bother polling the runs UI — just pull on the target server until the digest changes:
|
||||
|
||||
```bash
|
||||
ssh root@"$DEPLOY_HOST" 'until docker pull -q repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest 2>&1 | tail -1 | grep -qE "Image is up to date|Status: Downloaded"; do sleep 15; done'
|
||||
```
|
||||
|
||||
`-q` suppresses the noisy layer progress so the grep can match cleanly. If you started this command before the CI build finished, it'll loop until the new image lands; once the digest matches, it exits.
|
||||
|
||||
To confirm you got the new image, check the image-creation time vs your push:
|
||||
|
||||
```bash
|
||||
ssh root@"$DEPLOY_HOST" 'docker images repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base --format "{{.CreatedSince}}"'
|
||||
```
|
||||
|
||||
It should say "X minutes ago" matching the build wait, not "yesterday".
|
||||
|
||||
---
|
||||
|
||||
## Step 5 — Verify the new image has what you think it has, BEFORE recreating
|
||||
|
||||
The image is `gcr.io/distroless/static-debian12:nonroot`-based, no shell. To peek inside, run a one-shot with a sh entrypoint override (only works if you put one in the image — coraza-spoa is distroless and won't have sh; haproxy-manager-base is Python-based and does):
|
||||
|
||||
```bash
|
||||
# haproxy-manager-base (has sh):
|
||||
ssh root@"$DEPLOY_HOST" 'docker run --rm --entrypoint sh repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest -c "ls /haproxy/errors/ && head -5 /haproxy/errors/403-waf.html"'
|
||||
|
||||
# coraza-spoa (distroless, no sh) — use docker create + docker cp instead:
|
||||
ssh root@"$DEPLOY_HOST" 'docker create --name _peek repo.anhonesthost.net/cloud-hosting-platform/coraza-spoa:latest && docker cp _peek:/etc/coraza/config.yaml - | tar xO; docker rm _peek'
|
||||
```
|
||||
|
||||
This step exists because the CI build can succeed but ship the wrong file (wrong commit pulled, build cache issue, etc.). Catching it here is one step earlier than catching it from a customer report.
|
||||
|
||||
---
|
||||
|
||||
## Step 6 — Recreate the container
|
||||
|
||||
```bash
|
||||
ssh root@"$DEPLOY_HOST" '/root/whp/scripts/container-manager.sh recreate haproxy-manager'
|
||||
```
|
||||
|
||||
For coraza-spoa changes:
|
||||
```bash
|
||||
ssh root@"$DEPLOY_HOST" '/root/whp/scripts/container-manager.sh recreate coraza-spoa'
|
||||
```
|
||||
|
||||
`container-manager.sh recreate` does: stop, remove, docker pull (idempotent if already pulled), start with the right flags from settings.json. **It reads `/docker/whp/settings.json` for things like `coraza_waf.mode`**, so if the user has toggled mode while you were building, the recreated container reflects the current setting — not whatever it was when you started.
|
||||
|
||||
---
|
||||
|
||||
## Step 7 — Verify the deploy
|
||||
|
||||
For haproxy-manager:
|
||||
|
||||
```bash
|
||||
ssh root@"$DEPLOY_HOST" '
|
||||
echo "=== container ==="
|
||||
docker ps --filter name=haproxy-manager --format "image: {{.Image}} status: {{.Status}}"
|
||||
echo "=== healthy ==="
|
||||
docker inspect haproxy-manager --format "{{.State.Health.Status}}"
|
||||
echo "=== haproxy config valid ==="
|
||||
docker exec haproxy-manager haproxy -c -f /etc/haproxy/haproxy.cfg 2>&1 | tail -3
|
||||
echo "=== new asset reachable inside container ==="
|
||||
docker exec haproxy-manager ls -la /haproxy/errors/ 2>&1 | tail -3
|
||||
echo "=== panel health ==="
|
||||
curl -fsS -m 5 -o /dev/null -w "PANEL=%{http_code}\n" http://127.0.0.1:8000/health
|
||||
'
|
||||
```
|
||||
|
||||
Pass criteria:
|
||||
- Container status = healthy
|
||||
- haproxy config validates (warnings OK, errors not)
|
||||
- Your new asset (if any) is at the expected path inside the running container
|
||||
- Panel returns 200
|
||||
|
||||
If any check fails, the change still went out — diagnose immediately. Don't say "deploy complete" before this clears.
|
||||
|
||||
---
|
||||
|
||||
## Step 8 — End-to-end test if customer-visible
|
||||
|
||||
If your change affects what a visitor sees (block pages, redirects, security responses), do a synthetic test that exercises the actual path. For WAF block-page changes, the recipe is:
|
||||
|
||||
```bash
|
||||
# Inject a temporary ACL that forces the WAF deny path on a custom header,
|
||||
# fire one request, observe the rendered response, then revert + reload.
|
||||
ssh root@"$DEPLOY_HOST" '
|
||||
docker exec haproxy-manager cp /etc/haproxy/haproxy.cfg /tmp/cfg-bak
|
||||
docker exec haproxy-manager sh -c "sed -i \"/http-request send-spoe-group coraza coraza-req/a\\\\ http-request set-var(txn.coraza.action) str(deny) if { req.hdr(x-force-waf-block) -m str yes }\" /etc/haproxy/haproxy.cfg"
|
||||
docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null
|
||||
sleep 1
|
||||
curl -sSk -D - -H "x-force-waf-block: yes" -H "Host: <live-vhost>" "https://localhost/" | head -40
|
||||
# revert
|
||||
docker exec haproxy-manager cp /tmp/cfg-bak /etc/haproxy/haproxy.cfg
|
||||
docker exec haproxy-manager sh -c "echo reload | socat stdio /tmp/haproxy-cli" >/dev/null
|
||||
'
|
||||
```
|
||||
|
||||
**Pick a real `<live-vhost>`.** The `Host:` header must match a domain currently served by this haproxy-manager, or the request won't route to the WAF path. Don't hardcode a customer hostname in this skill — pull a live one at test time (any entry from the panel's domain list, or `docker exec haproxy-manager ls /etc/letsencrypt/live`) and substitute it.
|
||||
|
||||
**The injection point matters.** Insert AFTER `http-request send-spoe-group coraza coraza-req`, because the SPOE call overwrites `txn.coraza.action` based on the real Coraza verdict — if you inject before it, your override is wiped.
|
||||
|
||||
**The reload mechanism matters.** Use `echo reload | socat stdio /tmp/haproxy-cli` — the container is python-based but doesn't have `kill` in PATH, and `docker kill --signal=HUP` signals the python manager (PID 1), not haproxy.
|
||||
|
||||
---
|
||||
|
||||
## Recovery hints
|
||||
|
||||
- **`docker pull` exits "Image is up to date" but your change isn't there** — CI hasn't finished yet. Check `https://repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base/actions` for in-progress runs.
|
||||
- **Container recreates but new file is missing inside** — you put the file under `/etc/haproxy/` and the named volume shadows it. See Step 2. Move the file under `/haproxy/` (or another non-volumed path) and rebuild.
|
||||
- **HAProxy `lf-file` page renders but CSS is broken / percentages stripped** — literal `%` in the file body must be doubled (`100%%`). HAProxy log-format expansion eats single `%`. See `haproxy-lf-file-percent-escape` memory.
|
||||
- **Synthetic test returns 200 from gunicorn instead of the block page** — your test ACL is being overwritten by the SPOE call. Inject after `send-spoe-group`, not before.
|
||||
- **`docker exec haproxy-manager kill -HUP 1` fails** — the python-based container doesn't have `kill` in PATH. Use the haproxy admin socket: `echo reload | socat stdio /tmp/haproxy-cli`.
|
||||
|
||||
---
|
||||
|
||||
## Why this skill is rigid
|
||||
|
||||
The pipeline is short, but the volume-shadowing trap and the SPOE-overwrite trap during testing each cost a 5-10 minute debugging detour during the session this skill was authored from. Both are silent failures — your change goes out, the container is healthy, and you only notice the bug when a customer report (or a careful synthetic test) surfaces it. The verification steps exist to catch them before that happens.
|
||||
@@ -0,0 +1,54 @@
|
||||
name: Build and push coraza-spoa
|
||||
run-name: ${{ gitea.actor }} pushed a change to coraza-spoa/
|
||||
|
||||
# Triggers only on changes to the coraza-spoa subdirectory or this workflow
|
||||
# file itself — keeps the main haproxy-manager-base build and the coraza-spoa
|
||||
# build independent. workflow_dispatch lets us trigger manually after bumping
|
||||
# the upstream coraza-spoa version pin.
|
||||
on:
|
||||
push:
|
||||
branches:
|
||||
- main
|
||||
paths:
|
||||
- 'coraza-spoa/**'
|
||||
- '.gitea/workflows/build-push-coraza.yaml'
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
Build-and-Push:
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- name: Checkout
|
||||
uses: actions/checkout@v4
|
||||
|
||||
- name: Set up QEMU
|
||||
uses: docker/setup-qemu-action@v3
|
||||
|
||||
- name: Set up Docker Buildx
|
||||
uses: https://github.com/docker/setup-buildx-action@v3
|
||||
|
||||
- name: Login to Gitea
|
||||
uses: docker/login-action@v3
|
||||
with:
|
||||
registry: repo.anhonesthost.net
|
||||
username: ${{ secrets.CI_USER }}
|
||||
password: ${{ secrets.CI_TOKEN }}
|
||||
|
||||
# Mirror to GitHub Container Registry — see build-push.yaml for the
|
||||
# secret/username convention.
|
||||
- name: Login to GHCR
|
||||
uses: docker/login-action@v3
|
||||
with:
|
||||
registry: ghcr.io
|
||||
username: shadowdao
|
||||
password: ${{ secrets.GHCR_TOKEN }}
|
||||
|
||||
- name: Build Image
|
||||
uses: docker/build-push-action@v6
|
||||
with:
|
||||
context: ./coraza-spoa
|
||||
platforms: linux/amd64
|
||||
push: true
|
||||
tags: |
|
||||
repo.anhonesthost.net/cloud-hosting-platform/coraza-spoa:latest
|
||||
ghcr.io/shadowdao/coraza-spoa:latest
|
||||
@@ -25,10 +25,37 @@ jobs:
|
||||
username: ${{ secrets.CI_USER }}
|
||||
password: ${{ secrets.CI_TOKEN }}
|
||||
|
||||
# Second push target so the image is also available from GitHub Container
|
||||
# Registry under the user's account. The PAT only needs write:packages
|
||||
# (and read:packages if the package is private). Stored in Gitea as
|
||||
# secrets.GHCR_TOKEN; username is the literal GitHub login.
|
||||
- name: Login to GHCR
|
||||
uses: docker/login-action@v3
|
||||
with:
|
||||
registry: ghcr.io
|
||||
username: shadowdao
|
||||
password: ${{ secrets.GHCR_TOKEN }}
|
||||
|
||||
# Read the human-readable release version from the VERSION file so every
|
||||
# build is pinnable for rollback (alongside the immutable git SHA). Bump
|
||||
# VERSION (YYYY.MM.N) in the same commit as a release-worthy change.
|
||||
- name: Read version
|
||||
id: ver
|
||||
run: echo "version=$(cat VERSION)" >> "$GITHUB_OUTPUT"
|
||||
|
||||
- name: Build Image
|
||||
uses: docker/build-push-action@v6
|
||||
with:
|
||||
platforms: linux/amd64
|
||||
push: true
|
||||
build-args: |
|
||||
VERSION=${{ steps.ver.outputs.version }}
|
||||
# Three tags per registry: :latest (moving), :<version> (human-readable
|
||||
# release), :<sha> (immutable, guaranteed-unique rollback target).
|
||||
tags: |
|
||||
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:${{ steps.ver.outputs.version }}
|
||||
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:${{ gitea.sha }}
|
||||
ghcr.io/shadowdao/haproxy-manager-base:latest
|
||||
ghcr.io/shadowdao/haproxy-manager-base:${{ steps.ver.outputs.version }}
|
||||
ghcr.io/shadowdao/haproxy-manager-base:${{ gitea.sha }}
|
||||
|
||||
@@ -0,0 +1,65 @@
|
||||
name: Mirror base images
|
||||
run-name: weekly base-image mirror
|
||||
|
||||
# Pulls each declared base image from upstream and re-pushes to the in-house
|
||||
# registry, so any of our images that FROM these don't depend on docker.io's
|
||||
# Cloudflare R2 blob storage being reachable. The 2026-05-12 Cloudflare
|
||||
# incident motivated this for python:3.12-slim and again for golang:1.25
|
||||
# when the coraza-spoa build hit the same blob-fetch failure.
|
||||
#
|
||||
# Adding a new mirror = add one entry to the matrix below. The destination
|
||||
# tag is always cloud-hosting-platform/<image>:<tag>, matching upstream.
|
||||
|
||||
on:
|
||||
schedule:
|
||||
# Mondays 06:00 UTC — outside customer peak hours and well before the
|
||||
# typical Tuesday/Thursday push cycles. workflow_dispatch lets us trigger
|
||||
# manually from the Gitea UI when upstream publishes patches.
|
||||
- cron: '0 6 * * 1'
|
||||
workflow_dispatch:
|
||||
|
||||
jobs:
|
||||
Mirror-Base:
|
||||
runs-on: ubuntu-latest
|
||||
strategy:
|
||||
# fail-fast=false so one image's upstream being down doesn't block the
|
||||
# others from refreshing.
|
||||
fail-fast: false
|
||||
matrix:
|
||||
image:
|
||||
- { src: 'docker.io/library/python:3.12-slim', dst_path: 'cloud-hosting-platform/python', tag: '3.12-slim' }
|
||||
- { src: 'docker.io/library/golang:1.25', dst_path: 'cloud-hosting-platform/golang', tag: '1.25' }
|
||||
|
||||
steps:
|
||||
- name: Login to in-house registry
|
||||
uses: docker/login-action@v3
|
||||
with:
|
||||
registry: repo.anhonesthost.net
|
||||
username: ${{ secrets.CI_USER }}
|
||||
password: ${{ secrets.CI_TOKEN }}
|
||||
|
||||
- name: Pull, retag, push ${{ matrix.image.src }}
|
||||
run: |
|
||||
set -euo pipefail
|
||||
SRC="${{ matrix.image.src }}"
|
||||
DST="repo.anhonesthost.net/${{ matrix.image.dst_path }}:${{ matrix.image.tag }}"
|
||||
|
||||
echo "::group::Pulling ${SRC}"
|
||||
docker pull "${SRC}"
|
||||
echo "::endgroup::"
|
||||
|
||||
# Capture the upstream digest so the workflow log shows what we
|
||||
# actually pushed. Helps diagnose "did the mirror really update"
|
||||
# questions later.
|
||||
SRC_DIGEST=$(docker image inspect "${SRC}" -f '{{index .RepoDigests 0}}')
|
||||
echo "upstream digest: ${SRC_DIGEST}"
|
||||
|
||||
docker tag "${SRC}" "${DST}"
|
||||
|
||||
echo "::group::Pushing ${DST}"
|
||||
docker push "${DST}"
|
||||
echo "::endgroup::"
|
||||
|
||||
# Sanity: the in-house tag should now resolve to the same content.
|
||||
DST_DIGEST=$(docker image inspect "${DST}" -f '{{index .RepoDigests 0}}')
|
||||
echo "mirror digest: ${DST_DIGEST}"
|
||||
@@ -37,3 +37,7 @@ ENV/
|
||||
# OS
|
||||
.DS_Store
|
||||
Thumbs.db
|
||||
|
||||
# Local-only deploy config (never commit real hostnames)
|
||||
.claude/skills/haproxy-manager-deploy/target-host.local
|
||||
*.local
|
||||
@@ -39,10 +39,11 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
|
||||
- `backend_servers` - Individual servers within backend groups
|
||||
|
||||
3. **Template System** - Jinja2 templates for HAProxy configuration generation:
|
||||
- `hap_header.tpl` - Global HAProxy settings and defaults
|
||||
- `hap_header.tpl` - Global HAProxy settings, defaults, and HTTP/2 tuning
|
||||
- `hap_backend.tpl` - Backend server definitions
|
||||
- `hap_listener.tpl` - Frontend listener configurations
|
||||
- `hap_listener.tpl` - Frontend listener configurations with rate limiting
|
||||
- `hap_letsencrypt.tpl` - SSL certificate configurations
|
||||
- `hap_security_tables.tpl` - Stats frontend and security stick tables
|
||||
- Template override support for custom backend configurations
|
||||
|
||||
4. **Certificate Management** - Automated SSL certificate handling:
|
||||
@@ -73,10 +74,51 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co
|
||||
- Certificate private keys combined with certificates in HAProxy-compatible format
|
||||
- Default backend page for unmatched domains instead of exposing HAProxy errors
|
||||
|
||||
### Rate Limiting & Connection Limits (hap_listener.tpl)
|
||||
|
||||
- **Stick table**: `type ip size 200k expire 10m` tracking `conn_cur`, `conn_rate(10s)`, `http_req_rate(10s)`, `http_err_rate(30s)`
|
||||
- Tracks real client IP via `var(txn.real_ip)` to work correctly behind Cloudflare/proxies
|
||||
- **Rate limit thresholds**:
|
||||
- Tarpit at 3000 req/10s (300 req/s)
|
||||
- Hard block (deny) at 5000 req/10s (500 req/s)
|
||||
- Connection rate limit: 500/10s
|
||||
- Concurrent connection limit: 500
|
||||
- Error rate limit: 100/30s
|
||||
- **Whitelist bypasses** (exempt from rate limits):
|
||||
- `is_local` — RFC1918 private address ranges
|
||||
- `is_trusted_ip` — source IPs listed in `trusted_ips.list`
|
||||
- `is_whitelisted` — real IPs (from proxy headers) matched in `trusted_ips.map`
|
||||
|
||||
### Trusted IP Whitelist Files
|
||||
|
||||
- `trusted_ips.list` — Source IP whitelist for rate limit bypass (one CIDR/IP per line)
|
||||
- `trusted_ips.map` — Real IP whitelist for proxy-header matching (format: `<IP> 1`)
|
||||
- Both files are baked into the Docker image via `COPY` in the Dockerfile
|
||||
- Ship as comment-only templates (no real IPs). Add trusted IPs locally and do **not** commit them — this repo is mirrored publicly. Entries persist in the `/etc/haproxy` named volume across recreates
|
||||
|
||||
### Timeout Hardening (hap_header.tpl)
|
||||
|
||||
- `timeout http-request`: 300s -> 30s (slowloris protection)
|
||||
- `timeout connect`: 120s -> 10s
|
||||
- `timeout client`: 10m -> 5m
|
||||
- `timeout http-keep-alive`: 120s -> 30s
|
||||
|
||||
### HTTP/2 Protection (hap_header.tpl)
|
||||
|
||||
- `tune.h2.fe.max-total-streams 2000` — limits total streams per HTTP/2 connection
|
||||
- `tune.h2.fe.glitches-threshold 50` — CVE-2023-44487 Rapid Reset protection
|
||||
|
||||
### Stats Frontend (hap_security_tables.tpl)
|
||||
|
||||
- HAProxy stats page bound to `127.0.0.1:8404` (localhost only, accessible inside container)
|
||||
- Template: `templates/hap_security_tables.tpl`
|
||||
|
||||
### Deployment Context
|
||||
|
||||
- Designed to run as Docker container with persistent volumes for certificates and configurations
|
||||
- Exposes ports 80 (HTTP), 443 (HTTPS), and 8000 (management API/UI)
|
||||
- Stats page on port 8404 (localhost only inside container)
|
||||
- Management interface on port 8000 should be firewall-protected in production
|
||||
- Dockerfile HEALTHCHECK verifies both port 8000 (Flask API) and port 80 (HAProxy), with `start-period=60s` and `timeout=10s`
|
||||
- Supports deployment on servers with git directory at `/root/whp` and web file sync via rsync to `/docker/whp/web/`
|
||||
- HAProxy is version 3.0.11
|
||||
+35
-2
@@ -1,10 +1,41 @@
|
||||
FROM python:3.12-slim
|
||||
# Base image mirrored into the in-house registry to remove docker.io
|
||||
# (Cloudflare R2) as a single point of failure for CI builds. The 2026-05-12
|
||||
# Cloudflare incident took down docker.io blob pulls and broke this image's CI.
|
||||
# Refresh procedure (run on a workstation that can reach docker.io, e.g.
|
||||
# monthly or when Python patches drop):
|
||||
# docker pull docker.io/library/python:3.12-slim
|
||||
# docker tag docker.io/library/python:3.12-slim \
|
||||
# repo.anhonesthost.net/cloud-hosting-platform/python:3.12-slim
|
||||
# docker push repo.anhonesthost.net/cloud-hosting-platform/python:3.12-slim
|
||||
# Future improvement: a scheduled Gitea Action that does the above automatically.
|
||||
FROM repo.anhonesthost.net/cloud-hosting-platform/python:3.12-slim
|
||||
|
||||
# image.source is what ghcr.io uses to link the package to a GitHub repo
|
||||
# sidebar; pointing at the public GitHub mirror enables that linking. The
|
||||
# canonical source-of-truth git remote is still Gitea, but Gitea's registry
|
||||
# doesn't consume this label, so there's no contention.
|
||||
# Stamped from the VERSION file by CI (build-arg) so `docker inspect` reports
|
||||
# what's running on any host. Defaults to "dev" for local/manual builds.
|
||||
ARG VERSION=dev
|
||||
LABEL org.opencontainers.image.title="haproxy-manager-base" \
|
||||
org.opencontainers.image.description="HAProxy management API with Let's Encrypt automation, Coraza WAF integration, and template-driven config" \
|
||||
org.opencontainers.image.source="https://github.com/shadowdao/haproxy-manager-base" \
|
||||
org.opencontainers.image.version="${VERSION}" \
|
||||
org.opencontainers.image.licenses="MIT"
|
||||
|
||||
RUN apt update -y && apt dist-upgrade -y && apt install socat haproxy cron certbot curl jq net-tools -y && apt clean && rm -rf /var/lib/apt/lists/*
|
||||
WORKDIR /haproxy
|
||||
COPY ./templates /haproxy/templates
|
||||
COPY requirements.txt /haproxy/
|
||||
COPY haproxy_manager.py /haproxy/
|
||||
COPY scripts /haproxy/scripts
|
||||
COPY trusted_ips.list /etc/haproxy/trusted_ips.list
|
||||
COPY trusted_ips.map /etc/haproxy/trusted_ips.map
|
||||
# /etc/haproxy is a named volume in deployed containers, so baked-in files
|
||||
# under that path get shadowed by the volume on existing deployments.
|
||||
# Place errorfiles outside the volumed path; the HAProxy config references
|
||||
# them by absolute path.
|
||||
COPY errors /haproxy/errors
|
||||
RUN chmod +x /haproxy/scripts/*
|
||||
RUN pip install -r requirements.txt
|
||||
# Create log directories
|
||||
@@ -16,7 +47,9 @@ RUN mkdir -p /var/spool/cron/crontabs && \
|
||||
echo '0 */12 * * * /haproxy/scripts/renew-certificates.sh >> /var/log/haproxy-manager.log 2>&1' >> /var/spool/cron/crontabs/root && \
|
||||
chmod 600 /var/spool/cron/crontabs/root && \
|
||||
chown root:crontab /var/spool/cron/crontabs/root
|
||||
EXPOSE 80 443 8000
|
||||
# 443/udp carries HTTP/3 (QUIC). EXPOSE is documentation only — the container
|
||||
# must still be run with `-p 443:443/udp` for the UDP listener to be reachable.
|
||||
EXPOSE 80 443 443/udp 8000
|
||||
# Add health check
|
||||
HEALTHCHECK --interval=30s --timeout=10s --start-period=60s --retries=3 \
|
||||
CMD curl -sf --max-time 5 http://localhost:8000/health && curl -s --max-time 5 -o /dev/null http://localhost/ || exit 1
|
||||
|
||||
@@ -6,10 +6,10 @@ A Flask-based API service for managing HAProxy configurations with dynamic SSL c
|
||||
To run the container:
|
||||
```bash
|
||||
# Without API key authentication (default)
|
||||
docker run -d -p 80:80 -p 443:443 -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy --name haproxy-manager repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
docker run -d -p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy --name haproxy-manager your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
|
||||
# With API key authentication (recommended for production)
|
||||
docker run -d -p 80:80 -p 443:443 -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy -e HAPROXY_API_KEY=your-secure-api-key-here --name haproxy-manager repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
docker run -d -p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 -v lets-encrypt:/etc/letsencrypt -v haproxy:/etc/haproxy -e HAPROXY_API_KEY=your-secure-api-key-here --name haproxy-manager your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
```
|
||||
|
||||
## Features
|
||||
@@ -394,7 +394,7 @@ You can customize the default page by setting environment variables:
|
||||
|
||||
```bash
|
||||
docker run -d \
|
||||
-p 80:80 -p 443:443 -p 8000:8000 \
|
||||
-p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 \
|
||||
-v lets-encrypt:/etc/letsencrypt \
|
||||
-v haproxy:/etc/haproxy \
|
||||
-e HAPROXY_API_KEY=your-secure-api-key-here \
|
||||
@@ -402,7 +402,7 @@ docker run -d \
|
||||
-e HAPROXY_DEFAULT_MAIN_MESSAGE="This website is currently under construction and will be available soon." \
|
||||
-e HAPROXY_DEFAULT_SECONDARY_MESSAGE="Please check back later or contact us for more information." \
|
||||
--name haproxy-manager \
|
||||
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
```
|
||||
|
||||
## Example Usage
|
||||
@@ -411,12 +411,12 @@ docker run -d \
|
||||
```bash
|
||||
# Start container with API key
|
||||
docker run -d \
|
||||
-p 80:80 -p 443:443 -p 8000:8000 \
|
||||
-p 80:80 -p 443:443 -p 443:443/udp -p 8000:8000 \
|
||||
-v lets-encrypt:/etc/letsencrypt \
|
||||
-v haproxy:/etc/haproxy \
|
||||
-e HAPROXY_API_KEY=your-secure-api-key-here \
|
||||
--name haproxy-manager \
|
||||
repo.anhonesthost.net/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
your-registry.example.com/cloud-hosting-platform/haproxy-manager-base:latest
|
||||
|
||||
# Add a domain
|
||||
curl -X POST http://localhost:8000/api/domain \
|
||||
|
||||
@@ -0,0 +1,70 @@
|
||||
# Coraza-SPOA sidecar for haproxy-manager.
|
||||
#
|
||||
# Layout: built from upstream source. main.go is at the repo root; CRS rules
|
||||
# are bundled into the binary at build time (referenced as @owasp_crs/), so
|
||||
# the CRS version is whatever ships with the pinned coraza-spoa tag.
|
||||
#
|
||||
# Pin: review the upstream CHANGELOG (https://github.com/corazawaf/coraza-spoa/releases)
|
||||
# before bumping. New tags can ship newer CRS, which can introduce new rules
|
||||
# whose IDs fall into the "enforce day-one" ranges in overrides.conf — verify
|
||||
# those are still high-confidence before promoting a new tag to prod.
|
||||
|
||||
ARG CORAZA_SPOA_VERSION=v0.7.1
|
||||
|
||||
# golang:1.25 from the in-house mirror. The 2026-05-12 Cloudflare incident
|
||||
# took out docker.io blob pulls TWICE in one day (first for python:3.12-slim,
|
||||
# then for this image's golang:1.25), so both are mirrored at
|
||||
# repo.anhonesthost.net via the .gitea/workflows/mirror-base-image.yaml
|
||||
# weekly job.
|
||||
FROM repo.anhonesthost.net/cloud-hosting-platform/golang:1.25 AS build
|
||||
ARG CORAZA_SPOA_VERSION
|
||||
WORKDIR /src
|
||||
RUN apt-get update \
|
||||
&& apt-get install -y --no-install-recommends git \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
RUN git clone --depth 1 --branch "${CORAZA_SPOA_VERSION}" \
|
||||
https://github.com/corazawaf/coraza-spoa.git . \
|
||||
&& go mod download \
|
||||
&& CGO_ENABLED=0 go build -trimpath -ldflags='-s -w' -o /out/coraza-spoa .
|
||||
|
||||
# Catalog extractor: walks the bundled CRS at build time and emits
|
||||
# rules-catalog.json so WHP's UI can render rule metadata without parsing
|
||||
# .conf files at runtime. Uses the SAME coraza-coreruleset version pin as
|
||||
# the coraza-spoa binary above (drift between the two would mislabel rules).
|
||||
FROM repo.anhonesthost.net/cloud-hosting-platform/golang:1.25 AS catalog
|
||||
WORKDIR /src
|
||||
COPY catalog-extractor/ .
|
||||
RUN go build -trimpath -o /out/catalog-extractor . \
|
||||
&& /out/catalog-extractor > /out/rules-catalog.json
|
||||
|
||||
# Distroless runtime: no shell, no package manager, no /tmp by default —
|
||||
# smallest attack surface for an exposed service. Audit log directory is
|
||||
# bind-mounted; coraza-spoa writes to it via direct file I/O (no shell needed).
|
||||
FROM gcr.io/distroless/static-debian12:nonroot
|
||||
|
||||
LABEL org.opencontainers.image.title="coraza-spoa-whp" \
|
||||
org.opencontainers.image.description="Coraza WAF SPOA agent configured for WHP haproxy-manager integration" \
|
||||
org.opencontainers.image.source="https://github.com/shadowdao/haproxy-manager-base" \
|
||||
org.opencontainers.image.licenses="MIT"
|
||||
|
||||
COPY --from=build /out/coraza-spoa /coraza-spoa
|
||||
COPY config.yaml /etc/coraza-spoa/config.yaml
|
||||
COPY overrides.conf /etc/coraza/overrides.conf
|
||||
COPY pre-overrides.conf /etc/coraza/pre-overrides.conf
|
||||
COPY local-overrides.conf /etc/coraza/local-overrides.conf
|
||||
COPY host-exceptions/ /etc/coraza/host-exceptions/
|
||||
COPY --from=catalog /out/rules-catalog.json /etc/coraza/rules-catalog.json
|
||||
|
||||
# Audit log directory — bind-mount /var/log/coraza:/var/log/coraza from host
|
||||
# so logs persist across container restarts and AI Monitor can tail them.
|
||||
# Distroless nonroot user has UID 65532; the host directory must be writable
|
||||
# by that UID (install script will chown it appropriately).
|
||||
VOLUME ["/var/log/coraza"]
|
||||
|
||||
# SPOE TCP port — bound on 0.0.0.0:9000 inside the container. The host-side
|
||||
# port mapping is controlled by `docker run -p` (typically not exposed beyond
|
||||
# the internal docker network, since haproxy-manager reaches it by container
|
||||
# name on client-net).
|
||||
EXPOSE 9000
|
||||
|
||||
ENTRYPOINT ["/coraza-spoa", "--config", "/etc/coraza-spoa/config.yaml"]
|
||||
@@ -0,0 +1,78 @@
|
||||
# coraza-spoa sidecar
|
||||
|
||||
A sidecar container that runs [Coraza-SPOA](https://github.com/corazawaf/coraza-spoa) as a WAF engine for `haproxy-manager`. HAProxy consults it per-request via the SPOE/SPOP protocol; Coraza evaluates the request against OWASP CRS rules and tells HAProxy whether to allow or block.
|
||||
|
||||
## Design constraints
|
||||
|
||||
- **`haproxy-manager` does NOT depend on this sidecar.** The base image works standalone (used in other projects and home networks) without WAF. SPOE config in the generated `haproxy.cfg` is opt-in via an env var on `haproxy-manager`.
|
||||
- **Fail-open when the sidecar is unhealthy.** `option set-on-error continue` in the HAProxy SPOE config means request flow continues uninspected if coraza-spoa is unreachable, rather than 503-ing customer traffic.
|
||||
- **Detect-only globally; enforce explicitly.** See `overrides.conf` for the day-one enforce list. Most CRS rules log without blocking until we've tuned per-customer false positives.
|
||||
|
||||
## Deployment shape
|
||||
|
||||
Two containers per host, both on the `client-net` docker network:
|
||||
|
||||
```
|
||||
haproxy-manager (existing) — ports 80, 443, 8000
|
||||
│ SPOE TCP/9000 → reach coraza-spoa by container DNS
|
||||
▼
|
||||
coraza-spoa (this image)
|
||||
port 9000 (SPOE) — NOT exposed on host; internal network only
|
||||
/var/log/coraza — bind-mounted to host for AI Monitor consumption
|
||||
```
|
||||
|
||||
Typical `docker run`:
|
||||
|
||||
```bash
|
||||
mkdir -p /var/log/coraza
|
||||
chown 65532:65532 /var/log/coraza # distroless nonroot UID
|
||||
|
||||
docker run -d \
|
||||
--name coraza-spoa \
|
||||
--network client-net \
|
||||
--restart unless-stopped \
|
||||
-v /var/log/coraza:/var/log/coraza \
|
||||
your-registry.example.com/cloud-hosting-platform/coraza-spoa:latest
|
||||
```
|
||||
|
||||
Then on the `haproxy-manager` container, add the env var:
|
||||
|
||||
```
|
||||
-e HAPROXY_CORAZA_SPOE_BACKEND=coraza-spoa:9000
|
||||
```
|
||||
|
||||
The haproxy-manager template engine sees the env var and renders the SPOE config block pointing at this sidecar. Without the env var, no SPOE blocks render — the haproxy-manager image's behavior is unchanged.
|
||||
|
||||
## Files
|
||||
|
||||
| File | Purpose |
|
||||
|---|---|
|
||||
| `Dockerfile` | Multi-stage build (golang:1.25 → distroless), pinned to upstream coraza-spoa tag |
|
||||
| `config.yaml` | SPOA listener config + one named application `haproxy` |
|
||||
| `overrides.conf` | Day-one enforce list (`ctl:ruleEngine=On` for high-confidence rule IDs) |
|
||||
| `README.md` | This file |
|
||||
|
||||
## Audit log
|
||||
|
||||
`/var/log/coraza/audit.log` — JSON, one event per line, RelevantOnly (only requests that triggered ≥1 rule are logged). AI Monitor should be configured to tail this on each host.
|
||||
|
||||
Entries include rule IDs, matched patterns, request metadata, and action taken (`log` for detect-only, `deny` for enforced). Use the JSON `action` field to filter blocked vs. observed.
|
||||
|
||||
## Upgrading the pin
|
||||
|
||||
CRS rules are bundled into the coraza-spoa binary at build time, so the CRS version is whatever ships with the pinned coraza-spoa tag. To upgrade:
|
||||
|
||||
1. Check upstream releases: <https://github.com/corazawaf/coraza-spoa/releases>
|
||||
2. Skim the CHANGELOG for new/changed rules in the `overrides.conf` ID ranges.
|
||||
3. Bump `ARG CORAZA_SPOA_VERSION` in the Dockerfile.
|
||||
4. Push to `main` — the Gitea workflow at `.gitea/workflows/build-push-coraza.yaml` rebuilds + pushes `:latest`.
|
||||
5. On each host, run `container-manager.sh recreate coraza-spoa` to pull the new image.
|
||||
|
||||
## Tuning false positives
|
||||
|
||||
When a legitimate request triggers a blocked rule, the audit log shows the rule ID. Two ways to silence it:
|
||||
|
||||
1. **Per-rule exception** in `overrides.conf`: `SecRuleRemoveById <id>` (full disable) or `SecRuleRemoveTargetById <id> "<target>"` (targeted exception).
|
||||
2. **Drop from the enforce list**: remove the rule's ID range from the `ctl:ruleEngine=On` overrides; it falls back to detect-only.
|
||||
|
||||
After tuning, push the change — CI rebuilds, then `recreate coraza-spoa` on each host to apply.
|
||||
@@ -0,0 +1,7 @@
|
||||
module catalog-extractor
|
||||
|
||||
go 1.23
|
||||
|
||||
require github.com/corazawaf/coraza-coreruleset/v4 v4.25.0
|
||||
|
||||
require github.com/magefile/mage v1.17.0 // indirect
|
||||
@@ -0,0 +1,12 @@
|
||||
github.com/corazawaf/coraza-coreruleset/v4 v4.25.0 h1:tqFO1lfVpTiyWtlN618OXpZMfw+nnN0Q4///W5W+/HM=
|
||||
github.com/corazawaf/coraza-coreruleset/v4 v4.25.0/go.mod h1:nRuGXITxOPvsLF2VxaTB7pYok8QB8BitX3ZenXcUryY=
|
||||
github.com/davecgh/go-spew v1.1.1 h1:vj9j/u1bqnvCEfJOwUhtlOARqs3+rkHYY13jYWTU97c=
|
||||
github.com/davecgh/go-spew v1.1.1/go.mod h1:J7Y8YcW2NihsgmVo/mv3lAwl/skON4iLHjSsI+c5H38=
|
||||
github.com/magefile/mage v1.17.0 h1:dS4tkq997Ism03akafC8509iqDjeE7TNTexI25Y7sXM=
|
||||
github.com/magefile/mage v1.17.0/go.mod h1:Yj51kqllmsgFpvvSzgrZPK9WtluG3kUhFaBUVLo4feA=
|
||||
github.com/pmezard/go-difflib v1.0.0 h1:4DBwDE0NGyQoBHbLQYPwSUPoCMWR5BEzIk/f1lZbAQM=
|
||||
github.com/pmezard/go-difflib v1.0.0/go.mod h1:iKH77koFhYxTK1pcRnkKkqfTogsbg7gZNVY4sRDYZ/4=
|
||||
github.com/stretchr/testify v1.11.1 h1:7s2iGBzp5EwR7/aIZr8ao5+dra3wiQyKjjFuvgVKu7U=
|
||||
github.com/stretchr/testify v1.11.1/go.mod h1:wZwfW3scLgRK+23gO65QZefKpKQRnfz6sD981Nm4B6U=
|
||||
gopkg.in/yaml.v3 v3.0.1 h1:fxVm/GzAzEWqLHuvctI91KS9hhNmmWOoWu0XTYJS7CA=
|
||||
gopkg.in/yaml.v3 v3.0.1/go.mod h1:K4uyk7z7BCEPqu6E+C64Yfv1cQ7kz7rIZviUmN+EgEM=
|
||||
@@ -0,0 +1,80 @@
|
||||
package main
|
||||
|
||||
import (
|
||||
"encoding/json"
|
||||
"fmt"
|
||||
"io/fs"
|
||||
"os"
|
||||
"regexp"
|
||||
"strconv"
|
||||
"strings"
|
||||
|
||||
crs "github.com/corazawaf/coraza-coreruleset/v4"
|
||||
)
|
||||
|
||||
type Rule struct {
|
||||
ID int `json:"id"`
|
||||
Msg string `json:"msg"`
|
||||
Severity string `json:"severity"`
|
||||
Tags []string `json:"tags"`
|
||||
File string `json:"file"`
|
||||
}
|
||||
|
||||
var (
|
||||
idRe = regexp.MustCompile(`(?i)\bid:'?(\d+)'?`)
|
||||
msgRe = regexp.MustCompile(`(?i)\bmsg:'([^']+)'`)
|
||||
severityRe = regexp.MustCompile(`(?i)\bseverity:'?([A-Z]+)'?`)
|
||||
tagRe = regexp.MustCompile(`(?i)\btag:'([^']+)'`)
|
||||
)
|
||||
|
||||
func main() {
|
||||
out := []Rule{}
|
||||
err := fs.WalkDir(crs.FS, ".", func(path string, d fs.DirEntry, err error) error {
|
||||
if err != nil || d.IsDir() {
|
||||
return err
|
||||
}
|
||||
if !strings.HasSuffix(path, ".conf") {
|
||||
return nil
|
||||
}
|
||||
b, err := fs.ReadFile(crs.FS, path)
|
||||
if err != nil {
|
||||
return err
|
||||
}
|
||||
// Coalesce backslash-continuation lines so id/msg/etc on the same
|
||||
// logical rule are visible to the per-line scanner.
|
||||
text := regexp.MustCompile(`\\\s*\n\s*`).ReplaceAllString(string(b), " ")
|
||||
for _, line := range strings.Split(text, "\n") {
|
||||
line = strings.TrimSpace(line)
|
||||
if !strings.HasPrefix(line, "SecRule") && !strings.HasPrefix(line, "SecAction") {
|
||||
continue
|
||||
}
|
||||
m := idRe.FindStringSubmatch(line)
|
||||
if m == nil {
|
||||
continue
|
||||
}
|
||||
id, _ := strconv.Atoi(m[1])
|
||||
r := Rule{ID: id, File: path}
|
||||
if mm := msgRe.FindStringSubmatch(line); mm != nil {
|
||||
r.Msg = mm[1]
|
||||
}
|
||||
if mm := severityRe.FindStringSubmatch(line); mm != nil {
|
||||
r.Severity = strings.ToLower(mm[1])
|
||||
}
|
||||
for _, mm := range tagRe.FindAllStringSubmatch(line, -1) {
|
||||
r.Tags = append(r.Tags, mm[1])
|
||||
}
|
||||
out = append(out, r)
|
||||
}
|
||||
return nil
|
||||
})
|
||||
if err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
enc := json.NewEncoder(os.Stdout)
|
||||
enc.SetIndent("", " ")
|
||||
if err := enc.Encode(out); err != nil {
|
||||
fmt.Fprintln(os.Stderr, err)
|
||||
os.Exit(1)
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,72 @@
|
||||
# Coraza-SPOA configuration for WHP haproxy-manager integration.
|
||||
#
|
||||
# One named application "haproxy" — the haproxy-manager spoe template
|
||||
# references this same name in its spoe-agent block, so the SPOA knows
|
||||
# which rules to apply when HAProxy dispatches a request.
|
||||
#
|
||||
# Mode: SecRuleEngine DetectionOnly globally; overrides.conf promotes
|
||||
# specific high-confidence rule ID ranges to enforcement individually.
|
||||
# This is the safest posture for v1 — every rule logs, but only the
|
||||
# unambiguous ones (scanner UAs, RCE, LFI, webshells, Log4Shell) block.
|
||||
|
||||
bind: 0.0.0.0:9000
|
||||
|
||||
# Process-level logging (separate from per-request audit logging below)
|
||||
log_level: info
|
||||
log_file: /dev/stdout
|
||||
log_format: json
|
||||
|
||||
# Fallback when the request doesn't match a named application — we only
|
||||
# have one, so it's also the default.
|
||||
default_application: haproxy
|
||||
|
||||
applications:
|
||||
- name: haproxy
|
||||
directives: |
|
||||
# CRS-bundled defaults: recommended Coraza settings + CRS setup +
|
||||
# the rule pack itself (~16 MB of rules embedded in the binary).
|
||||
Include @coraza.conf-recommended
|
||||
Include @crs-setup.conf.example
|
||||
|
||||
# Runtime-managed PRE-CRS exclusions written by WHP UI. Empty by default.
|
||||
# Loaded BEFORE the CRS rules so per-host ctl:ruleRemoveById exemptions
|
||||
# fire in phase:1 BEFORE the CRS rule they're trying to exempt would
|
||||
# otherwise match. Server-wide overrides live in local-overrides.conf
|
||||
# (loaded after CRS) instead.
|
||||
Include /etc/coraza/pre-overrides.conf
|
||||
|
||||
Include @owasp_crs/*.conf
|
||||
|
||||
# WHP-specific overrides — day-one enforce list, plus tuning for
|
||||
# the customer mix (WordPress, WooCommerce, Divi). Read this file
|
||||
# to see exactly what blocks vs what's detect-only.
|
||||
Include /etc/coraza/overrides.conf
|
||||
|
||||
# Runtime-managed POST-CRS overrides written by WHP UI. Empty by default.
|
||||
Include /etc/coraza/local-overrides.conf
|
||||
|
||||
# Global mode: log all alerts, block only what overrides.conf
|
||||
# explicitly promotes via ctl:ruleEngine=On.
|
||||
SecRuleEngine DetectionOnly
|
||||
|
||||
# Audit log: JSON to a bind-mounted file so AI Monitor + log
|
||||
# rotation can pick it up. RelevantOnly means we don't log every
|
||||
# passing request, only ones that triggered at least one rule.
|
||||
SecAuditEngine RelevantOnly
|
||||
SecAuditLog /var/log/coraza/audit.log
|
||||
SecAuditLogFormat JSON
|
||||
SecAuditLogParts ABIJDEFHKZ
|
||||
|
||||
# HAProxy sends request-only events for v1. Response inspection adds
|
||||
# latency on every page render with marginal additional protection
|
||||
# for our customer mix; can be turned on later if we want it.
|
||||
response_check: false
|
||||
|
||||
# Transactions cache for 60s. SPOE protocol is fire-and-forget per
|
||||
# request, so this is just how long Coraza holds context for any
|
||||
# multi-stage processing.
|
||||
transaction_ttl_ms: 60000
|
||||
|
||||
log_level: info
|
||||
log_file: /var/log/coraza/spoa.log
|
||||
log_format: json
|
||||
@@ -0,0 +1,3 @@
|
||||
# AUTOGENERATED by WHP — do not hand-edit.
|
||||
# Source of truth: whp.security_db coraza_rule_overrides table.
|
||||
# Empty file = no runtime overrides; baked-in overrides.conf governs.
|
||||
@@ -0,0 +1,105 @@
|
||||
# WHP day-one enforce overrides for coraza-spoa.
|
||||
#
|
||||
# Global mode in config.yaml is SecRuleEngine DetectionOnly. The rule ID
|
||||
# ranges below are promoted to enforcement individually, chosen for very
|
||||
# low false-positive rate on the kinds of customer traffic seen on WHP
|
||||
# (WordPress, WooCommerce, Divi page builders).
|
||||
#
|
||||
# When bumping the upstream coraza-spoa pin (and thus the bundled CRS):
|
||||
# 1. Skim the CRS CHANGELOG for new/changed rules in these ID ranges.
|
||||
# 2. Verify they're still high-confidence before promoting the new image.
|
||||
# 3. Smoke-test in staging detect-only mode for 24h before flipping enforce.
|
||||
#
|
||||
# Per-customer false-positive tuning lives in a future per-customer
|
||||
# override mechanism; v1 is server-wide.
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 930120 — LFI: explicit traversal to sensitive system files
|
||||
# (/etc/passwd, /proc/self/, /.ssh/, /etc/shadow, /etc/group, etc.)
|
||||
# Unambiguous probe pattern; no legitimate site path leads here.
|
||||
# Note: 930xxx as a whole includes broader traversal patterns that can FP
|
||||
# on legitimate relative-path file browsers — keep those detect-only.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 930120 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 932100-932160 — RCE: Unix shell command injection
|
||||
# Patterns like `; cat /etc/passwd`, `|whoami`, backtick `\`uname\``,
|
||||
# $(...) substitution, &&/|| chaining with shell builtins.
|
||||
# Don't appear in normal POST bodies, URL params, or headers. Targeting
|
||||
# these is unambiguous attempted command execution.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 932100-932160 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 933170-933200 — PHP Webshell access patterns
|
||||
# Direct requests to known webshell paths: c99.php, r57.php, b374k.php,
|
||||
# wso.php, alfa.php, mini.php, etc. Almost universally reconnaissance
|
||||
# scanning for post-exploitation. Even legitimate WordPress installs
|
||||
# never serve these paths.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 933170-933200 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 944100-944300 — Log4Shell / JNDI injection
|
||||
# `${jndi:ldap://}`, `${jndi:rmi://}`, and obfuscated variants thereof
|
||||
# in headers, query strings, or bodies. Even our PHP/Node stack isn't
|
||||
# vulnerable, but blocking at the edge keeps logs clean and protects
|
||||
# any future Java workloads.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 944100-944300 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 920440 — URL file extension restricted by policy
|
||||
# Catches probes for backup / config / dump files: .bak, .old, .save,
|
||||
# .swp, .sql, .dist, .backup. Promoted to enforce after empirical
|
||||
# observation on whp01 (2026-05-12, first ~30 min of detect-only):
|
||||
# 124 events, all backup-file recon — `/wp-config.php.old`,
|
||||
# `/db_backup.sql`, `/.env.save`, `/releases.sql`, etc. — from a
|
||||
# single GCP-hosted scanner. Zero false positives observed; standard
|
||||
# WP/WooCommerce/Divi/HPR URLs do not end in these extensions.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 920440 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# 930130 — Restricted File Access Attempt
|
||||
# Catches dotfile / VCS / config-disclosure probes: .env (and .env.local /
|
||||
# .env.bak / .env.save variants), .git/config, config.php at root or under
|
||||
# /admin /backend, etc. Distinct from 930120 (system file paths like
|
||||
# /etc/passwd); this targets application secret files.
|
||||
#
|
||||
# Promoted to enforce on the same observation pass that justified 920440:
|
||||
# 117 events split across joshuaknapp.net (136), cgdannyb.com (51),
|
||||
# onlinesupplements.net (23) — all `.env`-class disclosure probes.
|
||||
# Zero false positives observed. Notably, HPR's `/ccdn.php?filename=...`
|
||||
# audio delivery path does NOT trigger this rule — verified empirically.
|
||||
# ---------------------------------------------------------------------------
|
||||
SecRuleUpdateActionById 930130 "ctl:ruleEngine=On"
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# Rule families intentionally kept at DETECT-ONLY for v1 — high FP rate
|
||||
# on customer mix. Promote individually after observation:
|
||||
#
|
||||
# 913xxx (Scanner UAs)— matches legitimate ActivityPub federation
|
||||
# (Mastodon's "...Bot" UA) and SiteLockSpider (a
|
||||
# paid customer-security service some sites use).
|
||||
# Observed on whp01 burn-in 2026-05-13:
|
||||
# 20/185 hits = ~11% FP rate on HPR + greggfranklin
|
||||
# + suchascream. Detection adds anomaly score
|
||||
# either way; enforce upside is low.
|
||||
# 941xxx (XSS) — Divi rich-text editor saves, TinyMCE submissions
|
||||
# 942xxx (SQLi) — WP admin queries reflected in params
|
||||
# 920xxx (other) — most 920xxx rules; 920440 specifically promoted above
|
||||
# 933150 — PHP injection FP on WooCommerce checkout
|
||||
# (`session_start` literal appearing in billing form data)
|
||||
# 950xxx-953xxx — Data leakage / backup-file disclosure (mixed FP)
|
||||
# ---------------------------------------------------------------------------
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
# RESERVED RULE-ID RANGE: 990000000 – 990999999
|
||||
# WHP's coraza_rule_manager generates per-host-exception rules in this range
|
||||
# (rule ID = 990000000 + target_rule_id). Do NOT add new rules in this range
|
||||
# from any other source. When bumping the coraza-spoa pin, check the CRS
|
||||
# changelog for new rules with 9-digit IDs (rare but possible) and re-namespace
|
||||
# if collision risk emerges.
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -0,0 +1,3 @@
|
||||
# AUTOGENERATED by WHP — do not hand-edit.
|
||||
# Source of truth: whp.security_db coraza_rule_host_exceptions table.
|
||||
# Loaded BEFORE the CRS rules. Empty file = no per-host exemptions active.
|
||||
@@ -0,0 +1,109 @@
|
||||
<!DOCTYPE html>
|
||||
<!--
|
||||
Served by HAProxy via `lf-file` on Coraza WAF deny.
|
||||
IMPORTANT: HAProxy's lf-file expansion treats `%` as the start of a
|
||||
log-format expression. Literal percent signs (CSS 100%, gradient stops,
|
||||
url-encoded data, etc.) MUST be doubled as `%%` or HAProxy will silently
|
||||
swallow them. Expressions like `%[unique-id]` / `%[req.hdr(host)]` stay
|
||||
single-`%` — those are the substitutions we want.
|
||||
-->
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="utf-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1">
|
||||
<meta name="robots" content="noindex, nofollow">
|
||||
<title>Request blocked · %[req.hdr(host)]</title>
|
||||
<style>
|
||||
*, *::before, *::after { box-sizing: border-box; }
|
||||
html, body { margin: 0; padding: 0; height: 100%%; }
|
||||
body {
|
||||
font-family: -apple-system, BlinkMacSystemFont, "Segoe UI", Roboto, "Helvetica Neue", Arial, sans-serif;
|
||||
color: #1f2937;
|
||||
background: linear-gradient(135deg, #f9fafb 0%%, #eef2f7 100%%);
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
padding: 24px;
|
||||
line-height: 1.5;
|
||||
}
|
||||
.card {
|
||||
background: #fff;
|
||||
border-radius: 12px;
|
||||
box-shadow: 0 1px 3px rgba(0,0,0,0.05), 0 12px 32px rgba(31,41,55,0.08);
|
||||
max-width: 560px;
|
||||
width: 100%%;
|
||||
padding: 36px 40px;
|
||||
}
|
||||
.badge {
|
||||
display: inline-block;
|
||||
background: #fef3c7;
|
||||
color: #92400e;
|
||||
font-size: 12px;
|
||||
font-weight: 600;
|
||||
letter-spacing: 0.04em;
|
||||
text-transform: uppercase;
|
||||
padding: 4px 10px;
|
||||
border-radius: 999px;
|
||||
margin-bottom: 16px;
|
||||
}
|
||||
h1 { font-size: 22px; margin: 0 0 12px; color: #111827; }
|
||||
p { margin: 0 0 14px; color: #374151; }
|
||||
.ref {
|
||||
background: #f3f4f6;
|
||||
border: 1px solid #e5e7eb;
|
||||
border-radius: 6px;
|
||||
padding: 12px 14px;
|
||||
margin: 20px 0;
|
||||
font-family: ui-monospace, "SF Mono", Menlo, Consolas, monospace;
|
||||
font-size: 13px;
|
||||
color: #111827;
|
||||
word-break: break-all;
|
||||
}
|
||||
.ref-label {
|
||||
display: block;
|
||||
color: #6b7280;
|
||||
font-size: 11px;
|
||||
font-weight: 600;
|
||||
letter-spacing: 0.05em;
|
||||
text-transform: uppercase;
|
||||
margin-bottom: 4px;
|
||||
font-family: inherit;
|
||||
}
|
||||
.owner {
|
||||
border-top: 1px solid #e5e7eb;
|
||||
margin-top: 24px;
|
||||
padding-top: 20px;
|
||||
color: #4b5563;
|
||||
font-size: 14px;
|
||||
}
|
||||
.owner h2 { font-size: 14px; font-weight: 600; color: #111827; margin: 0 0 8px; }
|
||||
a {
|
||||
color: #1d4ed8;
|
||||
text-decoration: none;
|
||||
border-bottom: 1px solid transparent;
|
||||
}
|
||||
a:hover, a:focus { border-bottom-color: #1d4ed8; }
|
||||
.small { font-size: 12px; color: #6b7280; margin-top: 16px; }
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<main class="card" role="main">
|
||||
<span class="badge">Access blocked</span>
|
||||
<h1>Your request was blocked by our security filter</h1>
|
||||
<p>The request to <strong>%[req.hdr(host)]</strong> looked suspicious to our web application firewall and was not delivered to the site.</p>
|
||||
<p>This is automated. No one has reviewed the request yet.</p>
|
||||
|
||||
<div class="ref">
|
||||
<span class="ref-label">Request reference</span>
|
||||
%[unique-id]
|
||||
</div>
|
||||
|
||||
<div class="owner">
|
||||
<h2>Site owner?</h2>
|
||||
<p>If you operate this site and believe this block is incorrect, please contact your hosting provider's support team and include the request reference above. They can look up exactly which rule fired and adjust it if it's a false positive.</p>
|
||||
</div>
|
||||
|
||||
<p class="small">Reference IDs expire from our active logs after 14 days, so please open a ticket promptly if you'd like this investigated.</p>
|
||||
</main>
|
||||
</body>
|
||||
</html>
|
||||
+843
-83
File diff suppressed because it is too large
Load Diff
@@ -1,3 +1,7 @@
|
||||
Flask==2.3.3
|
||||
Jinja2==3.1.2
|
||||
psutil
|
||||
# Production WSGI server. Replaces Flask's built-in werkzeug dev server, which
|
||||
# is single-threaded and leaks workers over long uptimes (root cause of the
|
||||
# 2026-05 haproxy-manager "healthy but stalled" incidents).
|
||||
gunicorn==23.0.0
|
||||
|
||||
@@ -0,0 +1,52 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Idempotent haproxy liveness check — driven by the in-container supervisor loop.
|
||||
|
||||
Why this exists
|
||||
---------------
|
||||
haproxy runs as a *background child of PID 1* (gunicorn) — it is started once at
|
||||
container init (scripts/init.py -> do_initial_setup -> start_haproxy) and then
|
||||
left running. Nothing supervises it after that. If the haproxy master process
|
||||
dies mid-life (SIGABRT -> exit 134, segfault, or an OOM of the haproxy master),
|
||||
the container stays "up" because gunicorn is still PID 1, so Docker's
|
||||
`--restart` policy never fires. haproxy then stays down until the *external*
|
||||
host watchdog (haproxy-watchdog.sh) notices port 80 is dead for ~3 minutes and
|
||||
does a full `docker restart` — which drops every in-flight connection.
|
||||
|
||||
This script closes that gap: called on a short interval by the supervisor loop
|
||||
in start-up.sh, it re-launches haproxy *in place* within one interval.
|
||||
|
||||
Safety
|
||||
------
|
||||
start_haproxy() is guarded by `is_process_running('haproxy')` (psutil-based, so
|
||||
it works in this container which has no `ps`), so calling this while haproxy is
|
||||
healthy is a cheap no-op. It only ever acts when haproxy is genuinely gone.
|
||||
"""
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, '/haproxy')
|
||||
import haproxy_manager # noqa: E402 (sys.path manipulation must come first)
|
||||
|
||||
|
||||
def main():
|
||||
if haproxy_manager.is_process_running('haproxy'):
|
||||
return 0
|
||||
|
||||
haproxy_manager.logger.warning(
|
||||
"[haproxy-supervisor] haproxy process not found — attempting in-place restart"
|
||||
)
|
||||
# start_haproxy() validates the config (and regenerates it if invalid)
|
||||
# before launching, and swallows its own errors, so it will not raise here.
|
||||
haproxy_manager.start_haproxy()
|
||||
|
||||
if haproxy_manager.is_process_running('haproxy'):
|
||||
haproxy_manager.logger.info("[haproxy-supervisor] haproxy restarted in place")
|
||||
return 0
|
||||
|
||||
haproxy_manager.logger.error(
|
||||
"[haproxy-supervisor] haproxy restart FAILED — still not running after start_haproxy()"
|
||||
)
|
||||
return 1
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
sys.exit(main())
|
||||
Executable
+13
@@ -0,0 +1,13 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Container init: DB schema, certbot account, config generation, HAProxy start.
|
||||
|
||||
Runs once per container start, BEFORE gunicorn workers spawn. Keeping init out
|
||||
of the WSGI app's module-load path avoids fork-time races (multiple workers
|
||||
attempting to start_haproxy() simultaneously, certbot lock contention, etc.).
|
||||
"""
|
||||
import sys
|
||||
|
||||
sys.path.insert(0, '/haproxy')
|
||||
import haproxy_manager # noqa: E402 (sys.path manipulation must come first)
|
||||
|
||||
haproxy_manager.do_initial_setup()
|
||||
@@ -6,22 +6,40 @@
|
||||
SOCKET="/tmp/haproxy-cli"
|
||||
MAP_FILE="/etc/haproxy/blocked_ips.map"
|
||||
|
||||
# HAProxy runs in master-worker mode here, and /tmp/haproxy-cli is the MASTER
|
||||
# socket. Data-plane commands (map/table manipulation) are NOT accepted on the
|
||||
# master socket — they must be routed to a worker with the "@<n>" prefix. "@1"
|
||||
# targets the current active worker. (A bare "add map ..." on the master socket
|
||||
# fails with "Unknown command: 'add'".)
|
||||
cli() { printf '@1 %s\n' "$*" | socat stdio "$SOCKET"; }
|
||||
|
||||
# Map lookup in haproxy.cfg is `map_ip(...,0) -m int gt 0`, so each entry MUST be
|
||||
# "<ip_or_cidr> 1" — a bare IP yields an empty value (0) and is NOT blocked once
|
||||
# the map file is re-read on reload. The runtime map and the file must agree.
|
||||
MAP_VALUE=1
|
||||
|
||||
# Ensure map file exists
|
||||
if [ ! -f "$MAP_FILE" ]; then
|
||||
touch "$MAP_FILE"
|
||||
echo "# Blocked IPs - Format: IP_ADDRESS" > "$MAP_FILE"
|
||||
echo "# Blocked IPs - Format: <ip_or_cidr> 1 (one per line)" > "$MAP_FILE"
|
||||
fi
|
||||
|
||||
# Escape regex metacharacters (notably dots) in an IP/CIDR for anchored matching.
|
||||
esc_re() { printf '%s' "$1" | sed 's/[.[\*^$/]/\\&/g'; }
|
||||
|
||||
case "$1" in
|
||||
block)
|
||||
if [ -z "$2" ]; then
|
||||
echo "Usage: $0 block IP_ADDRESS"
|
||||
exit 1
|
||||
fi
|
||||
# Add IP to map file
|
||||
grep -q "^$2" "$MAP_FILE" || echo "$2" >> "$MAP_FILE"
|
||||
# Add to runtime map
|
||||
echo "add map /etc/haproxy/blocked_ips.map $2 1" | socat stdio "$SOCKET"
|
||||
re="$(esc_re "$2")"
|
||||
# Persist (idempotent, anchored so 1.2.3.4 doesn't match 1.2.3.45),
|
||||
# always with the trailing value so the block survives a reload.
|
||||
if ! grep -qE "^${re}([[:space:]]|$)" "$MAP_FILE"; then
|
||||
echo "$2 $MAP_VALUE" >> "$MAP_FILE"
|
||||
fi
|
||||
# Apply at runtime immediately (no reload).
|
||||
cli "add map $MAP_FILE $2 $MAP_VALUE"
|
||||
echo "Blocked IP: $2"
|
||||
;;
|
||||
|
||||
@@ -30,31 +48,33 @@ case "$1" in
|
||||
echo "Usage: $0 unblock IP_ADDRESS"
|
||||
exit 1
|
||||
fi
|
||||
# Remove from map file
|
||||
sed -i "/^$2$/d" "$MAP_FILE"
|
||||
# Remove from runtime map
|
||||
echo "del map /etc/haproxy/blocked_ips.map $2" | socat stdio "$SOCKET"
|
||||
re="$(esc_re "$2")"
|
||||
# Remove from map file (match "<ip>" optionally followed by a value).
|
||||
sed -i -E "/^${re}([[:space:]]|$)/d" "$MAP_FILE"
|
||||
# Remove from runtime map.
|
||||
cli "del map $MAP_FILE $2"
|
||||
echo "Unblocked IP: $2"
|
||||
;;
|
||||
|
||||
list)
|
||||
echo "Currently blocked IPs:"
|
||||
echo "show map /etc/haproxy/blocked_ips.map" | socat stdio "$SOCKET" | awk '{print $1}'
|
||||
# `show map` output is "<ptr> <key> <value>" — the IP is field 2.
|
||||
cli "show map $MAP_FILE" | awk 'NF>=2 {print $2}'
|
||||
;;
|
||||
|
||||
clear)
|
||||
echo "Clearing all blocked IPs..."
|
||||
echo "clear map /etc/haproxy/blocked_ips.map" | socat stdio "$SOCKET"
|
||||
echo "# Blocked IPs - Format: IP_ADDRESS" > "$MAP_FILE"
|
||||
cli "clear map $MAP_FILE"
|
||||
echo "# Blocked IPs - Format: <ip_or_cidr> 1 (one per line)" > "$MAP_FILE"
|
||||
echo "All IPs unblocked"
|
||||
;;
|
||||
|
||||
stats)
|
||||
echo "=== HAProxy 3.0.11 Threat Intelligence Dashboard ==="
|
||||
echo "show table web" | socat stdio "$SOCKET" | awk 'NR<=21'
|
||||
cli "show table web" | awk 'NR<=21'
|
||||
echo ""
|
||||
echo "=== Top Threat Scores ==="
|
||||
echo "show table web" | socat stdio "$SOCKET" | awk '
|
||||
cli "show table web" | awk '
|
||||
NR>1 {
|
||||
ip = $1
|
||||
auth_fail = 0
|
||||
@@ -84,7 +104,7 @@ case "$1" in
|
||||
exit 1
|
||||
fi
|
||||
# Add to manual blacklist using GPC(13)
|
||||
echo "set table web key $2 data.gpc(13) 1" | socat stdio "$SOCKET"
|
||||
cli "set table web key $2 data.gpc(13) 1"
|
||||
echo "Manually blacklisted IP: $2 (GPC(13) = 1)"
|
||||
;;
|
||||
|
||||
@@ -94,7 +114,7 @@ case "$1" in
|
||||
exit 1
|
||||
fi
|
||||
# Clear manual blacklist flag
|
||||
echo "set table web key $2 data.gpc(13) 0" | socat stdio "$SOCKET"
|
||||
cli "set table web key $2 data.gpc(13) 0"
|
||||
echo "Removed manual blacklist for IP: $2"
|
||||
;;
|
||||
|
||||
@@ -104,7 +124,7 @@ case "$1" in
|
||||
exit 1
|
||||
fi
|
||||
# Add to auto-blacklist using GPC(14)
|
||||
echo "set table web key $2 data.gpc(14) 1" | socat stdio "$SOCKET"
|
||||
cli "set table web key $2 data.gpc(14) 1"
|
||||
echo "Auto-blacklisted IP: $2 (GPC(14) = 1)"
|
||||
;;
|
||||
|
||||
@@ -115,14 +135,14 @@ case "$1" in
|
||||
fi
|
||||
# Show detailed threat breakdown for specific IP
|
||||
echo "Threat analysis for $2:"
|
||||
echo "show table web key $2" | socat stdio "$SOCKET"
|
||||
cli "show table web key $2"
|
||||
;;
|
||||
|
||||
*)
|
||||
echo "Usage: $0 {block|unblock|list|clear|blacklist|unblacklist|auto-blacklist|threat-score|stats} [IP_ADDRESS]"
|
||||
echo ""
|
||||
echo "HAProxy 3.0.11 Enhanced Security Commands:"
|
||||
echo " block IP - Block IP via map file (immediate)"
|
||||
echo " block IP - Block IP via map file (immediate + persisted)"
|
||||
echo " unblock IP - Unblock IP from map file"
|
||||
echo " blacklist IP - Manual blacklist via GPC(13) array"
|
||||
echo " unblacklist IP - Remove manual blacklist flag"
|
||||
|
||||
@@ -20,12 +20,21 @@ log_error() {
|
||||
|
||||
log_info "Starting certificate renewal process"
|
||||
|
||||
# Run certbot renewal
|
||||
if certbot renew --quiet --no-random-sleep-on-renew; then
|
||||
log_info "Certbot renewal completed"
|
||||
# Run certbot renewal — don't exit on failure, some certs may have
|
||||
# renewed successfully even if others failed (e.g., domain no longer
|
||||
# pointed here). Continue to copy/combine whatever succeeded.
|
||||
CERTBOT_OUTPUT=$(certbot renew --no-random-sleep-on-renew 2>&1)
|
||||
CERTBOT_EXIT=$?
|
||||
|
||||
if [ $CERTBOT_EXIT -eq 0 ]; then
|
||||
log_info "Certbot renewal completed successfully"
|
||||
else
|
||||
log_error "Certbot renewal failed with exit code $?"
|
||||
exit 1
|
||||
log_error "Certbot renewal had failures (exit code $CERTBOT_EXIT):"
|
||||
# Log the specific failures
|
||||
echo "$CERTBOT_OUTPUT" | grep -E "Failed to renew|failure" | while read -r line; do
|
||||
log_error " $line"
|
||||
done
|
||||
log_info "Continuing to process successfully renewed certificates..."
|
||||
fi
|
||||
|
||||
# Copy all certificates to HAProxy format
|
||||
|
||||
Regular → Executable
+79
-2
@@ -1,6 +1,83 @@
|
||||
#!/usr/bin/env bash
|
||||
# Container entrypoint. Two-phase startup:
|
||||
# 1. One-shot init (init.py): DB schema, certbot register, config gen, start HAProxy.
|
||||
# Runs synchronously and to completion so haproxy is up before the API binds.
|
||||
# 2. WSGI serving via gunicorn (replacing the Flask dev server). Two gunicorn
|
||||
# instances:
|
||||
# - port 8080 -> default_app (default page + blocked-ip page; HAProxy
|
||||
# proxies unmatched / blocked traffic here)
|
||||
# - port 8000 -> app (management API)
|
||||
#
|
||||
# Why gunicorn:
|
||||
# Flask's built-in werkzeug "development server" is single-threaded and leaks
|
||||
# workers under sustained load. It carried haproxy-manager for a long time but
|
||||
# stalled out around 24-48h uptime ("healthy" health-check, but every request
|
||||
# queued behind a stuck worker). Gunicorn with --max-requests cycles workers
|
||||
# periodically, which prevents the slow-leak failure mode entirely.
|
||||
|
||||
# Exit on error
|
||||
set -eo pipefail
|
||||
|
||||
# Ensure trusted IP whitelist files exist (volume-mounted /etc/haproxy may shadow image defaults)
|
||||
mkdir -p /etc/haproxy
|
||||
[ -f /etc/haproxy/trusted_ips.list ] || : > /etc/haproxy/trusted_ips.list
|
||||
[ -f /etc/haproxy/trusted_ips.map ] || : > /etc/haproxy/trusted_ips.map
|
||||
|
||||
cron &
|
||||
python /haproxy/haproxy_manager.py
|
||||
|
||||
# Phase 1: container init
|
||||
python /haproxy/scripts/init.py
|
||||
|
||||
# Phase 1.5: in-container haproxy supervisor.
|
||||
# haproxy runs as a background child of PID 1 (gunicorn) with NOTHING watching
|
||||
# it after init. If the haproxy master dies mid-life (e.g. SIGABRT -> exit 134,
|
||||
# segfault), the container stays "up" (gunicorn is PID 1), Docker's --restart
|
||||
# policy never fires, and haproxy is down until the external host watchdog
|
||||
# full-restarts the whole container minutes later (dropping every connection).
|
||||
# This loop revives haproxy in place within one interval. ensure_haproxy.py is
|
||||
# idempotent — a cheap no-op whenever haproxy is already running.
|
||||
HAPROXY_SUPERVISOR_INTERVAL="${HAPROXY_SUPERVISOR_INTERVAL:-15}"
|
||||
(
|
||||
while true; do
|
||||
sleep "${HAPROXY_SUPERVISOR_INTERVAL}"
|
||||
python /haproxy/scripts/ensure_haproxy.py 2>&1 || true
|
||||
done
|
||||
) &
|
||||
|
||||
# Phase 2: WSGI servers
|
||||
# Tunable via env: HAPROXY_MGR_API_WORKERS (default 1), HAPROXY_MGR_API_TIMEOUT
|
||||
# (default 120 — API can do slow ACME calls), HAPROXY_MGR_MAX_REQUESTS (default
|
||||
# 1000 — worker recycle frequency).
|
||||
#
|
||||
# API_WORKERS default is 2 (was 1). A single worker is a single point of
|
||||
# failure: if its gthread pool ever wedges (see the 2026-07-07 subprocess-hang
|
||||
# incident — now bounded by DEFAULT_SUBPROCESS_TIMEOUT in haproxy_manager.py),
|
||||
# the entire management API goes dark. A second worker keeps the API answering
|
||||
# (config regenerate, health, SSL) while the other recycles via --max-requests.
|
||||
API_WORKERS="${HAPROXY_MGR_API_WORKERS:-2}"
|
||||
API_TIMEOUT="${HAPROXY_MGR_API_TIMEOUT:-120}"
|
||||
MAX_REQ="${HAPROXY_MGR_MAX_REQUESTS:-1000}"
|
||||
MAX_REQ_JITTER="${HAPROXY_MGR_MAX_REQUESTS_JITTER:-100}"
|
||||
|
||||
# Default page server on :8080. Stays in the background.
|
||||
# --threads 4 lets one worker handle bursts of blocked-IP/default-page hits
|
||||
# without forking. --max-requests recycles the worker to bound memory drift.
|
||||
gunicorn \
|
||||
--bind 0.0.0.0:8080 \
|
||||
--workers 1 --threads 4 --worker-class gthread \
|
||||
--max-requests "${MAX_REQ}" --max-requests-jitter "${MAX_REQ_JITTER}" \
|
||||
--timeout 30 \
|
||||
--access-logfile - --error-logfile - --log-level info \
|
||||
--pythonpath /haproxy \
|
||||
'haproxy_manager:default_app' &
|
||||
|
||||
# Main API server on :8000 in the foreground. exec so signals propagate
|
||||
# correctly and the container exits if the API dies (docker --restart picks it
|
||||
# up). Longer --timeout because cert issuance hits ACME and can take a while.
|
||||
exec gunicorn \
|
||||
--bind 0.0.0.0:8000 \
|
||||
--workers "${API_WORKERS}" --threads 4 --worker-class gthread \
|
||||
--max-requests "${MAX_REQ}" --max-requests-jitter "${MAX_REQ_JITTER}" \
|
||||
--timeout "${API_TIMEOUT}" \
|
||||
--access-logfile - --error-logfile - --log-level info \
|
||||
--pythonpath /haproxy \
|
||||
'haproxy_manager:app'
|
||||
|
||||
Executable
+467
@@ -0,0 +1,467 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Regression tests for HAProxy config backup / rollback ordering.
|
||||
|
||||
Why this file exists
|
||||
--------------------
|
||||
generate_config() used to write the new haproxy.cfg and only THEN call
|
||||
reload_haproxy_safely() -> create_backup(), so the "backup" was a copy of the
|
||||
config that had just been written. On a validation failure restore_backup()
|
||||
restored the identical broken bytes: the advertised rollback was a no-op and a
|
||||
fatal haproxy.cfg stayed on disk, where start_haproxy() refuses to launch.
|
||||
|
||||
These tests pin the ordering invariant (backup predates the write) and the
|
||||
observable end-to-end behaviour (after a failed validation the file on disk is
|
||||
the previous working config and HAProxy will start with it).
|
||||
|
||||
Running
|
||||
-------
|
||||
python3 scripts/test-config-rollback.py # tests the repo checkout
|
||||
HAPROXY_MANAGER_DIR=/some/other/tree \
|
||||
python3 scripts/test-config-rollback.py # tests another tree
|
||||
|
||||
The repo has no Python test framework (scripts/test-*.sh are curl-based
|
||||
integration scripts against a running API), so this is a self-contained
|
||||
stdlib-unittest script - no pytest, no venv, no new dependencies beyond the
|
||||
application's own requirements.txt (Flask/Jinja2/psutil), which are already
|
||||
present in the container image.
|
||||
|
||||
No HAProxy binary is required: a stub `haproxy` is put on PATH that mimics
|
||||
`haproxy -c -f <file>` by rejecting any config containing the token
|
||||
__BROKEN__, which is how the tests inject an invalid configuration.
|
||||
"""
|
||||
|
||||
import os
|
||||
import sys
|
||||
import shutil
|
||||
import sqlite3
|
||||
import logging
|
||||
import tempfile
|
||||
import textwrap
|
||||
import unittest
|
||||
|
||||
BROKEN_TOKEN = '__BROKEN__'
|
||||
|
||||
MODULE_DIR = os.path.abspath(
|
||||
os.environ.get('HAPROXY_MANAGER_DIR',
|
||||
os.path.join(os.path.dirname(os.path.abspath(__file__)), '..'))
|
||||
)
|
||||
|
||||
# haproxy_manager builds its Jinja2 environment from the relative path
|
||||
# Path('templates'), so it has to be imported with the module dir as cwd.
|
||||
os.chdir(MODULE_DIR)
|
||||
sys.path.insert(0, MODULE_DIR)
|
||||
|
||||
# The module opens /var/log/haproxy-manager.log at import time via
|
||||
# logging.FileHandler. Redirect that one call so the suite runs unprivileged.
|
||||
_LOG_DIR = tempfile.mkdtemp(prefix='haproxy-mgr-test-logs-')
|
||||
_real_file_handler = logging.FileHandler
|
||||
logging.FileHandler = (
|
||||
lambda fn, *a, **kw: _real_file_handler(
|
||||
os.path.join(_LOG_DIR, os.path.basename(fn)), *a, **kw)
|
||||
)
|
||||
try:
|
||||
import haproxy_manager as hm
|
||||
except ImportError as exc: # pragma: no cover - environment problem, not a failure
|
||||
sys.stderr.write(
|
||||
f"SKIP: cannot import haproxy_manager ({exc}).\n"
|
||||
"Install the application requirements first: pip install -r requirements.txt\n"
|
||||
)
|
||||
raise SystemExit(77)
|
||||
finally:
|
||||
logging.FileHandler = _real_file_handler
|
||||
|
||||
logging.getLogger('haproxy_manager').setLevel(logging.CRITICAL)
|
||||
|
||||
FAKE_HAPROXY = textwrap.dedent(f"""\
|
||||
#!/bin/sh
|
||||
# Test stub for the haproxy binary.
|
||||
# haproxy -c -f FILE -> exit 1 if FILE contains {BROKEN_TOKEN}, else 0
|
||||
# haproxy -W -S ... -f FILE (start) -> same validation, then exit 0
|
||||
cfg=""
|
||||
while [ $# -gt 0 ]; do
|
||||
case "$1" in -f) cfg="$2"; shift ;; esac
|
||||
shift
|
||||
done
|
||||
if [ -n "$cfg" ] && grep -q '{BROKEN_TOKEN}' "$cfg" 2>/dev/null; then
|
||||
echo "[ALERT] parsing [$cfg:1] : unknown keyword '{BROKEN_TOKEN}'" >&2
|
||||
exit 1
|
||||
fi
|
||||
exit 0
|
||||
""")
|
||||
|
||||
|
||||
class RollbackTestCase(unittest.TestCase):
|
||||
"""Base fixture: an isolated fake /etc/haproxy plus a stub haproxy binary."""
|
||||
|
||||
def setUp(self):
|
||||
self.tmp = tempfile.mkdtemp(prefix='haproxy-rollback-test-')
|
||||
self.addCleanup(shutil.rmtree, self.tmp, True)
|
||||
|
||||
bindir = os.path.join(self.tmp, 'bin')
|
||||
os.makedirs(bindir)
|
||||
stub = os.path.join(bindir, 'haproxy')
|
||||
with open(stub, 'w') as fh:
|
||||
fh.write(FAKE_HAPROXY)
|
||||
os.chmod(stub, 0o755)
|
||||
self._old_path = os.environ['PATH']
|
||||
os.environ['PATH'] = bindir + os.pathsep + self._old_path
|
||||
self.addCleanup(lambda: os.environ.__setitem__('PATH', self._old_path))
|
||||
|
||||
self.etc = os.path.join(self.tmp, 'etc')
|
||||
os.makedirs(self.etc)
|
||||
|
||||
overrides = {
|
||||
'DB_FILE': os.path.join(self.etc, 'haproxy_config.db'),
|
||||
'HAPROXY_CONFIG_PATH': os.path.join(self.etc, 'haproxy.cfg'),
|
||||
'HAPROXY_BACKUP_PATH': os.path.join(self.etc, 'haproxy.cfg.backup'),
|
||||
'BLOCKED_IPS_MAP_PATH': os.path.join(self.etc, 'blocked_ips.map'),
|
||||
'BLOCKED_IPS_MAP_BACKUP_PATH': os.path.join(self.etc, 'blocked_ips.map.backup'),
|
||||
'CLUSTER_SECRET_PATH': os.path.join(self.etc, 'cluster-secret'),
|
||||
'SSL_CERTS_DIR': os.path.join(self.etc, 'certs'),
|
||||
'HAPROXY_SOCKET_PATH': os.path.join(self.etc, 'haproxy.sock'),
|
||||
# Added by the rollback fix; older trees do not have it.
|
||||
'CORAZA_SPOE_CONFIG_PATH': os.path.join(self.etc, 'coraza-spoe.cfg'),
|
||||
'CORAZA_SPOE_BACKUP_PATH': os.path.join(self.etc, 'coraza-spoe.cfg.backup'),
|
||||
}
|
||||
self._saved = {}
|
||||
for name, value in overrides.items():
|
||||
self._saved[name] = getattr(hm, name, None)
|
||||
setattr(hm, name, value)
|
||||
self.addCleanup(self._restore_globals)
|
||||
os.makedirs(hm.SSL_CERTS_DIR)
|
||||
|
||||
# log_operation() appends to a hardcoded /var/log path. Injecting `open`
|
||||
# into the module namespace shadows the builtin for that module only
|
||||
# (module globals are searched before builtins), so the real
|
||||
# log_operation code still runs.
|
||||
real_open = open
|
||||
log_dir = self.tmp
|
||||
|
||||
def _redirecting_open(path, *args, **kwargs):
|
||||
if isinstance(path, str) and path.startswith('/var/log/'):
|
||||
path = os.path.join(log_dir, os.path.basename(path))
|
||||
return real_open(path, *args, **kwargs)
|
||||
|
||||
hm.open = _redirecting_open
|
||||
self.addCleanup(lambda: hm.__dict__.pop('open', None))
|
||||
|
||||
hm.init_db()
|
||||
|
||||
def _restore_globals(self):
|
||||
for name, value in self._saved.items():
|
||||
if value is None:
|
||||
hm.__dict__.pop(name, None)
|
||||
else:
|
||||
setattr(hm, name, value)
|
||||
|
||||
# -- helpers ---------------------------------------------------------
|
||||
def add_domain(self, domain, backend_name, address='10.0.0.1'):
|
||||
with sqlite3.connect(hm.DB_FILE) as conn:
|
||||
cur = conn.cursor()
|
||||
cur.execute('INSERT INTO domains (domain, ssl_enabled) VALUES (?, 0)',
|
||||
(domain,))
|
||||
domain_id = cur.lastrowid
|
||||
cur.execute('INSERT INTO backends (name, domain_id) VALUES (?, ?)',
|
||||
(backend_name, domain_id))
|
||||
backend_id = cur.lastrowid
|
||||
cur.execute(
|
||||
'INSERT INTO backend_servers '
|
||||
'(backend_id, server_name, server_address, server_port) '
|
||||
'VALUES (?, ?, ?, ?)',
|
||||
(backend_id, 'srv1', address, 8080))
|
||||
conn.commit()
|
||||
|
||||
def block_ip(self, ip):
|
||||
with sqlite3.connect(hm.DB_FILE) as conn:
|
||||
conn.execute('INSERT INTO blocked_ips (ip_address, reason) VALUES (?, ?)',
|
||||
(ip, 'test'))
|
||||
conn.commit()
|
||||
|
||||
def read(self, path):
|
||||
with open(path) as fh:
|
||||
return fh.read()
|
||||
|
||||
def config_is_loadable(self):
|
||||
"""True if HAProxy would accept the config currently on disk."""
|
||||
import subprocess
|
||||
return subprocess.run(
|
||||
['haproxy', '-c', '-f', hm.HAPROXY_CONFIG_PATH],
|
||||
capture_output=True).returncode == 0
|
||||
|
||||
def generate_good_config(self):
|
||||
self.add_domain('good.example.com', 'good_backend')
|
||||
hm.generate_config()
|
||||
self.assertTrue(self.config_is_loadable(),
|
||||
'fixture precondition: first generated config must be valid')
|
||||
return self.read(hm.HAPROXY_CONFIG_PATH)
|
||||
|
||||
def break_the_config(self):
|
||||
"""Queue a domain whose rendered backend the validator rejects."""
|
||||
self.add_domain('bad.example.com', BROKEN_TOKEN + '_backend', '10.0.0.2')
|
||||
|
||||
|
||||
class TestBackupOrdering(RollbackTestCase):
|
||||
|
||||
def test_backup_is_taken_before_the_new_config_is_written(self):
|
||||
"""The ordering invariant, asserted directly.
|
||||
|
||||
Whatever create_backup() sees on disk must be the OLD config; if the
|
||||
write happens first the backup is a copy of the new config and rollback
|
||||
is meaningless.
|
||||
"""
|
||||
good = self.generate_good_config()
|
||||
|
||||
seen = {}
|
||||
real_create_backup = hm.create_backup
|
||||
|
||||
def spy(*args, **kwargs):
|
||||
seen['config_on_disk'] = self.read(hm.HAPROXY_CONFIG_PATH)
|
||||
return real_create_backup(*args, **kwargs)
|
||||
|
||||
hm.create_backup = spy
|
||||
self.addCleanup(setattr, hm, 'create_backup', real_create_backup)
|
||||
|
||||
self.add_domain('second.example.com', 'second_backend', '10.0.0.3')
|
||||
hm.generate_config()
|
||||
|
||||
self.assertIn('config_on_disk', seen,
|
||||
'create_backup() was never called during generate_config()')
|
||||
self.assertEqual(
|
||||
seen['config_on_disk'], good,
|
||||
'create_backup() ran AFTER the new config was written - the backup '
|
||||
'is a copy of the new config, so rollback cannot undo anything')
|
||||
|
||||
def test_backup_tracks_the_last_known_good_config(self):
|
||||
"""After a change that validated AND loaded, the backup is that config.
|
||||
|
||||
The rollback target is "the last configuration HAProxy actually ran",
|
||||
not "the file that happened to be there last time".
|
||||
"""
|
||||
good = self.generate_good_config()
|
||||
self.add_domain('second.example.com', 'second_backend', '10.0.0.3')
|
||||
hm.generate_config()
|
||||
|
||||
live = self.read(hm.HAPROXY_CONFIG_PATH)
|
||||
self.assertNotEqual(live, good, 'fixture sanity: the new config should differ')
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), live,
|
||||
'the successful config was not recorded as known-good')
|
||||
|
||||
def test_backup_is_not_promoted_when_the_change_fails(self):
|
||||
"""A config that never loaded must not become the rollback target."""
|
||||
good = self.generate_good_config()
|
||||
self.break_the_config()
|
||||
with self.assertRaises(Exception):
|
||||
hm.generate_config()
|
||||
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
|
||||
'a config that failed validation was promoted to backup')
|
||||
|
||||
|
||||
class TestRollbackEndToEnd(RollbackTestCase):
|
||||
|
||||
def test_failed_validation_leaves_the_last_good_config_on_disk(self):
|
||||
good = self.generate_good_config()
|
||||
|
||||
self.break_the_config()
|
||||
with self.assertRaises(Exception):
|
||||
hm.generate_config()
|
||||
|
||||
on_disk = self.read(hm.HAPROXY_CONFIG_PATH)
|
||||
self.assertNotIn(BROKEN_TOKEN, on_disk,
|
||||
'the rejected config is still on disk - rollback was a no-op')
|
||||
self.assertEqual(on_disk, good,
|
||||
'on-disk config is not byte-identical to the last good one')
|
||||
|
||||
def test_haproxy_would_still_start_after_a_failed_change(self):
|
||||
"""The operational consequence: the edge can still come up."""
|
||||
self.generate_good_config()
|
||||
self.break_the_config()
|
||||
with self.assertRaises(Exception):
|
||||
hm.generate_config()
|
||||
|
||||
self.assertTrue(self.config_is_loadable(),
|
||||
'HAProxy would refuse to start with the config left on disk')
|
||||
with self.assertLogs('haproxy_manager', level='INFO') as captured:
|
||||
hm.start_haproxy()
|
||||
self.assertTrue(
|
||||
any('HAProxy started successfully' in line for line in captured.output),
|
||||
f'start_haproxy() did not succeed after rollback: {captured.output}')
|
||||
|
||||
def test_blocked_ips_map_is_rolled_back_too(self):
|
||||
"""generate_config() rewrites the map file before writing haproxy.cfg."""
|
||||
self.block_ip('192.0.2.10')
|
||||
self.generate_good_config()
|
||||
good_map = self.read(hm.BLOCKED_IPS_MAP_PATH)
|
||||
|
||||
self.block_ip('198.51.100.20')
|
||||
self.break_the_config()
|
||||
with self.assertRaises(Exception):
|
||||
hm.generate_config()
|
||||
|
||||
self.assertEqual(self.read(hm.BLOCKED_IPS_MAP_PATH), good_map,
|
||||
'blocked IPs map was not rolled back with the config')
|
||||
|
||||
def test_first_run_failure_reports_that_rollback_was_impossible(self):
|
||||
"""No prior config: there is nothing to restore, and that must be said.
|
||||
|
||||
A missing backup must never be reported as a successful restore, and it
|
||||
must never be turned into "restore an empty file".
|
||||
"""
|
||||
self.break_the_config()
|
||||
with self.assertRaises(Exception) as ctx:
|
||||
hm.generate_config()
|
||||
|
||||
self.assertIn('ROLLBACK FAILED', str(ctx.exception),
|
||||
'a failed change with no backup was not reported as such')
|
||||
self.assertFalse(os.path.exists(hm.HAPROXY_BACKUP_PATH),
|
||||
'a backup was fabricated from the broken config')
|
||||
# The broken config is deliberately left in place: start_haproxy() can
|
||||
# then detect it and try to regenerate. It must not be blanked.
|
||||
self.assertGreater(os.path.getsize(hm.HAPROXY_CONFIG_PATH), 0,
|
||||
'config file was emptied instead of left for diagnosis')
|
||||
|
||||
|
||||
class TestBackupPrimitives(RollbackTestCase):
|
||||
|
||||
def test_restore_backup_distinguishes_missing_backup_from_success(self):
|
||||
restored, message = hm.restore_backup()
|
||||
self.assertFalse(restored,
|
||||
'restore_backup() reported success with no backup present')
|
||||
self.assertIn('cannot roll back', message.lower())
|
||||
|
||||
good = self.generate_good_config()
|
||||
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
|
||||
fh.write('scribbled over\n')
|
||||
|
||||
restored, message = hm.restore_backup()
|
||||
self.assertTrue(restored, message)
|
||||
self.assertEqual(self.read(hm.HAPROXY_CONFIG_PATH), good)
|
||||
|
||||
def test_a_successful_generation_records_a_rollback_target(self):
|
||||
"""Even the first-ever generation must leave something to roll back to."""
|
||||
good = self.generate_good_config()
|
||||
self.assertTrue(
|
||||
os.path.exists(hm.HAPROXY_BACKUP_PATH),
|
||||
'after a successful reload there is still no known-good backup')
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good)
|
||||
|
||||
def test_a_broken_current_config_does_not_replace_a_good_backup(self):
|
||||
"""The known-good marker.
|
||||
|
||||
If the config already on disk is broken (previous failed write, manual
|
||||
edit), snapshotting it would make "rollback" mean "restore a different
|
||||
broken config". The older validated backup must survive.
|
||||
"""
|
||||
good = self.generate_good_config()
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
|
||||
'fixture: a good backup should exist by now')
|
||||
|
||||
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
|
||||
fh.write(f'garbage {BROKEN_TOKEN} config\n')
|
||||
|
||||
ok, status = hm.create_backup()
|
||||
self.assertTrue(ok)
|
||||
self.assertEqual(status, 'kept_previous')
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
|
||||
'a broken config overwrote the known-good backup')
|
||||
|
||||
def test_reload_does_not_take_its_own_backup(self):
|
||||
"""reload_haproxy_safely() runs after the write, so it must not back up."""
|
||||
good = self.generate_good_config()
|
||||
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
|
||||
fh.write(f'broken {BROKEN_TOKEN}\n')
|
||||
|
||||
success, message = hm.reload_haproxy_safely(backup_status='created')
|
||||
|
||||
self.assertFalse(success)
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good,
|
||||
'reload_haproxy_safely() overwrote the good backup')
|
||||
self.assertEqual(self.read(hm.HAPROXY_CONFIG_PATH), good,
|
||||
'reload_haproxy_safely() did not roll the config back')
|
||||
|
||||
def test_unchanged_config_is_not_revalidated(self):
|
||||
"""Fast path: if the backup already is the live config, do no work.
|
||||
|
||||
generate_config() runs inside customer-facing API calls and
|
||||
`haproxy -c` is expensive on an edge with hundreds of certificates.
|
||||
"""
|
||||
self.generate_good_config()
|
||||
|
||||
calls = []
|
||||
real_validate = hm.validate_config_file
|
||||
hm.validate_config_file = lambda path: (calls.append(path),
|
||||
real_validate(path))[1]
|
||||
self.addCleanup(setattr, hm, 'validate_config_file', real_validate)
|
||||
|
||||
ok, status = hm.create_backup()
|
||||
self.assertTrue(ok)
|
||||
self.assertEqual(status, 'created')
|
||||
self.assertEqual(calls, [],
|
||||
'the unchanged live config was re-validated needlessly')
|
||||
|
||||
def test_fast_path_does_not_hide_a_drifted_broken_config(self):
|
||||
"""If the live config drifted from the backup, the gate must still run."""
|
||||
good = self.generate_good_config()
|
||||
with open(hm.HAPROXY_CONFIG_PATH, 'w') as fh:
|
||||
fh.write(f'hand edited {BROKEN_TOKEN}\n')
|
||||
|
||||
ok, status = hm.create_backup()
|
||||
self.assertTrue(ok)
|
||||
self.assertEqual(status, 'kept_previous',
|
||||
'a drifted broken config was silently accepted')
|
||||
self.assertEqual(self.read(hm.HAPROXY_BACKUP_PATH), good)
|
||||
|
||||
def test_backup_set_covers_every_file_generate_config_writes(self):
|
||||
pairs = dict(hm._config_backup_pairs())
|
||||
for path in (hm.HAPROXY_CONFIG_PATH, hm.BLOCKED_IPS_MAP_PATH,
|
||||
hm.CORAZA_SPOE_CONFIG_PATH):
|
||||
self.assertIn(path, pairs,
|
||||
f'{path} is written by generate_config() but is not '
|
||||
'part of the backed-up config set')
|
||||
|
||||
def test_coraza_spoe_config_round_trips(self):
|
||||
self.generate_good_config()
|
||||
with open(hm.CORAZA_SPOE_CONFIG_PATH, 'w') as fh:
|
||||
fh.write('spoe-good\n')
|
||||
hm.create_backup()
|
||||
with open(hm.CORAZA_SPOE_CONFIG_PATH, 'w') as fh:
|
||||
fh.write('spoe-broken\n')
|
||||
restored, message = hm.restore_backup()
|
||||
self.assertTrue(restored, message)
|
||||
self.assertEqual(self.read(hm.CORAZA_SPOE_CONFIG_PATH), 'spoe-good\n')
|
||||
|
||||
|
||||
class TestAtomicWrite(RollbackTestCase):
|
||||
|
||||
def test_write_is_atomic_and_preserves_mode(self):
|
||||
path = os.path.join(self.etc, 'atomic.cfg')
|
||||
with open(path, 'w') as fh:
|
||||
fh.write('old')
|
||||
os.chmod(path, 0o644)
|
||||
|
||||
hm.write_config_atomically(path, 'new content\n')
|
||||
|
||||
self.assertEqual(self.read(path), 'new content\n')
|
||||
self.assertEqual(oct(os.stat(path).st_mode & 0o777), oct(0o644))
|
||||
leftovers = [n for n in os.listdir(self.etc) if n.endswith('.tmp')]
|
||||
self.assertEqual(leftovers, [], f'temp files left behind: {leftovers}')
|
||||
|
||||
def test_failed_write_leaves_the_previous_file_intact(self):
|
||||
path = os.path.join(self.etc, 'atomic.cfg')
|
||||
with open(path, 'w') as fh:
|
||||
fh.write('old content\n')
|
||||
|
||||
# Anything that makes f.write() blow up mid-flight stands in for a full
|
||||
# disk / killed container.
|
||||
with self.assertRaises(Exception):
|
||||
hm.write_config_atomically(path, object())
|
||||
|
||||
self.assertEqual(self.read(path), 'old content\n',
|
||||
'a failed write clobbered the previous config')
|
||||
leftovers = [n for n in os.listdir(self.etc) if n.endswith('.tmp')]
|
||||
self.assertEqual(leftovers, [], f'temp files left behind: {leftovers}')
|
||||
|
||||
|
||||
if __name__ == '__main__':
|
||||
print(f"testing haproxy_manager from: {MODULE_DIR}")
|
||||
unittest.main(verbosity=2)
|
||||
@@ -11,7 +11,7 @@ backend {{ name }}-backend
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
|
||||
# SSE-specific backend - optimized for Server-Sent Events long-lived connections
|
||||
@@ -36,5 +36,5 @@ backend {{ name }}-sse-backend
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
|
||||
@@ -0,0 +1,41 @@
|
||||
# Long-lived backend for {{ name }} (template_override='hap_backend_longlived').
|
||||
# Use for apps whose PRIMARY traffic holds connections open: media streaming,
|
||||
# large up/downloads, or persistent viewer/streaming sessions. Both the primary
|
||||
# and the SSE backend are tuned long-lived here (no http-server-close,
|
||||
# http-no-delay, 6h server/tunnel/keep-alive timeouts).
|
||||
#
|
||||
# Compare hap_backend_websocket.tpl, which keeps the PRIMARY backend standard
|
||||
# and only makes the -sse-backend long-lived. Pick this one when the main path
|
||||
# itself needs long-lived connections, not just an SSE side-channel.
|
||||
backend {{ name }}-backend
|
||||
no option http-server-close
|
||||
option http-no-delay
|
||||
timeout server 6h
|
||||
timeout tunnel 6h
|
||||
timeout http-keep-alive 6h
|
||||
option forwardfor
|
||||
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Real-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-Proto https if { ssl_fc }
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
|
||||
# SSE variant (Accept: text/event-stream / ?action=stream auto-routes here)
|
||||
backend {{ name }}-sse-backend
|
||||
no option http-server-close
|
||||
option http-no-delay
|
||||
timeout server 6h
|
||||
timeout tunnel 6h
|
||||
timeout http-keep-alive 6h
|
||||
option forwardfor
|
||||
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Real-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-Proto https if { ssl_fc }
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
@@ -0,0 +1,34 @@
|
||||
# Long-lived / websocket-safe backend for {{ name }} (template_override)
|
||||
# For apps with persistent WebSocket/streaming connections (e.g. Jitsi /xmpp-websocket, /colibri-ws).
|
||||
backend {{ name }}-backend
|
||||
no option http-server-close
|
||||
option http-no-delay
|
||||
timeout server 6h
|
||||
timeout tunnel 6h
|
||||
timeout http-keep-alive 6h
|
||||
option forwardfor
|
||||
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Real-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-Proto https if { ssl_fc }
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
|
||||
# SSE variant (Accept: text/event-stream / ?action=stream auto-routes here)
|
||||
backend {{ name }}-sse-backend
|
||||
no option http-server-close
|
||||
option http-no-delay
|
||||
timeout server 6h
|
||||
timeout tunnel 6h
|
||||
timeout http-keep-alive 6h
|
||||
option forwardfor
|
||||
http-request add-header X-CLIENT-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Real-IP %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-For %[var(txn.real_ip)]
|
||||
http-request set-header X-Forwarded-Proto https if { ssl_fc }
|
||||
http-request set-header X-Forwarded-Proto http if !{ ssl_fc }
|
||||
{% for server in servers %}
|
||||
server {{ server.server_name }} {{ server.server_address }}:{{ server.server_port }} {{ server.server_options }} resolvers docker_dns init-addr last,libc,none
|
||||
{% endfor %}
|
||||
@@ -0,0 +1,16 @@
|
||||
# Coraza-SPOA backend.
|
||||
# Only rendered into haproxy.cfg when HAPROXY_CORAZA_SPOE_BACKEND env var is
|
||||
# set on the haproxy-manager container. SPOE traffic to this backend is TCP,
|
||||
# not HTTP. The agent target comes from the env var so a single image can be
|
||||
# deployed against different sidecar host:port pairs (typically the sidecar
|
||||
# container's name + 9000 inside the shared docker network).
|
||||
backend coraza-spoa-backend
|
||||
mode tcp
|
||||
# spop-check actually speaks the SPOE protocol against the agent —
|
||||
# confirms the agent can negotiate a session, not just that the TCP
|
||||
# port is open. Required to detect a half-broken SPOA that's listening
|
||||
# but not actually processing.
|
||||
option spop-check
|
||||
timeout connect 5s
|
||||
timeout server 30s
|
||||
server coraza-spoa {{ agent_target }} check
|
||||
@@ -0,0 +1,60 @@
|
||||
# Coraza SPOE engine configuration.
|
||||
#
|
||||
# Written to /etc/haproxy/coraza-spoe.cfg by haproxy_manager.generate_config()
|
||||
# when HAPROXY_CORAZA_SPOE_BACKEND env var is set. Referenced from haproxy.cfg
|
||||
# via `filter spoe engine coraza config /etc/haproxy/coraza-spoe.cfg`.
|
||||
#
|
||||
# Engine name "coraza" must match the engine name in the filter line in the
|
||||
# main config; group name "coraza-req" must match the send-spoe-group action.
|
||||
# Application name "haproxy" must match the application block in coraza-spoa's
|
||||
# config.yaml.
|
||||
#
|
||||
# Reference: this config follows the shape from coraza-spoa's upstream
|
||||
# example/haproxy/coraza.cfg (v0.7.1). Arg names + ordering are required by
|
||||
# Coraza-SPOA exactly as specified — DO NOT reorder or rename without
|
||||
# coordinating with the agent.
|
||||
|
||||
[coraza]
|
||||
|
||||
spoe-agent coraza
|
||||
# `groups` (not `messages`) lists the spoe-group names this engine offers
|
||||
# via `send-spoe-group` actions. The same group name appears below in a
|
||||
# spoe-group block, which in turn references the actual message.
|
||||
groups coraza-req
|
||||
|
||||
# Prefix for variables the agent sets back on the request transaction —
|
||||
# e.g. var(txn.coraza.error) when set-on-error triggers.
|
||||
option var-prefix coraza
|
||||
|
||||
# On agent error/timeout, set var(txn.coraza.error). We DON'T add a
|
||||
# corresponding `http-request deny if { var(txn.coraza.error) -m bool }`
|
||||
# in the frontend, so the request continues uninspected. This is the
|
||||
# fail-open posture: WAF outage shouldn't 503 customer traffic.
|
||||
option set-on-error error
|
||||
|
||||
timeout hello 2s
|
||||
timeout idle 2m
|
||||
timeout processing 100ms
|
||||
|
||||
use-backend coraza-spoa-backend
|
||||
log global
|
||||
|
||||
# Per-request inspection message. No `event` directive — fires only when
|
||||
# explicitly invoked from haproxy.cfg via `http-request send-spoe-group`.
|
||||
# Arg order/names are mandatory: Coraza-SPOA parses positionally and renames
|
||||
# break the agent. `app=str(haproxy)` is the literal application name from
|
||||
# coraza-spoa's config.yaml `applications:` block.
|
||||
#
|
||||
# src-ip uses var(txn.real_ip) — HAProxy resolves the real client IP at the
|
||||
# top of the frontend (CF-Connecting-IP > X-Real-IP > X-Forwarded-For > src,
|
||||
# only honoring those headers from trusted proxies). Falls back to `src`
|
||||
# when no proxy headers are present. This is what shows up as `client_ip`
|
||||
# in /var/log/coraza/audit.log and the rule manager's view-matches panel.
|
||||
spoe-message coraza-req
|
||||
args app=str(haproxy) src-ip=var(txn.real_ip) src-port=src_port dst-ip=dst dst-port=dst_port method=method path=path query=query version=req.ver headers=req.hdrs body=req.body
|
||||
|
||||
# Group binding for send-spoe-group invocation in the frontend. One group,
|
||||
# one message; could add more in the future (e.g. coraza-res for response
|
||||
# inspection — currently disabled in coraza-spoa's config.yaml).
|
||||
spoe-group coraza-req
|
||||
messages coraza-req
|
||||
@@ -27,12 +27,47 @@ global
|
||||
# SSL and Performance
|
||||
tune.ssl.default-dh-param 2048
|
||||
|
||||
# HTTP/3 over QUIC. The Debian haproxy package is built against system
|
||||
# OpenSSL via the compatibility shim (USE_QUIC_OPENSSL_COMPAT), which is
|
||||
# not a native QUIC TLS stack. HAProxy therefore rejects `quic*@` binds
|
||||
# unless this opt-in is set. `limited-quic` enables QUIC through the compat
|
||||
# layer (no 0-RTT — that needs quictls/aws-lc or native OpenSSL 3.5 QUIC).
|
||||
# Without this, the quic bind in the frontend fails to start: "this SSL
|
||||
# library does not support the QUIC protocol".
|
||||
limited-quic
|
||||
{%- if cluster_secret %}
|
||||
|
||||
# Stable secret keying QUIC Retry/address-validation tokens. Self-healed
|
||||
# to /etc/haproxy/cluster-secret (named volume) by the manager so it
|
||||
# survives recreates; without it haproxy picks a random one per process
|
||||
# and tokens don't survive reloads (benign, just a startup notice).
|
||||
cluster-secret "{{ cluster_secret }}"
|
||||
{%- endif %}
|
||||
|
||||
# HTTP/2 protection against Rapid Reset (CVE-2023-44487) and stream abuse
|
||||
tune.h2.fe.max-total-streams 2000
|
||||
tune.h2.fe.glitches-threshold 50
|
||||
|
||||
# Stats persistence for zero-downtime reloads
|
||||
stats-file /var/lib/haproxy/stats.dat
|
||||
|
||||
#---------------------------------------------------------------------
|
||||
# DNS resolver for Docker container name resolution
|
||||
# Re-resolves backend server addresses so container IP changes
|
||||
# (from restarts, recreations, scaling) are picked up automatically
|
||||
#---------------------------------------------------------------------
|
||||
resolvers docker_dns
|
||||
nameserver dns1 127.0.0.11:53
|
||||
resolve_retries 3
|
||||
timeout resolve 1s
|
||||
timeout retry 1s
|
||||
hold valid 10s
|
||||
hold other 10s
|
||||
hold refused 10s
|
||||
hold nx 10s
|
||||
hold timeout 10s
|
||||
hold obsolete 10s
|
||||
|
||||
#---------------------------------------------------------------------
|
||||
# common defaults that all the 'listen' and 'backend' sections will
|
||||
# use if not designated in their block
|
||||
@@ -56,3 +91,14 @@ defaults
|
||||
timeout tarpit 10s # Tarpit delay for low-level scanners (before silent-drop)
|
||||
maxconn 3000
|
||||
|
||||
# Per-request unique reference, used:
|
||||
# - in the log line (httplog includes %ID)
|
||||
# - echoed to clients in the X-Request-Reference response header on
|
||||
# WAF blocks so a customer can quote it when opening a support ticket
|
||||
# - embedded in /etc/haproxy/errors/403-waf.html so a blocked visitor
|
||||
# sees it on the rendered 403 page
|
||||
# Support correlates ref → /var/log/haproxy.log line → timestamp+client+host
|
||||
# → /var/log/coraza/audit.log entry → rule_id.
|
||||
unique-id-format %[uuid()]
|
||||
unique-id-header X-Request-Reference
|
||||
|
||||
+150
-15
@@ -4,6 +4,21 @@ frontend web
|
||||
# crt can now be a path, so it will load all .pem files in the path
|
||||
bind 0.0.0.0:443 ssl crt {{ crt_path }} alpn h2,http/1.1
|
||||
|
||||
# HTTP/3 over QUIC (UDP/443). Same cert path as the TCP listener above.
|
||||
# The Debian haproxy package is built +QUIC (QUIC_OPENSSL_COMPAT), so this
|
||||
# is config-only — no source build. Requires UDP/443 published on the
|
||||
# container (`-p 443:443/udp`) and open at the host firewall. `h3` is the
|
||||
# only ALPN QUIC negotiates; h2/http1 stay on the TCP bind above. Sharing
|
||||
# the frontend means all the real-IP, rate-limit, IP-block and Coraza
|
||||
# rules below apply identically to H3 traffic.
|
||||
bind quic4@0.0.0.0:443 ssl crt {{ crt_path }} alpn h3
|
||||
|
||||
# Advertise H3 so browsers upgrade their existing TCP (h2) connection to
|
||||
# QUIC on the next request. `ma` is how long (seconds) the client may
|
||||
# cache the advertisement. http-after-response applies it to every
|
||||
# response, including haproxy-generated ones (blocks, default page).
|
||||
http-after-response set-header alt-svc "h3=\":443\"; ma=86400"
|
||||
|
||||
# Capture Host header so it appears in httplog output (in %hr field)
|
||||
http-request capture req.hdr(Host) len 64
|
||||
|
||||
@@ -13,31 +28,102 @@ frontend web
|
||||
acl has_x_real_ip req.hdr(X-Real-IP) -m found
|
||||
acl has_x_forwarded_for req.hdr(X-Forwarded-For) -m found
|
||||
|
||||
# Set the real IP based on available headers
|
||||
http-request set-var(txn.real_ip) req.hdr(CF-Connecting-IP) if has_cf_connecting_ip
|
||||
http-request set-var(txn.real_ip) req.hdr(X-Real-IP) if !has_cf_connecting_ip has_x_real_ip
|
||||
http-request set-var(txn.real_ip) req.hdr(X-Forwarded-For) if !has_cf_connecting_ip !has_x_real_ip has_x_forwarded_for
|
||||
# Set the real IP based on available headers. Use hdr_ip (not hdr) so the
|
||||
# variable is typed as IP — required by the Coraza SPOE arg `src-ip` which
|
||||
# decodes binary IP bytes (passing a string IP panics the SPOA goroutine).
|
||||
# `hdr_ip(X-Forwarded-For,1)` extracts the FIRST address from a possibly
|
||||
# comma-separated chain (original client, not intermediate proxies).
|
||||
http-request set-var(txn.real_ip) req.hdr_ip(CF-Connecting-IP) if has_cf_connecting_ip
|
||||
http-request set-var(txn.real_ip) req.hdr_ip(X-Real-IP) if !has_cf_connecting_ip has_x_real_ip
|
||||
http-request set-var(txn.real_ip) req.hdr_ip(X-Forwarded-For,1) if !has_cf_connecting_ip !has_x_real_ip has_x_forwarded_for
|
||||
http-request set-var(txn.real_ip) src if !has_cf_connecting_ip !has_x_real_ip !has_x_forwarded_for
|
||||
|
||||
# --- Connection & rate tracking ---
|
||||
stick-table type ip size 200k expire 10m store conn_cur,conn_rate(10s),http_req_rate(10s),http_err_rate(30s)
|
||||
http-request track-sc0 var(txn.real_ip)
|
||||
|
||||
# Whitelist: let health checks and local traffic bypass rate limits
|
||||
# Whitelist: let health checks, local, and trusted traffic bypass rate limits
|
||||
acl is_local src 127.0.0.0/8 10.0.0.0/8 172.16.0.0/12 192.168.0.0/16
|
||||
acl is_trusted_ip src -f /etc/haproxy/trusted_ips.list
|
||||
acl is_health_check path_beg /.well-known/acme-challenge
|
||||
acl is_whitelisted var(txn.real_ip),map_ip(/etc/haproxy/trusted_ips.map,0) -m int gt 0
|
||||
|
||||
# --- Rate limit rules (applied in order, first match wins) ---
|
||||
# Hard block: >500 req/10s per IP (sustained flood)
|
||||
http-request deny deny_status 429 if { sc_http_req_rate(0) gt 500 } !is_local !is_health_check
|
||||
# Tarpit: >200 req/10s per IP (aggressive scraping / light flood)
|
||||
http-request tarpit deny_status 429 if { sc_http_req_rate(0) gt 200 } !is_local !is_health_check
|
||||
# Connection rate limit: >150 new connections per 10s per IP
|
||||
http-request deny deny_status 429 if { sc_conn_rate(0) gt 150 } !is_local !is_health_check
|
||||
# Concurrent connection limit: >100 simultaneous connections per IP
|
||||
http-request deny deny_status 429 if { sc_conn_cur(0) gt 100 } !is_local !is_health_check
|
||||
# High error rate: >20 errors in 30s (scanner/fuzzer behavior)
|
||||
http-request tarpit deny_status 403 if { sc_http_err_rate(0) gt 20 } !is_local !is_health_check
|
||||
# Thresholds are generous to accommodate media-heavy sites where a
|
||||
# single page can load 100+ images/assets. These only trigger on
|
||||
# obvious automated abuse, not real users.
|
||||
#
|
||||
# Hard block: >5000 req/10s per IP (500 req/s — sustained flood)
|
||||
http-request deny deny_status 429 if { sc_http_req_rate(0) gt 5000 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
|
||||
# Tarpit: >3000 req/10s per IP (300 req/s — aggressive bot/scraper)
|
||||
http-request tarpit deny_status 429 if { sc_http_req_rate(0) gt 3000 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
|
||||
# Connection rate limit: >500 new connections per 10s per IP
|
||||
http-request deny deny_status 429 if { sc_conn_rate(0) gt 500 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
|
||||
# Concurrent connection limit: >500 simultaneous connections per IP
|
||||
http-request deny deny_status 429 if { sc_conn_cur(0) gt 500 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
|
||||
# High error rate: >100 errors in 30s (scanner/fuzzer behavior)
|
||||
http-request tarpit deny_status 403 if { sc_http_err_rate(0) gt 100 } !is_local !is_trusted_ip !is_whitelisted !is_health_check
|
||||
|
||||
# --- WordPress wp-login.php brute-force protection ---
|
||||
# The generic limits above are deliberately high (media-heavy sites), so a
|
||||
# slow credential-stuffing run (dozens of login POSTs/min) slips under them.
|
||||
# Track POSTs to wp-login.php per real client IP in a DEDICATED 60s table
|
||||
# (sc1 / backend wp_bruteforce, defined in hap_security_tables.tpl) and
|
||||
# tarpit once an IP exceeds 30/min. Only login POSTs are counted — GETs of
|
||||
# the login form, normal browsing, and the handful of POSTs a legit user
|
||||
# makes are unaffected; an offending IP can still browse, just not keep
|
||||
# hammering login. path_end also covers subdirectory WP installs. Honors the
|
||||
# same whitelist (RFC1918 / trusted_ips.list / trusted_ips.map).
|
||||
acl wp_login_path path_end /wp-login.php
|
||||
http-request track-sc1 var(txn.real_ip) table wp_bruteforce if METH_POST wp_login_path
|
||||
http-request tarpit deny_status 429 if METH_POST wp_login_path { sc_http_req_rate(1) gt 30 } !is_local !is_trusted_ip !is_whitelisted
|
||||
|
||||
# --- WordPress wp-login.php "must-load-the-form-first" cookie challenge ---
|
||||
# Defeats DISTRIBUTED credential-stuffing (hundreds of thousands of unique
|
||||
# IPs, each low-and-slow, so the per-IP rule above can't see them). Such
|
||||
# bots POST straight to /wp-login.php without ever GETting the form — on
|
||||
# these sites the login POST:GET ratio is ~15:1. We hand out a cookie when
|
||||
# the form is actually fetched (GET) and require it on POST; direct-POST
|
||||
# bots lack it and are denied AT THE EDGE before reaching PHP. Real logins
|
||||
# are unaffected — WordPress login already requires loading the page and
|
||||
# accepting cookies. Immediate deny (NOT tarpit) — under a 300k-POST flood,
|
||||
# holding tarpit connections would exhaust HAProxy. Honors the whitelist.
|
||||
# Mark login-form GETs at REQUEST time (method/path are reliably evaluable
|
||||
# here; in the response phase they are not) so the cookie is emitted on the
|
||||
# form's own response.
|
||||
http-request set-var(txn.wp_login_form) int(1) if METH_GET wp_login_path
|
||||
http-after-response add-header set-cookie "whplc=1; Path=/; Max-Age=1800; HttpOnly; Secure; SameSite=Lax" if { var(txn.wp_login_form) -m found }
|
||||
acl has_login_cookie req.cook(whplc) -m found
|
||||
http-request deny deny_status 403 if METH_POST wp_login_path !has_login_cookie !is_local !is_trusted_ip !is_whitelisted
|
||||
|
||||
# WordPress REST batch endpoint lockdown ("wp2shell": CVE-2026-63030 +
|
||||
# CVE-2026-60137). Chaining a core SQL injection with REST batch-route
|
||||
# confusion gives unauthenticated RCE on WP 6.9.0-6.9.4 and 7.0.0-7.0.1
|
||||
# (fixed in 6.9.5 / 7.0.2). Exploits are public and were used against this
|
||||
# fleet on 2026-07-19/20; one site was compromised via this path before
|
||||
# patching. This is a virtual patch: it does not repair the vulnerable
|
||||
# application logic, it only removes reachability, so it stays until every
|
||||
# site is confirmed on a fixed release.
|
||||
#
|
||||
# Both routing forms must be covered -- a rule matching only the pretty
|
||||
# permalink path leaves the ?rest_route= fallback wide open, and urlp()
|
||||
# does not URL-decode, hence the third ACL for the %2F spelling.
|
||||
#
|
||||
# Anonymous-only. batch/v1 is used legitimately by the block editor for
|
||||
# multi-entity saves, so a blanket deny would break wp-admin for real
|
||||
# users; requiring a wordpress_logged_in_* cookie costs them nothing.
|
||||
# req.cook() needs an exact name and WordPress suffixes a per-site hash,
|
||||
# so this substring-matches the raw Cookie header instead.
|
||||
#
|
||||
# Immediate deny, not tarpit -- holding connections open helps an attacker
|
||||
# who is already scripting this. Honors the same whitelist as above.
|
||||
acl wp_batch_path path_beg /wp-json/batch/v1
|
||||
acl wp_batch_route urlp(rest_route) -i -m beg /batch/v1
|
||||
acl wp_batch_route_enc query -i -m sub rest_route=%2Fbatch%2Fv1
|
||||
acl has_wp_logged_in req.hdr(Cookie) -i -m sub wordpress_logged_in_
|
||||
http-request deny deny_status 403 if wp_batch_path !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
|
||||
http-request deny deny_status 403 if wp_batch_route !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
|
||||
http-request deny deny_status 403 if wp_batch_route_enc !has_wp_logged_in !is_local !is_trusted_ip !is_whitelisted
|
||||
|
||||
# IP blocking using map file (manual blocks only)
|
||||
# Map file format: /etc/haproxy/blocked_ips.map contains "<ip_or_cidr> 1" per line
|
||||
@@ -47,3 +133,52 @@ frontend web
|
||||
acl is_blocked_ip var(txn.real_ip),map_ip(/etc/haproxy/blocked_ips.map,0) -m int gt 0
|
||||
http-request set-path /blocked-ip if is_blocked_ip
|
||||
use_backend default-backend if is_blocked_ip
|
||||
{%- if suspension_enabled %}
|
||||
|
||||
# Site suspension routing. Any Host header listed in
|
||||
# /etc/haproxy/suspended_domains.list is rewritten to /suspended and
|
||||
# routed through default-backend, which is the same Flask app that
|
||||
# serves the default page and blocked-ip page (port 8080 inside this
|
||||
# container). The `/suspended` route returns HTTP 503 with a static
|
||||
# suspension page. External tooling (e.g. WHP's site_disable.php)
|
||||
# maintains the list file via `docker cp`. An empty list is safe —
|
||||
# the ACL simply doesn't match. Sits after IP-blocking so 429/403
|
||||
# still trigger first.
|
||||
acl is_suspended_domain hdr(host),lower -f /etc/haproxy/suspended_domains.list
|
||||
http-request set-path /suspended if is_suspended_domain
|
||||
use_backend default-backend if is_suspended_domain
|
||||
{%- endif %}
|
||||
{%- if coraza_spoe_backend %}
|
||||
|
||||
# Coraza WAF inspection via SPOE. Runs AFTER rate-limit and IP-block
|
||||
# guards (no point asking the WAF about requests we're already dropping)
|
||||
# and AFTER the real-client-IP resolution (so Coraza sees the right src).
|
||||
filter spoe engine coraza config /etc/haproxy/coraza-spoe.cfg
|
||||
http-request send-spoe-group coraza coraza-req
|
||||
|
||||
# Enforce Coraza's verdict. The SPOA sets var(txn.coraza.action) to
|
||||
# "deny" / "drop" / "redirect" when a rule with the corresponding
|
||||
# disruptive action fires (depends on SecRuleEngine mode + per-rule
|
||||
# ctl:ruleEngine overrides). Without these rules, Coraza would inspect
|
||||
# but never block.
|
||||
#
|
||||
# On request-phase deny we return a rendered HTML page that surfaces the
|
||||
# request reference (the unique-id) so a customer who's been blocked
|
||||
# incorrectly can open a support ticket and quote it. lf-file expands
|
||||
# log-format expressions inside the file at response time, so
|
||||
# %[unique-id] / %[req.hdr(host)] / etc. get substituted live.
|
||||
# Response-phase deny stays as a bare 403 — outbound blocks are rare in
|
||||
# our config (Coraza response inspection is disabled by default) and
|
||||
# an HTML body on a 403 generated mid-response could land mid-stream.
|
||||
http-request return status 403 content-type "text/html; charset=utf-8" hdr waf-block "request" hdr x-request-reference "%[unique-id]" lf-file /haproxy/errors/403-waf.html if { var(txn.coraza.action) -m str deny }
|
||||
http-response deny deny_status 403 hdr waf-block "response" hdr x-request-reference "%[unique-id]" if { var(txn.coraza.action) -m str deny }
|
||||
http-request silent-drop if { var(txn.coraza.action) -m str drop }
|
||||
http-response silent-drop if { var(txn.coraza.action) -m str drop }
|
||||
http-request redirect code 302 location %[var(txn.coraza.data)] if { var(txn.coraza.action) -m str redirect }
|
||||
http-response redirect code 302 location %[var(txn.coraza.data)] if { var(txn.coraza.action) -m str redirect }
|
||||
|
||||
# FAIL-OPEN on SPOA error. Upstream's example does the opposite — denies
|
||||
# 500 if var(txn.coraza.error) is set — but for a hosting platform we'd
|
||||
# rather lose WAF coverage briefly than 503 customer sites. The error
|
||||
# variable still gets set, so monitoring can observe it.
|
||||
{%- endif %}
|
||||
|
||||
@@ -6,3 +6,11 @@ frontend stats
|
||||
stats refresh 30s
|
||||
stats show-legends
|
||||
stats show-node
|
||||
|
||||
# Dedicated stick-table for WordPress wp-login.php brute-force tracking.
|
||||
# Tracked via track-sc1 from the `web` frontend (hap_listener.tpl); counts only
|
||||
# login POSTs per real client IP over a 60s window. Separate from the generic
|
||||
# sc0 connection/rate table so the login-attempt threshold is independent of
|
||||
# the (much higher) flood thresholds.
|
||||
backend wp_bruteforce
|
||||
stick-table type ip size 100k expire 30m store http_req_rate(60s)
|
||||
@@ -0,0 +1,58 @@
|
||||
<!DOCTYPE html>
|
||||
<html lang="en">
|
||||
<head>
|
||||
<meta charset="UTF-8">
|
||||
<meta name="viewport" content="width=device-width, initial-scale=1.0">
|
||||
<meta name="robots" content="noindex,nofollow">
|
||||
<title>Site temporarily unavailable</title>
|
||||
<style>
|
||||
body {
|
||||
font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, sans-serif;
|
||||
text-align: center;
|
||||
padding: 50px 20px;
|
||||
background: linear-gradient(135deg, #1e293b 0%, #0f172a 100%);
|
||||
margin: 0;
|
||||
min-height: 100vh;
|
||||
display: flex;
|
||||
align-items: center;
|
||||
justify-content: center;
|
||||
color: #e2e8f0;
|
||||
}
|
||||
.container {
|
||||
background: #1e293b;
|
||||
border: 1px solid #334155;
|
||||
padding: 40px;
|
||||
border-radius: 12px;
|
||||
box-shadow: 0 10px 30px rgba(0,0,0,0.4);
|
||||
max-width: 560px;
|
||||
width: 100%;
|
||||
}
|
||||
h1 {
|
||||
color: #f1f5f9;
|
||||
margin: 0 0 20px;
|
||||
font-size: 1.75em;
|
||||
font-weight: 600;
|
||||
}
|
||||
p {
|
||||
color: #cbd5e1;
|
||||
line-height: 1.7;
|
||||
margin: 0 0 12px;
|
||||
font-size: 1.05em;
|
||||
}
|
||||
.note {
|
||||
color: #94a3b8;
|
||||
font-size: 0.9em;
|
||||
margin-top: 24px;
|
||||
padding-top: 24px;
|
||||
border-top: 1px solid #334155;
|
||||
}
|
||||
</style>
|
||||
</head>
|
||||
<body>
|
||||
<div class="container">
|
||||
<h1>This site is temporarily unavailable.</h1>
|
||||
<p>The site you are trying to reach is currently offline.</p>
|
||||
<p class="note">Site owners: please contact support to restore service.</p>
|
||||
</div>
|
||||
</body>
|
||||
</html>
|
||||
@@ -0,0 +1,9 @@
|
||||
# Source-IP whitelist — exempt from HAProxy rate limits (one IP or CIDR per line).
|
||||
# Referenced by templates/hap_listener.tpl:
|
||||
# acl is_trusted_ip src -f /etc/haproxy/trusted_ips.list
|
||||
#
|
||||
# Add trusted source IPs below. Do NOT commit real/personal IPs to this repo —
|
||||
# it is mirrored publicly. Keep real entries in an untracked local copy, or add
|
||||
# them directly on the server (the file lives in the /etc/haproxy named volume
|
||||
# and persists across container recreates).
|
||||
127.0.0.1
|
||||
@@ -0,0 +1,9 @@
|
||||
# Real-IP whitelist for proxy-header matching — exempt from HAProxy rate limits.
|
||||
# Format: "<IP> 1" (one per line). Referenced by templates/hap_listener.tpl:
|
||||
# acl is_whitelisted var(txn.real_ip),map_ip(/etc/haproxy/trusted_ips.map,0) -m int gt 0
|
||||
#
|
||||
# Add trusted real IPs below. Do NOT commit real/personal IPs to this repo —
|
||||
# it is mirrored publicly. Keep real entries in an untracked local copy, or add
|
||||
# them directly on the server (the file lives in the /etc/haproxy named volume
|
||||
# and persists across container recreates).
|
||||
127.0.0.1 1
|
||||
Reference in New Issue
Block a user