Ship a firewall backend, and correct three claims review disproved
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 3s
Build App (Preview) / build-macos (pull_request) Successful in 2m46s
Build App (Preview) / build-linux (pull_request) Successful in 7m22s
Build App (Preview) / build-windows (pull_request) Successful in 7m43s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
Build Container / build-container (pull_request) Successful in 11m20s
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 3s
Build App (Preview) / build-macos (pull_request) Successful in 2m46s
Build App (Preview) / build-linux (pull_request) Successful in 7m22s
Build App (Preview) / build-windows (pull_request) Successful in 7m43s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
Build Container / build-container (pull_request) Successful in 11m20s
Review of #28 found the iptables exclusion was justified by a false premise, and I confirmed it: `wireguard-tools` declares `Recommends: nftables | iptables`, `--no-install-recommends` strips it, and `wg-quick`'s add_default() shells out to a firewall backend with no `type -p` guard. Measured on the image as this PR shipped it: [#] iptables-restore -n /usr/bin/wg-quick: line 32: iptables-restore: command not found wg-quick EXIT=127 That fires for `AllowedIPs = 0.0.0.0/0` — every stock full-tunnel config from every provider — not for a desktop client's killswitch as the comment claimed. Split tunnels are unaffected. Ship `nftables` rather than `iptables`: wg-quick prefers it (`type -p nft`, so with both installed iptables is dead weight), it is first in the package's own Recommends, and it is half the size. The review's proposed fix stopped there; it does not hold. Adding nftables does not make wg-quick work on this host, and neither does iptables: Warning: Extension CONNMARK revision 0 not supported, missing kernel module? `Table=auto` routes by fwmark and needs xt_CONNMARK from the *host* kernel. WSL2 has none and containers have no /lib/modules to load one from. So this fixes native Linux and Docker Desktop for Mac — which other WHP users are on — and cannot fix Docker Desktop for Windows, where the answer is to add routes with `ip route` directly. Documented rather than left to be rediscovered. Also from review: - "`ip` and `wg` are always present" was false. A project keeps the base image it was first built from, so this reaches new projects only. Reworded to match the wording already used for the Playwright libraries, and `/usr/bin/wg` added to FEATURE_PROBES so an existing project is *told* it is missing VPN tooling and prompted to migrate, rather than finding out via `wg: command not found`. - "no client is installed" contradicted shipping `wg` four lines earlier. The true claim is that no tunnel is configured or started. - The size figure measured against bare ubuntu:24.04, which over-counts by the ~209 kB of libelf1t64 the real base already has, and covered one arch. Now measured against the current base on amd64 and stated for arm64 too, per the standard CLAUDE.md sets for the Playwright layer. - `/run` persistence conflated two mechanisms: same-container files on a stop/start, `docker commit` on a recreation. Both stated, plus the corollary that key material written to /run ends up inside a snapshot image — observed, a `wg.priv` was already sitting in one. - The DNS bullet presented a Docker Desktop address as the general case. Now leads with the mechanism, notes 127.0.0.11 on a user-defined network is unaffected, and adds the two things the advice omitted: a resolver the tunnel can reach (or it leaks every query), and pinning the endpoint via the old gateway (or the tunnel routes through itself). Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
+32
-18
@@ -475,13 +475,18 @@ When enabled, the host Docker socket is mounted into the container so Claude Cod
|
||||
|
||||
When enabled, the container is given the three things a VPN client needs to build a tunnel:
|
||||
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
|
||||
sysctl that WireGuard requires. The `ip` and `wg` commands are always present to use them. This is
|
||||
**off by default**.
|
||||
sysctl that WireGuard requires. This is **off by default**.
|
||||
|
||||
The `ip`, `wg` and `nft` commands ship in the container image so there is something able to use
|
||||
them. If your project's container was created from an older base image it will not have them, and
|
||||
`wg` will simply not be found — **migrating the project onto the current base image** is what picks
|
||||
them up. Installing them by hand with `sudo apt install wireguard-tools` works in the meantime, but
|
||||
lives in the writable layer, so it is undone by a **Reset** and by a migration.
|
||||
|
||||
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
|
||||
redirected, and no client is installed or started on your behalf. Enabling it and expecting the
|
||||
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
|
||||
container's traffic to start leaving through a VPN is the most common misreading of what it does —
|
||||
installing a client and routing traffic into it remains yours to do.
|
||||
configuring a tunnel and routing traffic into it remains yours to do.
|
||||
|
||||
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
|
||||
the connection attempt **hangs until it times out** — a default container has no tun device to open
|
||||
@@ -504,20 +509,29 @@ Things worth knowing:
|
||||
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
|
||||
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
|
||||
- **No tunnel survives a restart.** The network namespace is built fresh every time the container
|
||||
starts, and there is no service manager inside to reconnect anything. Files under `/run` may
|
||||
persist via the snapshot and make it *look* like the tunnel is still configured, but after any
|
||||
stop/start, Reset or recreation the interface and its routes are gone and traffic goes out your
|
||||
real address again — with no error and nothing visibly different. Re-establish it after every
|
||||
start, and check rather than assume.
|
||||
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Containers
|
||||
resolve through an address on the Docker network (`192.168.65.7` under Docker Desktop) that sits
|
||||
outside the container's own subnet, so a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus
|
||||
`128.0.0.0/1` pair — captures it and sends every lookup into a tunnel that cannot carry it. The
|
||||
symptom is total: Claude Code reports it cannot connect, because it cannot resolve
|
||||
`api.anthropic.com`. Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16`
|
||||
via the original gateway, and use the VPN provider's own resolver for everything else. Note that
|
||||
a health check which fetches an IP literal such as `1.1.1.1` passes cleanly while this is broken —
|
||||
resolve a name instead.
|
||||
starts, and there is no service manager inside to reconnect anything. Leftover state under `/run`
|
||||
makes it *look* like the tunnel is still configured — that directory is in the container's
|
||||
writable layer, so it is simply still there after a stop/start, and `docker commit` carries it
|
||||
into the snapshot that a recreation is built from. Either way the interface and its routes are
|
||||
gone and traffic goes out your real address again, with no error and nothing visibly different.
|
||||
Re-establish it after every start, and check rather than assume.
|
||||
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your
|
||||
resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own
|
||||
subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures
|
||||
it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is
|
||||
`192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`,
|
||||
which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is
|
||||
total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`.
|
||||
Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original
|
||||
gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own,
|
||||
or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own
|
||||
address via the original gateway, or the tunnel's encrypted packets try to route through the
|
||||
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
|
||||
while DNS is broken — resolve a name instead.
|
||||
- **`wg-quick` cannot bring up a full tunnel on Docker Desktop for Windows.** Its `Table=auto` mode
|
||||
routes by firewall mark and needs `xt_CONNMARK` from the host kernel, which WSL2's does not have
|
||||
and a container cannot load. Split tunnels (a specific `AllowedIPs`) work fine, as does adding
|
||||
the routes yourself with `ip route`. Native Linux and Docker Desktop for Mac are unaffected.
|
||||
|
||||
> This setting can only be changed when the container is stopped. Capabilities and devices are
|
||||
> fixed when a container is created, so toggling it recreates the container on the next start.
|
||||
|
||||
Reference in New Issue
Block a user