diff --git a/CLAUDE.md b/CLAUDE.md index a7b9296..251f2c5 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -203,7 +203,8 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li ### Container (`container/`) - **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared - libraries a browser links against (see below) + libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled` + toggle grants capability for (`iproute2`, `wireguard-tools`, `iptables`) - **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs `npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium @@ -307,6 +308,39 @@ container is created once by a very long function where a dropped capability is container and make the switch impossible to turn off. - Off is byte-identical to a container created before the feature existed, and a missing label reads as `false`, so no existing project is churned. +- **The toggle grants capability and stops there — it routes nothing.** `vpn_host_config()` returns + a cap, a device and a sysctl; no client is installed, no route is touched, no tunnel is started + or restored. Users read the name as "turn the VPN on" and report the default network not routing + through it as a bug. It isn't, and the docs say so explicitly; keep it that way. +- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in + `container/Dockerfile` because a runtime install lands in the writable layer and is lost on + base-image migration — leaving a project holding the capability with nothing able to exercise it, + and no error that points at why. `iptables` is included and `nftables` deliberately is not; see + the Dockerfile comment for why that way round. +- **Anything built on this fails open.** The network namespace is rebuilt on every start and no + service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover + `/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the + writable layer, so on a stop/start it is simply the same container's files, and on a recreation + `docker commit` has carried it into the snapshot. Traffic silently reverts to the real address. + Any future autostart or killswitch work starts here. +- **`/run` riding the snapshot means a VPN client's key material can end up in an image.** Verified: + a fresh container off the whp snapshot already contained the `wg.priv` a previous tunnel left in + `/run`. Anything writing key material there inherits the problem — the same `docker commit` + hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks + ephemeral and is not. A VPN client that does this should delete its key on teardown. +- **`iptables` is baked, and picking `nftables` instead would have been wrong.** `Recommends: + nftables | iptables` is stripped by `--no-install-recommends`, and `wg-quick` needs a backend for + any `AllowedIPs = 0.0.0.0/0`. `nftables` is the tempting choice — preferred by `wg-quick`, half + the size — but `wg-quick` picks nft *unconditionally* when present, and its nft ruleset needs + `nft_fib_ipv4`, which LinuxKit (Docker Desktop for Mac) does not build while it *does* build + `xt_CONNMARK`. Shipping nftables would therefore have forfeited Mac. See the Dockerfile comment; + the kernel-config evidence is quoted there. +- **Two `wg-quick` failures remain, and only one is ours to fix.** Full tunnels still need + `xt_CONNMARK`, which WSL2 before 6.6 lacks — nothing installable changes that. And every + provider's stock config carries a `DNS =` line that fails in `set_dns()` before any routing, so it + breaks split tunnels too; `openresolv` has no candidate on noble and `resolvconf` drags in + systemd-resolved, so that one is documented rather than fixed. Driving `wg` and `ip route` + directly avoids both, which is what the skill does. ### Container Lifecycle diff --git a/HOW-TO-USE.md b/HOW-TO-USE.md index 166292b..85f53b6 100644 --- a/HOW-TO-USE.md +++ b/HOW-TO-USE.md @@ -477,8 +477,19 @@ When enabled, the container is given the three things a VPN client needs to buil the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark` sysctl that WireGuard requires. This is **off by default**. -Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but -the connection attempt **hangs until it times out** — a default container has no tun device to open +The `ip`, `wg` and `iptables` commands ship in the container image so there is something able to use +them. If your project's container was created from an older base image it will not have them, and +`wg` will simply not be found — **migrating the project onto the current base image** is what picks +them up. `sudo apt install iproute2 wireguard-tools iptables` works in the meantime, but lives in +the writable layer, so it is undone by a **Reset** and by a migration. + +**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is +redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the +container's traffic to start leaving through a VPN is the most common misreading of what it does — +configuring a tunnel and routing traffic into it remains yours to do. + +With the setting **off**, a client such as PIA or OpenVPN installs and its daemon starts normally, +but the connection attempt **hangs until it times out** — a default container has no tun device to open and no permission to add an interface or a route, and most clients report that as a generic timeout rather than a permissions error. @@ -497,6 +508,40 @@ Things worth knowing: **start**, with an error naming `/dev/net/tun` and pointing back at this setting. - A VPN client's kill switch applies to everything in the container, Claude Code included. If the tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off. +- **No tunnel survives a restart.** The network namespace is built fresh every time the container + starts, and there is no service manager inside to reconnect anything. Leftover state under `/run` + makes it *look* like the tunnel is still configured — that directory is in the container's + writable layer, so it is simply still there after a stop/start, and `docker commit` carries it + into the snapshot that a recreation is built from. Either way the interface and its routes are + gone and traffic goes out your real address again, with no error and nothing visibly different. + Re-establish it after every start, and check rather than assume. +- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your + resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own + subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures + it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is + `192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`, + which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is + total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`. + Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original + gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own, + or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own + address via the original gateway, or the tunnel's encrypted packets try to route through the + tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly + while DNS is broken — resolve a name instead. +- **Delete a client's key material when you tear a tunnel down.** Anything written under `/run` is + in the container's writable layer, and recreating or migrating the project runs `docker commit` + over it — so a WireGuard private key left there gets baked into the project's snapshot image and + copied forward from then on. This is not hypothetical; it has already happened here. +- **Strip the `DNS =` line from a provider's `.conf` before `wg-quick up`.** Every commercial + provider ships one, and `wg-quick` hands it to `resolvconf`, which is not installed — so it fails + at `resolvconf: command not found` and deletes the interface again. This happens before any + routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive + `wg` and `ip route` directly rather than going through `wg-quick`. +- **`wg-quick` full tunnels additionally need `xt_CONNMARK` from the host kernel.** WSL2 kernels + before 6.6 do not have it and a container cannot load one — on Windows, `wsl --update` moves you + to a current kernel, which does. Failing that, add the routes yourself with `ip route`, which + needs no firewall backend on any platform. Note this is the *second* hurdle: clear the `DNS =` + one above first, or you will not reach this. > This setting can only be changed when the container is stopped. Capabilities and devices are > fixed when a container is created, so toggling it recreates the container on the next start. @@ -1256,6 +1301,7 @@ The sandbox container (Ubuntu 24.04) comes pre-installed with: | ruff | Latest | Python linter/formatter | | Rust | Stable | Rust development (via rustup) | | Docker CLI | Latest | Container management (when spawning is enabled) | +| iproute2, WireGuard tools, iptables | Latest | Building a tunnel (when VPN Support is enabled) | | git | Latest | Version control | | GitHub CLI (gh) | Latest | GitHub integration | | AWS CLI | v2 | AWS services and Bedrock | diff --git a/app/src-tauri/src/docker/migration.rs b/app/src-tauri/src/docker/migration.rs index 6e9fecf..a08776a 100644 --- a/app/src-tauri/src/docker/migration.rs +++ b/app/src-tauri/src/docker/migration.rs @@ -135,6 +135,7 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[ ("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"), ("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"), ("/opt/mission-control", "Mission Control (Flight Control)"), + ("/usr/bin/wg", "VPN tooling for the VPN Support toggle (WireGuard)"), ]; /// Headroom demanded on Docker's storage backend on top of the measured diff --git a/container/Dockerfile b/container/Dockerfile index cefd436..b94b76c 100644 --- a/container/Dockerfile +++ b/container/Dockerfile @@ -34,6 +34,9 @@ RUN for i in 1 2 3 4 5; do \ cron \ bubblewrap \ socat \ + iproute2 \ + wireguard-tools \ + iptables \ && rm -rf /var/lib/apt/lists/* # `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither @@ -42,6 +45,83 @@ RUN for i in 1 2 3 4 5; do \ # corporate CA, no matter what the system trust store says. entrypoint.sh # degrades to a warning if it is ever missing. +# `iproute2`, `wireguard-tools` and `iptables` above are what the VPN support +# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a +# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to +# add a route, and without `wg` no way to build the tunnel those two exist to +# serve — a capability with nothing able to use it. +# +# They are baked rather than left to a runtime `apt-get install` for the same +# reason as the Playwright libraries below: the writable layer is re-paid after +# every Reset and lost on base-image migration. A hand-installed `wg` therefore +# works right up until an upgrade, then disappears and takes the tunnel with it +# — silently, since a VPN that fails to come up looks exactly like one that was +# never started. +# +# Measured against the *current base image*, since a bare ubuntu:24.04 also +# pulls libelf1t64 and netbase, which this base already has, and so over-reports +# by ~258 kB: **+12 packages, 7,203 kB on amd64**. The same set on arm64 is +# ~14.4 MB — the package list is identical on both arches, the binaries are +# simply larger (measured as 14.7 MB on arm64 ubuntu:24.04, less that 258 kB). +# +# ## Why `iptables`, and not `nftables` +# +# `wireguard-tools` declares `Recommends: nftables | iptables`, which the +# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s +# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e. +# every stock full-tunnel config every provider hands out — and it shells out to +# a firewall backend with no `type -p` guard. Measured with neither installed: +# +# [#] iptables-restore -n +# /usr/bin/wg-quick: line 32: iptables-restore: command not found +# wg-quick EXIT=127 +# +# `nftables` looks like the better pick — wg-quick prefers it, it is first in +# that Recommends, it is half the size — and it is the wrong one. wg-quick picks +# nft *unconditionally* when present (`if type -p nft`, line 241), so installing +# it makes the iptables path unreachable; and its nft ruleset needs three +# expression families where the iptables path needs one. Isolating them on a +# WSL2 host, the two connmark rules install fine and this is what fails: +# +# nft add rule ... fib saddr type != local drop +# Error: Could not process rule: No such file or directory +# ^^^^^^^^^^^^^^ needs nft_fib_ipv4 +# +# The choice therefore turns on which kernel symbol each path needs, and the two +# are not equally safe to bet on. `xt_CONNMARK` (iptables) was present in every +# kernel config examined — LinuxKit's for both arches, and WSL2's from 6.6. +# `nft_fib_ipv4` (nftables) was absent from the LinuxKit config read here, and a +# later review argued Docker Desktop has since enabled it and no longer builds +# from that config at all. That may well be true; it could not be settled from a +# Linux host, and it is the point: nftables' viability varies by Docker Desktop +# version in a way nobody here can pin down, while iptables' requirement did not +# vary anywhere it was checked. +# +# So `iptables` is chosen for being robust to that uncertainty rather than for +# beating nftables on any particular host. If nft_fib_ipv4 is present, wg-quick +# never reaches the iptables path and this costs 1.6 MB and nothing else; if it +# is absent, this is the difference between a working full tunnel and none. +# +# The residual gap is WSL2 before 6.6, which has neither symbol. Nothing +# installable in the container changes that — but `wsl --update` does, and moves +# the host to a far newer kernel. Add the routes with `ip route` in the meantime; +# that needs no firewall backend on any platform. +# +# ## What this still does not fix +# +# `wireguard-tools` only *Suggests* `openresolv | resolvconf`, so neither is +# installed, and every provider's stock config carries a `DNS =` line. That +# fails in `set_dns()`, *before* the firewall step, so it takes split tunnels +# down too: +# +# [#] resolvconf -a wg0 -m 0 -x +# /usr/bin/wg-quick: line 32: resolvconf: command not found +# +# Deliberately not fixed here: `openresolv` has no installation candidate on +# noble, and `resolvconf` resolves only by pulling in systemd-resolved — a +# resolver daemon and systemd units, into a container with no systemd. Strip the +# `DNS =` line and set the resolver another way. Documented in HOW-TO-USE.md. + # Remove default ubuntu user to free UID 1000 for host-user remapping RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \ && if getent group ubuntu >/dev/null 2>&1; then groupdel ubuntu 2>/dev/null || true; fi