Compare commits
3
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
49ac673045 | ||
|
|
cb848e110e | ||
|
|
87184a4be9 |
@@ -204,7 +204,7 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li
|
||||
|
||||
- **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared
|
||||
libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled`
|
||||
toggle grants capability for (`iproute2`, `wireguard-tools`, `iptables`)
|
||||
toggle grants capability for (`iproute2`, `wireguard-tools`, `nftables`)
|
||||
- **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs
|
||||
`npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own
|
||||
dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium
|
||||
@@ -315,8 +315,7 @@ container is created once by a very long function where a dropped capability is
|
||||
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
|
||||
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
|
||||
base-image migration — leaving a project holding the capability with nothing able to exercise it,
|
||||
and no error that points at why. `iptables` is included and `nftables` deliberately is not; see
|
||||
the Dockerfile comment for why that way round.
|
||||
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
|
||||
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
|
||||
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
|
||||
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
|
||||
@@ -328,19 +327,12 @@ container is created once by a very long function where a dropped capability is
|
||||
`/run`. Anything writing key material there inherits the problem — the same `docker commit`
|
||||
hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks
|
||||
ephemeral and is not. A VPN client that does this should delete its key on teardown.
|
||||
- **`iptables` is baked, and picking `nftables` instead would have been wrong.** `Recommends:
|
||||
nftables | iptables` is stripped by `--no-install-recommends`, and `wg-quick` needs a backend for
|
||||
any `AllowedIPs = 0.0.0.0/0`. `nftables` is the tempting choice — preferred by `wg-quick`, half
|
||||
the size — but `wg-quick` picks nft *unconditionally* when present, and its nft ruleset needs
|
||||
`nft_fib_ipv4`, which LinuxKit (Docker Desktop for Mac) does not build while it *does* build
|
||||
`xt_CONNMARK`. Shipping nftables would therefore have forfeited Mac. See the Dockerfile comment;
|
||||
the kernel-config evidence is quoted there.
|
||||
- **Two `wg-quick` failures remain, and only one is ours to fix.** Full tunnels still need
|
||||
`xt_CONNMARK`, which WSL2 before 6.6 lacks — nothing installable changes that. And every
|
||||
provider's stock config carries a `DNS =` line that fails in `set_dns()` before any routing, so it
|
||||
breaks split tunnels too; `openresolv` has no candidate on noble and `resolvconf` drags in
|
||||
systemd-resolved, so that one is documented rather than fixed. Driving `wg` and `ip route`
|
||||
directly avoids both, which is what the skill does.
|
||||
- **`wg-quick` full tunnels need `xt_CONNMARK` from the host kernel**, which WSL2 does not have and
|
||||
a container cannot load; `Recommends: nftables | iptables` is also stripped by
|
||||
`--no-install-recommends`, so `nftables` is baked explicitly. See the Dockerfile comment — the
|
||||
short version is that shipping the backend fixes native Linux and Docker Desktop for Mac, nothing
|
||||
fixes Docker Desktop for Windows, and adding the routes directly with `ip route` sidesteps it on
|
||||
all three.
|
||||
- **The `pia-vpn` skill is installed *and removed* from `VPN_SUPPORT_ENABLED`.** `container/skills/`
|
||||
is baked to `/opt/triple-c-skills` and `install_feature_skill()` in `entrypoint.sh` copies it into
|
||||
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches any project whose base
|
||||
|
||||
+11
-22
@@ -477,11 +477,11 @@ When enabled, the container is given the three things a VPN client needs to buil
|
||||
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
|
||||
sysctl that WireGuard requires. This is **off by default**.
|
||||
|
||||
The `ip`, `wg` and `iptables` commands ship in the container image so there is something able to use
|
||||
The `ip`, `wg` and `nft` commands ship in the container image so there is something able to use
|
||||
them. If your project's container was created from an older base image it will not have them, and
|
||||
`wg` will simply not be found — **migrating the project onto the current base image** is what picks
|
||||
them up. `sudo apt install iproute2 wireguard-tools iptables` works in the meantime, but lives in
|
||||
the writable layer, so it is undone by a **Reset** and by a migration.
|
||||
them up. Installing them by hand with `sudo apt install wireguard-tools` works in the meantime, but
|
||||
lives in the writable layer, so it is undone by a **Reset** and by a migration.
|
||||
|
||||
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
|
||||
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
|
||||
@@ -496,11 +496,11 @@ It needs your PIA credentials in `~/pia-creds`, two lines, username then passwor
|
||||
different provider, ignore it and set up your own client; nothing else depends on it.
|
||||
|
||||
Like the VPN tooling above, the skill ships in the container image, so a project whose container
|
||||
predates it will not get one by toggling the setting — **migrate the project** and it appears; the
|
||||
migration pre-flight lists it among what you would gain.
|
||||
predates it will not get one by toggling the setting — **migrate the project** and it appears. The
|
||||
container says so on start when that is the case, rather than leaving you to wonder where it went.
|
||||
|
||||
With the setting **off**, a client such as PIA or OpenVPN installs and its daemon starts normally,
|
||||
but the connection attempt **hangs until it times out** — a default container has no tun device to open
|
||||
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
|
||||
the connection attempt **hangs until it times out** — a default container has no tun device to open
|
||||
and no permission to add an interface or a route, and most clients report that as a generic timeout
|
||||
rather than a permissions error.
|
||||
|
||||
@@ -539,20 +539,10 @@ Things worth knowing:
|
||||
address via the original gateway, or the tunnel's encrypted packets try to route through the
|
||||
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
|
||||
while DNS is broken — resolve a name instead.
|
||||
- **Delete a client's key material when you tear a tunnel down.** Anything written under `/run` is
|
||||
in the container's writable layer, and recreating or migrating the project runs `docker commit`
|
||||
over it — so a WireGuard private key left there gets baked into the project's snapshot image and
|
||||
copied forward from then on. This is not hypothetical; it has already happened here.
|
||||
- **Strip the `DNS =` line from a provider's `.conf` before `wg-quick up`.** Every commercial
|
||||
provider ships one, and `wg-quick` hands it to `resolvconf`, which is not installed — so it fails
|
||||
at `resolvconf: command not found` and deletes the interface again. This happens before any
|
||||
routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive
|
||||
`wg` and `ip route` directly rather than going through `wg-quick`.
|
||||
- **`wg-quick` full tunnels additionally need `xt_CONNMARK` from the host kernel.** WSL2 kernels
|
||||
before 6.6 do not have it and a container cannot load one — on Windows, `wsl --update` moves you
|
||||
to a current kernel, which does. Failing that, add the routes yourself with `ip route`, which
|
||||
needs no firewall backend on any platform. Note this is the *second* hurdle: clear the `DNS =`
|
||||
one above first, or you will not reach this.
|
||||
- **`wg-quick` cannot bring up a full tunnel on Docker Desktop for Windows.** Its `Table=auto` mode
|
||||
routes by firewall mark and needs `xt_CONNMARK` from the host kernel, which WSL2's does not have
|
||||
and a container cannot load. Split tunnels (a specific `AllowedIPs`) work fine, as does adding
|
||||
the routes yourself with `ip route`. Native Linux and Docker Desktop for Mac are unaffected.
|
||||
|
||||
> This setting can only be changed when the container is stopped. Capabilities and devices are
|
||||
> fixed when a container is created, so toggling it recreates the container on the next start.
|
||||
@@ -1312,7 +1302,6 @@ The sandbox container (Ubuntu 24.04) comes pre-installed with:
|
||||
| ruff | Latest | Python linter/formatter |
|
||||
| Rust | Stable | Rust development (via rustup) |
|
||||
| Docker CLI | Latest | Container management (when spawning is enabled) |
|
||||
| iproute2, WireGuard tools, iptables | Latest | Building a tunnel (when VPN Support is enabled) |
|
||||
| git | Latest | Version control |
|
||||
| GitHub CLI (gh) | Latest | GitHub integration |
|
||||
| AWS CLI | v2 | AWS services and Bedrock |
|
||||
|
||||
@@ -135,8 +135,8 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
|
||||
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
|
||||
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
|
||||
("/opt/mission-control", "Mission Control (Flight Control)"),
|
||||
("/usr/bin/wg", "VPN tooling for the VPN Support toggle (WireGuard)"),
|
||||
("/opt/triple-c-skills", "Bundled skills for the VPN Support toggle (PIA VPN)"),
|
||||
("/usr/bin/wg", "VPN support (WireGuard tools)"),
|
||||
("/opt/triple-c-skills", "Feature skills (PIA VPN)"),
|
||||
];
|
||||
|
||||
/// Headroom demanded on Docker's storage backend on top of the measured
|
||||
|
||||
+24
-51
@@ -36,7 +36,7 @@ RUN for i in 1 2 3 4 5; do \
|
||||
socat \
|
||||
iproute2 \
|
||||
wireguard-tools \
|
||||
iptables \
|
||||
nftables \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither
|
||||
@@ -45,7 +45,7 @@ RUN for i in 1 2 3 4 5; do \
|
||||
# corporate CA, no matter what the system trust store says. entrypoint.sh
|
||||
# degrades to a warning if it is ever missing.
|
||||
|
||||
# `iproute2`, `wireguard-tools` and `iptables` above are what the VPN support
|
||||
# `iproute2`, `wireguard-tools` and `nftables` above are what the VPN support
|
||||
# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a
|
||||
# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to
|
||||
# add a route, and without `wg` no way to build the tunnel those two exist to
|
||||
@@ -58,69 +58,42 @@ RUN for i in 1 2 3 4 5; do \
|
||||
# — silently, since a VPN that fails to come up looks exactly like one that was
|
||||
# never started.
|
||||
#
|
||||
# Measured against the *current base image*, since a bare ubuntu:24.04 also
|
||||
# pulls libelf1t64 and netbase, which this base already has, and so over-reports
|
||||
# by ~258 kB: **+12 packages, 7,203 kB on amd64**. The same set on arm64 is
|
||||
# ~14.4 MB — the package list is identical on both arches, the binaries are
|
||||
# simply larger (measured as 14.7 MB on arm64 ubuntu:24.04, less that 258 kB).
|
||||
# Measured against the *current base image*, not a bare ubuntu:24.04 — the base
|
||||
# already ships libelf1t64, so measuring on bare ubuntu over-counts by ~209 kB:
|
||||
# +9 packages, 5,614 kB on amd64 (4,153 kB of that is iproute2+wireguard-tools,
|
||||
# 1,461 kB is nftables). The same set on arm64 is 7,422 kB, measured against
|
||||
# ubuntu:24.04 since the arm64 base is not cached here.
|
||||
#
|
||||
# ## Why `iptables`, and not `nftables`
|
||||
# ## Why `nftables` specifically
|
||||
#
|
||||
# `wireguard-tools` declares `Recommends: nftables | iptables`, which the
|
||||
# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s
|
||||
# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e.
|
||||
# every stock full-tunnel config every provider hands out — and it shells out to
|
||||
# a firewall backend with no `type -p` guard. Measured with neither installed:
|
||||
# a firewall backend with no `type -p` guard. Measured without one:
|
||||
#
|
||||
# [#] iptables-restore -n
|
||||
# /usr/bin/wg-quick: line 32: iptables-restore: command not found
|
||||
# wg-quick EXIT=127
|
||||
# wg-quick EXIT=127 (interface rolled back, split tunnels unaffected)
|
||||
#
|
||||
# `nftables` looks like the better pick — wg-quick prefers it, it is first in
|
||||
# that Recommends, it is half the size — and it is the wrong one. wg-quick picks
|
||||
# nft *unconditionally* when present (`if type -p nft`, line 241), so installing
|
||||
# it makes the iptables path unreachable; and its nft ruleset needs three
|
||||
# expression families where the iptables path needs one. Isolating them on a
|
||||
# WSL2 host, the two connmark rules install fine and this is what fails:
|
||||
# `nftables` rather than `iptables` because `wg-quick` prefers it (`if type -p
|
||||
# nft`, so with both installed iptables is dead weight), it is the first
|
||||
# alternative in the package's own Recommends, and it is roughly half the size.
|
||||
#
|
||||
# nft add rule ... fib saddr type != local drop
|
||||
# Error: Could not process rule: No such file or directory
|
||||
# ^^^^^^^^^^^^^^ needs nft_fib_ipv4
|
||||
# This does NOT make `wg-quick`'s full-tunnel mode work everywhere. `Table=auto`
|
||||
# routes by fwmark and needs connection-mark tracking from the *host* kernel:
|
||||
#
|
||||
# The choice therefore turns on which kernel symbol each path needs, and the two
|
||||
# are not equally safe to bet on. `xt_CONNMARK` (iptables) was present in every
|
||||
# kernel config examined — LinuxKit's for both arches, and WSL2's from 6.6.
|
||||
# `nft_fib_ipv4` (nftables) was absent from the LinuxKit config read here, and a
|
||||
# later review argued Docker Desktop has since enabled it and no longer builds
|
||||
# from that config at all. That may well be true; it could not be settled from a
|
||||
# Linux host, and it is the point: nftables' viability varies by Docker Desktop
|
||||
# version in a way nobody here can pin down, while iptables' requirement did not
|
||||
# vary anywhere it was checked.
|
||||
# Warning: Extension CONNMARK revision 0 not supported, missing kernel module?
|
||||
#
|
||||
# So `iptables` is chosen for being robust to that uncertainty rather than for
|
||||
# beating nftables on any particular host. If nft_fib_ipv4 is present, wg-quick
|
||||
# never reaches the iptables path and this costs 1.6 MB and nothing else; if it
|
||||
# is absent, this is the difference between a working full tunnel and none.
|
||||
# WSL2's kernel has no `xt_CONNMARK` and containers have no /lib/modules to load
|
||||
# one from, so on Docker Desktop for Windows `wg-quick up` on a full tunnel fails
|
||||
# regardless of what is installed here. Native Linux and Docker Desktop for Mac
|
||||
# have it. Shipping the backend is what makes the difference on those two;
|
||||
# nothing shipped here can make the difference on WSL2, where the way out is to
|
||||
# add the routes with `ip route` instead of going through `wg-quick` at all.
|
||||
#
|
||||
# The residual gap is WSL2 before 6.6, which has neither symbol. Nothing
|
||||
# installable in the container changes that — but `wsl --update` does, and moves
|
||||
# the host to a far newer kernel. Add the routes with `ip route` in the meantime;
|
||||
# that needs no firewall backend on any platform.
|
||||
#
|
||||
# ## What this still does not fix
|
||||
#
|
||||
# `wireguard-tools` only *Suggests* `openresolv | resolvconf`, so neither is
|
||||
# installed, and every provider's stock config carries a `DNS =` line. That
|
||||
# fails in `set_dns()`, *before* the firewall step, so it takes split tunnels
|
||||
# down too:
|
||||
#
|
||||
# [#] resolvconf -a wg0 -m 0 -x
|
||||
# /usr/bin/wg-quick: line 32: resolvconf: command not found
|
||||
#
|
||||
# Deliberately not fixed here: `openresolv` has no installation candidate on
|
||||
# noble, and `resolvconf` resolves only by pulling in systemd-resolved — a
|
||||
# resolver daemon and systemd units, into a container with no systemd. Strip the
|
||||
# `DNS =` line and set the resolver another way. Documented in HOW-TO-USE.md.
|
||||
# `iptables` is deliberately still NOT here: with `nftables` present `wg-quick`
|
||||
# never reaches for it, so it would add size and firewall surface for nothing.
|
||||
|
||||
# Remove default ubuntu user to free UID 1000 for host-user remapping
|
||||
RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \
|
||||
|
||||
+6
-24
@@ -359,40 +359,22 @@ install_feature_skill() {
|
||||
local _src="/opt/triple-c-skills/$1"
|
||||
local _dest="/home/claude/.claude/skills/$1"
|
||||
|
||||
# Reject anything that is not a plain directory name. The disabled branch
|
||||
# `rm -rf`s $_dest under a *persisted volume*, so a blank name would take the
|
||||
# whole skills directory (Mission Control's included) and `../x` would escape
|
||||
# it entirely. Only the literal `pia-vpn` is passed today; this is so that
|
||||
# stays true.
|
||||
case "$_name" in
|
||||
''|*/*|.*) echo "entrypoint: install_feature_skill: bad skill name '$_name'"; return 1 ;;
|
||||
esac
|
||||
# A blank name would make the disabled branch `rm -rf` the whole skills
|
||||
# directory, Mission Control's included, under a persisted volume.
|
||||
[ -n "$_name" ] || { echo "entrypoint: install_feature_skill called with no name"; return 1; }
|
||||
|
||||
if [ "$_enabled" = "1" ]; then
|
||||
if [ ! -d "$_src" ]; then
|
||||
echo "entrypoint: $_name skill unavailable — this container's base image predates it; migrate the project to get it"
|
||||
return 0
|
||||
fi
|
||||
# Checked, not assumed: with no `set -e` in this script every step here
|
||||
# can fail (full volume, read-only mount, a file where the directory
|
||||
# should be) and the success line would still print.
|
||||
mkdir -p /home/claude/.claude/skills || {
|
||||
echo "entrypoint: $_name skill install FAILED (cannot create ~/.claude/skills)"; return 1; }
|
||||
mkdir -p /home/claude/.claude/skills
|
||||
# Not just $_dest: when Mission Control is off nothing else creates the
|
||||
# parent, so root would own it and `claude` could not add a skill there.
|
||||
chown claude:claude /home/claude/.claude/skills
|
||||
# Stage then swap. Copying over the live path meant a failure (full
|
||||
# volume, read-only mount) left a truncated SKILL.md and no script
|
||||
# behind, root-owned, on a persisted volume — which Claude Code then
|
||||
# discovers and loads.
|
||||
rm -rf "$_dest.new"
|
||||
cp -r "$_src" "$_dest.new" || {
|
||||
rm -rf "$_dest.new"
|
||||
echo "entrypoint: $_name skill install FAILED (copy from $_src); previous copy left intact"
|
||||
return 1; }
|
||||
chown -R claude:claude "$_dest.new"
|
||||
rm -rf "$_dest"
|
||||
mv "$_dest.new" "$_dest"
|
||||
cp -r "$_src" "$_dest"
|
||||
chown -R claude:claude "$_dest"
|
||||
echo "entrypoint: $_name skill installed to ~/.claude/skills/"
|
||||
elif [ -e "$_dest" ] || [ -L "$_dest" ]; then
|
||||
# -e/-L rather than -d: a leftover *file* at that path must go too.
|
||||
|
||||
@@ -9,7 +9,7 @@ Bring this container's traffic out through Private Internet Access over
|
||||
WireGuard, using the API PIA documents for headless use.
|
||||
|
||||
Run `sudo ~/.claude/skills/pia-vpn/pia-wg.sh` with `up`, `up --full`, `down` or
|
||||
`status`. Read the rest of this page before the first `up --full` — four of the
|
||||
`status`. Read the rest of this page before the first `up --full` — three of the
|
||||
behaviours below are actively misleading if you meet them without warning, and
|
||||
each one presents as "the VPN is fine" or "Claude is broken" rather than as
|
||||
what it is.
|
||||
@@ -107,21 +107,7 @@ mode: test route only (1.1.1.1 through the tunnel, nothing else)
|
||||
Two different addresses there is correct and expected in test mode. If you want
|
||||
the second line to change, you want `up --full`.
|
||||
|
||||
## Trap 4: a full tunnel hides the Docker host unless the name is pinned
|
||||
|
||||
`host.docker.internal` is answered *only* by the resolver that `up --full`
|
||||
replaces — it is not in `/etc/hosts`. Triple-C hands that name to the container
|
||||
for the LiteLLM gateway, and host-side Ollama and custom endpoints default to
|
||||
it, so losing the name takes the project's model backend down with it.
|
||||
|
||||
The nasty part is what a naive check reports. PIA's resolvers answer public
|
||||
names perfectly well, so a probe of `api.anthropic.com` says everything is fine
|
||||
while the Docker host has vanished. `pia-wg.sh` pins the address into
|
||||
`/etc/hosts` before swapping the resolver and restores the file on teardown, and
|
||||
`status` probes both names — but if you ever rewrite `resolv.conf` by hand, this
|
||||
is the one that will not announce itself.
|
||||
|
||||
## Trap 5: no tunnel survives a restart, and it fails open
|
||||
## Trap 4: no tunnel survives a restart, and it fails open
|
||||
|
||||
The network namespace is rebuilt every time the container starts, and nothing
|
||||
inside reconnects anything. After a stop/start, Reset or any config change that
|
||||
@@ -197,35 +183,22 @@ true in test mode too, and means much less than it sounds like.
|
||||
`down` restores `resolv.conf` from its backup (only if that backup still looks
|
||||
like a resolver file — restoring a truncated one would leave the container with
|
||||
no DNS at all), removes exactly the routes that were added, in reverse order,
|
||||
and deletes the interface. It is safe to run when nothing is up. Confirm
|
||||
afterwards that the public address is back to the container's own.
|
||||
deletes the interface, and shreds the WireGuard private key. It is safe to run
|
||||
when nothing is up, and `up` runs it first so a repeat `up` cannot stack state.
|
||||
Confirm afterwards that the public address is back to the container's own.
|
||||
|
||||
`up` calls it too, but only after the last network fetch — the key registration
|
||||
— has succeeded, so a failed `up` leaves an existing tunnel alone rather than
|
||||
tearing it down to report a bad password or an unreachable gateway. From that
|
||||
point on a rollback is armed: if any step of the setup fails, the tunnel is torn
|
||||
down rather than left half-configured.
|
||||
|
||||
The private key is deleted earlier still — the moment `wg set` has read it,
|
||||
while the tunnel is being built. That is not housekeeping: `/run` is in the
|
||||
container's writable layer, and recreating or migrating the project runs
|
||||
`docker commit` over it *without* tearing the tunnel down first. A key that
|
||||
lived for the tunnel's lifetime would be baked into the snapshot image and
|
||||
copied forward from then on. The kernel keeps its own copy, so nothing is lost.
|
||||
The key deletion is not housekeeping: `/run` is in the container's writable
|
||||
layer, so `docker commit` bakes whatever is there into the project's snapshot
|
||||
image. A key left behind rides that image into every future container.
|
||||
|
||||
## What this deliberately does not do
|
||||
|
||||
- **No killswitch.** `iptables` *is* in the image, so one is buildable — this
|
||||
is a deliberate omission, not a missing dependency. Blocking non-tunnel egress
|
||||
cuts Claude Code's own API traffic the moment the tunnel drops, which ends the
|
||||
session that would otherwise fix it. If the user needs guaranteed egress
|
||||
rather than convenient egress, say so plainly and let them decide, rather than
|
||||
improvising one.
|
||||
- **No killswitch.** Blocking non-tunnel egress needs `iptables`, which is not
|
||||
in the image, and would cut Claude Code's API traffic whenever the tunnel is
|
||||
down. If the user needs guaranteed egress rather than convenient egress, say
|
||||
so plainly rather than improvising one — it is a real design decision.
|
||||
- **No autostart.** There is no service manager in the container and Triple-C
|
||||
has no start hook, so nothing re-establishes the tunnel on its own. `cron` is
|
||||
in the image and `triple-c-scheduler` runs on it, so a scheduled reconnect is
|
||||
possible if the user wants one — it is just not set up, and a tunnel that
|
||||
reconnects unattended deserves an explicit decision.
|
||||
has no start hook, so nothing can re-establish the tunnel automatically.
|
||||
- **Not PIA's desktop client.** `pia-daemon` and `piactl` are installable but
|
||||
cannot work headless: the daemon never accepts a client connection without
|
||||
the GUI, and `piactl --help` states that connecting requires it. If you find
|
||||
|
||||
@@ -27,9 +27,6 @@
|
||||
set -euo pipefail
|
||||
|
||||
# Not ~/pia-creds: under sudo, HOME is /root.
|
||||
# Read by up()'s EXIT trap, which runs after the function's locals are gone.
|
||||
SETUP_OK=0
|
||||
|
||||
CREDS=${PIA_CREDS:-/home/claude/pia-creds}
|
||||
REGION=${PIA_REGION:-us_chicago}
|
||||
IFACE=pia0
|
||||
@@ -91,7 +88,7 @@ preflight() {
|
||||
# tunnel captures everything while the exclusions that keep DNS and the Docker
|
||||
# host reachable are quietly missing -- and `status` still says "full tunnel".
|
||||
add_route() {
|
||||
ip route add "$@" || die "could not add route '$*'"
|
||||
ip route add "$@" || die "could not add route '$*'. Run 'down' to undo the partial setup."
|
||||
printf '%s\n' "$*" >> "$STATE/routes"
|
||||
}
|
||||
|
||||
@@ -103,6 +100,11 @@ up() {
|
||||
esac
|
||||
preflight
|
||||
|
||||
# Always start from a known state. Without this a second `up` overwrites the
|
||||
# saved resolv.conf with PIA's own resolvers, so the later `down` "restores"
|
||||
# those and leaves the container with no working DNS and no way back.
|
||||
down >/dev/null 2>&1 || true
|
||||
|
||||
mkdir -p "$STATE"; cd "$STATE"
|
||||
|
||||
# `curl -o` creates the file before it knows the request failed, so a plain
|
||||
@@ -121,14 +123,8 @@ up() {
|
||||
u=$(sed -n 1p "$CREDS"); p=$(sed -n 2p "$CREDS")
|
||||
[ -n "$u" ] && [ -n "$p" ] || die "$CREDS needs two lines: username, then password"
|
||||
|
||||
# Via stdin, not `-u`. curl does blank the password in its own argv, but only
|
||||
# once it is running: sampling /proc/<pid>/cmdline in a tight loop caught the
|
||||
# plaintext in 3 of 200 tries, in the window between exec and the overwrite.
|
||||
# Small, but this is the permanent account password, and the mechanism to
|
||||
# avoid it entirely is already here for the token.
|
||||
tok=$(printf -- '--user "%s:%s"\n' "$u" "$p" \
|
||||
| run "PIA rejected the credentials in $CREDS, or could not be reached" \
|
||||
curl -sf -m 25 -K - \
|
||||
tok=$(run "PIA rejected the credentials in $CREDS, or could not be reached" \
|
||||
curl -sf -m 25 -u "$u:$p" \
|
||||
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
|
||||
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA returned no token - check the credentials in $CREDS"
|
||||
|
||||
@@ -139,9 +135,10 @@ up() {
|
||||
sip=$(echo "$srv" | jq -r .ip); scn=$(echo "$srv" | jq -r .cn)
|
||||
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region '$REGION'"
|
||||
|
||||
# The key is generated but NOT written yet -- `down` below deletes wg.priv, and
|
||||
# the teardown has to come after every fetch that can fail.
|
||||
priv=$(wg genkey); pub=$(printf '%s' "$priv" | wg pubkey)
|
||||
# umask, not a later chmod: the file is created under the inherited 0022
|
||||
# otherwise, so the key is world-readable for the moment in between.
|
||||
( umask 077; priv=$(wg genkey); printf '%s' "$priv" > wg.priv )
|
||||
priv=$(cat wg.priv); pub=$(printf '%s' "$priv" | wg pubkey)
|
||||
|
||||
# The token goes in on stdin as a curl config rather than in the argv, where
|
||||
# `ps` and /proc/*/cmdline expose it to every process in the container --
|
||||
@@ -154,32 +151,6 @@ up() {
|
||||
--cacert ca.rsa.4096.crt "https://$scn:1337/addKey")
|
||||
[ "$(echo "$resp" | jq -r .status)" = OK ] || die "key registration failed: $resp"
|
||||
|
||||
# Only now tear down any previous tunnel. Every network call above this line
|
||||
# can fail, and an earlier version tore down first -- so a failed token fetch,
|
||||
# an unreachable server list, or a refused key registration took a *working*
|
||||
# tunnel with it and silently reverted the container to its real address while
|
||||
# the error talked about credentials. Nothing above this line has touched the
|
||||
# network stack. addKey is the most failure-prone of the three: it reaches one
|
||||
# individual gateway by CN with a pinned certificate.
|
||||
#
|
||||
# It also still does the job it was added for: clearing a stale resolv.conf
|
||||
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
|
||||
down >/dev/null 2>&1 || true
|
||||
|
||||
# From here on the network stack is being modified, so any failure has to put
|
||||
# it back rather than exit half-configured.
|
||||
#
|
||||
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
|
||||
# ERR trap is not inherited by shell functions without `set -E`, so a failure
|
||||
# inside add_route would not fire it, and `die` exits explicitly, which is not
|
||||
# an error and would not fire it either. EXIT catches both.
|
||||
SETUP_OK=0
|
||||
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1 || true; }' EXIT
|
||||
|
||||
# umask, not a later chmod: created under the inherited 0022 otherwise, so the
|
||||
# key would be world-readable for the moment in between.
|
||||
( umask 077; printf '%s' "$priv" > wg.priv )
|
||||
|
||||
: > "$STATE/routes"
|
||||
ip link add "$IFACE" type wireguard 2>/dev/null || \
|
||||
die "could not create a WireGuard interface." \
|
||||
@@ -188,12 +159,6 @@ up() {
|
||||
peer "$(echo "$resp" | jq -r .server_key)" \
|
||||
endpoint "$(echo "$resp" | jq -r .server_ip):$(echo "$resp" | jq -r .server_port)" \
|
||||
allowed-ips 0.0.0.0/0 persistent-keepalive 25
|
||||
# The kernel holds the key from here, so the file has no reason to outlive
|
||||
# this line -- and every reason not to: /run is in the writable layer, and a
|
||||
# recreate or migrate runs `docker commit` over it without tearing the tunnel
|
||||
# down first, baking the key into the project's snapshot image. `down` also
|
||||
# removes it, for the case where `up` never got this far.
|
||||
rm -f wg.priv
|
||||
ip addr add "$(echo "$resp" | jq -r .peer_ip)/32" dev "$IFACE"
|
||||
ip link set "$IFACE" up
|
||||
|
||||
@@ -226,19 +191,6 @@ up() {
|
||||
|
||||
# PIA's resolvers live inside 10/8, so pin them back through the tunnel with
|
||||
# /32s -- longer still, so they beat the exclusion just added.
|
||||
# `host.docker.internal` is answered only by the resolver about to be
|
||||
# replaced -- it is not in /etc/hosts. Triple-C hands that name to the
|
||||
# container for the LiteLLM gateway and defaults host-side Ollama and custom
|
||||
# endpoints to it, so losing it takes the project's model backend with it.
|
||||
# The *route* to it is already excluded above; only the name needs pinning.
|
||||
# Resolve it with the old resolver and write it into /etc/hosts first.
|
||||
local hdi
|
||||
hdi=$(getent ahostsv4 host.docker.internal 2>/dev/null | awk '{print $1; exit}')
|
||||
if [ -n "$hdi" ]; then
|
||||
cp /etc/hosts "$STATE/hosts.bak"
|
||||
printf '%s host.docker.internal\n' "$hdi" >> /etc/hosts
|
||||
fi
|
||||
|
||||
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
|
||||
for d in $dns; do add_route "$d/32" dev "$IFACE"; done
|
||||
# resolv.conf is a bind mount: write through it, never replace it.
|
||||
@@ -249,34 +201,12 @@ up() {
|
||||
echo "test route only: 1.1.1.1 goes via PIA, everything else unchanged"
|
||||
fi
|
||||
|
||||
# A tunnel with no handshake still routes -- into a black hole. Without this
|
||||
# `up --full` would exit 0 having pointed all traffic *and* resolv.conf at a
|
||||
# peer that never answered, and `status` would print "mode: full tunnel".
|
||||
# Demand a number, not just "different from 0". `wg show` prints nothing at
|
||||
# all when the interface has no peer, and writes to stderr when the interface
|
||||
# is gone -- both leave $2 empty, and `[ "" != 0 ]` is true, so the original
|
||||
# form treated a missing tunnel as a completed handshake and exited 0. `until`
|
||||
# suspends both `set -e` and `pipefail`, so nothing else was going to catch it.
|
||||
local waited=0 hs
|
||||
until hs=$(wg show "$IFACE" latest-handshakes 2>/dev/null | awk 'NR==1{print $2}')
|
||||
[[ $hs =~ ^[0-9]+$ ]] && [ "$hs" -gt 0 ]; do
|
||||
waited=$((waited + 1))
|
||||
[ "$waited" -lt 40 ] || die "no handshake from $REGION after 20s"
|
||||
sleep 0.5
|
||||
done
|
||||
|
||||
SETUP_OK=1
|
||||
trap - EXIT
|
||||
sleep 2
|
||||
status
|
||||
}
|
||||
|
||||
down() {
|
||||
[ "$(id -u)" = 0 ] || die "run with sudo"
|
||||
# Teardown must finish even if a step fails; a half-rollback is the state this
|
||||
# exists to prevent. Deliberately not inherited from the caller's `set -e`.
|
||||
set +e
|
||||
# First, because removing the interface removes every route that points at it.
|
||||
ip link del "$IFACE" 2>/dev/null
|
||||
# Only restore something that actually looks like a resolver file. Restoring
|
||||
# an empty or truncated backup leaves the container with no DNS at all, which
|
||||
# is worse than leaving the current one alone.
|
||||
@@ -288,10 +218,6 @@ down() {
|
||||
fi
|
||||
rm -f "$STATE/resolv.conf.bak"
|
||||
fi
|
||||
if [ -f "$STATE/hosts.bak" ]; then
|
||||
cat "$STATE/hosts.bak" > /etc/hosts
|
||||
rm -f "$STATE/hosts.bak"
|
||||
fi
|
||||
if [ -f "$STATE/routes" ]; then
|
||||
# Reverse order: the specific overrides go before the ranges they sit in.
|
||||
tac "$STATE/routes" | while read -r r; do
|
||||
@@ -299,6 +225,7 @@ down() {
|
||||
done
|
||||
rm -f "$STATE/routes"
|
||||
fi
|
||||
ip link del "$IFACE" 2>/dev/null || true
|
||||
# /run is in the writable layer and `docker commit` bakes it into the
|
||||
# project's snapshot image, so a key left here rides that image into every
|
||||
# future container. Verified: a snapshot already carried one.
|
||||
@@ -315,24 +242,15 @@ TRACE_DIRECT=https://1.0.0.1/cdn-cgi/trace
|
||||
exit_ip() { curl -s -m 20 "$1" | sed -n 's/^ip=//p'; }
|
||||
|
||||
status() {
|
||||
# `wg show` needs root; `ip route`/`ip link` do not. Without this guard an
|
||||
# unprivileged run prints "no tunnel up" and then "mode: full tunnel" in the
|
||||
# same breath, and an agent reading the first line re-runs `up`.
|
||||
[ "$(id -u)" = 0 ] || die "run with sudo"
|
||||
wg show "$IFACE" 2>/dev/null | grep -E "latest handshake|transfer" || echo "no tunnel up"
|
||||
|
||||
# Resolve a name, not an IP literal. A curl to 1.1.1.1 succeeds while DNS is
|
||||
# completely broken, which is exactly how a dead resolver goes unnoticed.
|
||||
printf 'DNS: '
|
||||
if ! timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
|
||||
echo "BROKEN - cannot resolve api.anthropic.com"
|
||||
elif ! timeout 10 getent hosts host.docker.internal >/dev/null 2>&1; then
|
||||
# PIA's resolvers answer public names happily, so probing only
|
||||
# api.anthropic.com reports "ok" on a container that has just lost the
|
||||
# Docker host -- and with it the LiteLLM gateway and any host-side Ollama.
|
||||
echo "public ok, but host.docker.internal is UNRESOLVABLE (gateway/Ollama backends will fail)"
|
||||
else
|
||||
if timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
|
||||
echo "ok (via $(sed -n 's/^nameserver //p' /etc/resolv.conf | tr '\n' ' '))"
|
||||
else
|
||||
echo "BROKEN - cannot resolve api.anthropic.com"
|
||||
fi
|
||||
|
||||
# Report the exit per mode. In test mode the probe address is itself the one
|
||||
@@ -356,5 +274,5 @@ case "${1:-}" in
|
||||
up) shift; up "${1:-}" ;;
|
||||
down) down ;;
|
||||
status) status ;;
|
||||
*) sed -n '2,26p' "$0" | sed 's/^# \{0,1\}//'; exit 1 ;;
|
||||
*) sed -n '2,27p' "$0" | sed 's/^# \{0,1\}//'; exit 1 ;;
|
||||
esac
|
||||
|
||||
Reference in New Issue
Block a user