Compare commits

..
6 Commits
Author SHA1 Message Date
shadow-testandClaude Opus 5 48d0c3249a Fix two bugs in last round's fixes, and stop --full hiding the Docker host
Round 3 found defects in code written an hour earlier. Both reproduced.

**The handshake poll accepted empty output as a completed handshake.**
`[ "$(… | awk '{print $2}')" != 0 ]` is *true* when `wg show` prints nothing —
which it does when the interface has no peer, and when the interface is gone
(that message goes to stderr). `until` suspends `set -e` and `pipefail`, so
nothing else caught it. The poll added last round to make "success without a
tunnel" impossible produced exactly that. Now requires a number greater than
zero, and waits 20s rather than 10 so a slow link is not rolled back needlessly.

**`down` still sat above the key registration.** Last round moved it below the
token and server-list fetches but not below `addKey`, which is the most
failure-prone of the three — one gateway, by CN, pinned certificate. So a
refused registration still tore down a working tunnel. It now runs after the
last fetch; the key is generated before but written after, since `down` deletes
it. SKILL.md said "after every network fetch has succeeded", which was false;
corrected.

**`up --full` made `host.docker.internal` unresolvable — and `status` said DNS
was fine.** That name is answered only by the resolver being replaced; it is not
in `/etc/hosts`. `gateway.rs` hands it to every container for the LiteLLM
gateway, and Ollama and custom endpoints default to it, so an agent running
`up --full` silently removed the project's model backend. The route was already
excluded; only the name was lost. Now resolved with the old resolver and pinned
into `/etc/hosts` before the swap, restored on teardown, and `status` probes it
— PIA answers public names happily, which is precisely why probing only
`api.anthropic.com` reported "ok". Documented as Trap 4.

**The rollback could abort halfway.** The trap's `{ … }` is not exempt from
`set -e`, and `down`'s `cat`/`tac`/`rm` had no `|| true` — so one failure left
the interface up with all traffic captured, after printing "rolling back".
`down` now runs under `set +e`, the trap tolerates its failure, and the
interface is deleted *first*, since that removes every route pointing at it.

**The account password had a real argv window.** curl does blank `-u`, but only
once running: sampling /proc/<pid>/cmdline caught the plaintext in 2 of 400
tries, between exec and the overwrite. Small, but it is the permanent password
and the token already had the fix. Moved onto the same stdin config — 0 of 400.
Review reported this as a 25-second exposure; that was a wrapper's argv, not
curl's.

entrypoint: the skill install stages into `$_dest.new` and swaps, so a failed
copy leaves the previous copy intact instead of a truncated SKILL.md and no
script, root-owned, on a persisted volume. Verified against a size-limited
filesystem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:15:52 -07:00
shadow-testandClaude Opus 5 5b96ad4823 Stop a failed up from tearing down a working tunnel
The previous commit moved `down` to the top of `up` to fix a resolv.conf
idempotency bug, and in doing so put it *before* every network fetch that can
fail. Re-review caught it and I reproduced it: with a tunnel up, an `up` that
fails on bad credentials left `pia0` gone and traffic silently back on the real
address, while the error talked only about credentials. A privacy regression
introduced by a correctness fix.

`down` now runs after the token, server list and key registration have all
succeeded — nothing above that line touches the network stack — and still
clears the stale backup it was added for.

Everything after it is covered by a rollback. Note this is an EXIT trap with a
flag, not `trap ... ERR`: my first attempt used ERR and did not fire at all,
because ERR is not inherited by shell functions without `set -E`, so a failure
inside add_route missed it, and `die` exits explicitly, which is not an error.
Verified by forcing a route collision — the tunnel is torn down and DNS is
intact, where before the fix it was left half-configured with DNS dead.

Also from the review:

- **The private key lived on disk for the whole life of the tunnel.**
  `commit_container_snapshot` blanks env vars, never files, and nothing tears
  the tunnel down before a recreate or migrate — so the `down`-time cleanup
  never covered the path that put a key in a snapshot in the first place. It is
  now deleted the moment `wg set` has read it; the kernel keeps its own copy,
  verified by checking the interface still works afterwards.
- **`up` claimed success without a handshake.** An unreachable peer still
  routes — into a black hole — so `up --full` could exit 0 having pointed all
  traffic and resolv.conf at a peer that never answered, with `status` printing
  "mode: full tunnel". Now polls for a handshake and rolls back if none arrives.
- **`status` needed root and did not check.** `wg show` fails unprivileged and
  was swallowed, so an unprivileged run printed "no tunnel up" and then "mode:
  full tunnel" in the same breath. An agent reading the first line would re-run
  `up` — which, before the fix above, destroyed the tunnel it failed to see.
- **The killswitch bullet was false.** It said `iptables` is not in the image;
  it is, so a killswitch is buildable. It stays unbuilt because it would cut
  Claude Code's own API traffic — an honest reason, unlike the previous one.
- `install_feature_skill` rejects path-traversal names, not just blank ones —
  verified `../skills` would have deleted the whole skills directory including
  Mission Control's — and reports `mkdir`/`cp` failures instead of printing a
  success line regardless.
- The usage text ended by printing `set -euo pipefail`, off by one line.
- HOW-TO-USE claimed "the container says so on start". entrypoint prints to
  PID 1's stdout, which no terminal or UI surfaces — `docker logs` appears
  nowhere in the repo. Now points at the migration pre-flight, which does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:11:50 -07:00
shadow-testandClaude Opus 5 dcb13d23ea Fix what review found in the skill: five real defects
Adversarial review of #29 found bugs I confirmed by reproducing each one.

**Every hand-written error message was unreachable.** `tok=$(curl ...)` is a
plain assignment, so `set -e` acts on the command substitution before the
following `|| die` can run. A wrong password produced exit 22 and no output at
all — the most likely way this gets used wrongly, and the least explained.
All four captures now go through a `run` helper that takes a *description*
rather than echoing the command, because one of them carries the account
password in `-u`.

**`up` was not idempotent, and the second run destroyed DNS.** The resolv.conf
backup was copied unconditionally, so `up --full` twice overwrote the good
backup with PIA's own resolvers; the later `down` then "restored" those and
left the container with no working DNS and no way back. `up` now runs `down`
first. Verified: two `up --full` runs, then `down`, and the backup still holds
the original 192.168.65.7.

**An empty gateway produced total connectivity loss, reported as healthy.**
`$gw` was never validated and `add_route` swallowed every failure to /dev/null.
The two half-routes need no gateway and would succeed, so the tunnel captured
everything while the exclusions keeping DNS and the Docker host reachable
silently did not exist — and `status` still printed "full tunnel". Routes are
now fatal on failure, and a via-less default (`$3` is the literal "eth0") is
rejected.

**The PIA session token was in the process arguments** — confirmed in `ps` and
/proc/*/cmdline, a ~24h bearer credential for the account readable by anything
in the container. It now goes to curl on stdin as a config. Verified: 60 polls
across a full `up`, zero sightings.

**The preflight diagnosed the wrong kernel module.** It checked /dev/net/tun
and blamed the tun module, but kernel WireGuard is a netlink interface and does
not use it — verified by creating one with NET_ADMIN and no tun device. The
check is dropped (the container could not have started without the device
anyway) and `ip link add` now reports the real dependency.

Also: a full tunnel with no DNS servers from PIA used to warn and carry on,
which is a tunnel leaking every lookup while reporting itself healthy — now
fatal. `down` validates the backup before restoring it, so a truncated one
cannot leave the container with no resolver at all. `wg.priv` is shredded on
teardown and created under umask 077, because /run rides `docker commit` into
the snapshot image. A mistyped `up --ful` is rejected instead of silently
giving a test route.

entrypoint: `install_feature_skill` gets `local`, a blank-name guard (the
disabled branch would otherwise `rm -rf` the whole skills directory under a
persisted volume), `-e`/`-L` so a leftover *file* at the destination is cleaned
up, and a chown of the parent so `claude` can still add skills of their own
when Mission Control is off. When the base image predates the skill it now says
so instead of returning silently — and `/opt/triple-c-skills` joins
FEATURE_PROBES so the migration pre-flight reports it. Docs corrected to match:
neither half reaches an existing project without a migration.

`vpn_env_var` extracted and tested, pinning the property the whole removal path
rests on — that the variable is emitted as 0 rather than omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:11:50 -07:00
shadow-testandClaude Opus 5 7a8bbcbef7 Report which exit is which, instead of one ambiguous "public IP"
`status` probed https://1.1.1.1/cdn-cgi/trace and printed the answer as
"public IP". In test mode 1.1.1.1 is the *only* address routed into the tunnel,
so that line reported a PIA exit while every other packet left directly — a
test tunnel reading exactly like a full one.

Found on a live container: default route still via eth0, one 1.1.1.1/32 route
through pia0, and the old status line claiming a PIA public IP. This is a
plausible route to concluding the VPN is on when it is not, which is close to
the confusion this skill exists to prevent.

Status now names the mode and, in test mode, prints both exits with the real
address called out. 1.0.0.1 serves the same trace endpoint as 1.1.1.1 and is
never routed into the tunnel, so the direct exit can be probed without DNS.

Verified against all four states: no tunnel, test mode on a live tunnel that
was already up, full tunnel, and after teardown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:11:23 -07:00
shadow-testandClaude Opus 5 3bd3caa101 Ship a pia-vpn skill with the VPN support toggle
The toggle grants CAP_NET_ADMIN and /dev/net/tun and stops there, which users
reasonably read as "turn the VPN on" — the gap between the two is the reported
bug that the default network does not route through a VPN. Close it by giving
the container an agent-usable way to build the tunnel, rather than leaving
each project to rediscover it.

container/skills/ is baked to /opt/triple-c-skills and installed into
~/.claude/skills/ by entrypoint.sh from VPN_SUPPORT_ENABLED, mirroring how
Mission Control installs its own. Staged under /opt because ~/.claude is a
volume mount that would mask an image copy from first start.

Three details that are not incidental:

- The variable is sent as 0 rather than omitted when off, because ~/.claude
  persists: entrypoint has to be *told* to remove a skill left by an earlier
  run with the toggle on, and an absent variable cannot say that. A stale skill
  is worse than none, since it instructs an agent to use a capability the
  container no longer has.
- It is reserved in RESERVED_ENV_EXACT alongside MISSION_CONTROL_ENABLED, or a
  custom env var of the same name could claim the skill without the capability
  behind it. Covered by a test.
- The skill is re-copied on every start, rm -rf'd first, so fixes reach existing
  projects and files dropped from a later version do not linger.

The skill itself carries the three things that are easy to get wrong: that a
full tunnel captures the Docker resolver and takes DNS down with it, that an
IP-literal health check cannot see a dead resolver, and that no tunnel survives
a restart while /run state riding the snapshot makes it look as though one did.
It also states what it deliberately does not do — no killswitch, no autostart —
so an agent proposes those as decisions rather than improvising them.

pia-wg.sh preflights CAP_NET_ADMIN by capability bit rather than letting the
first `ip` call fail with a bare EPERM that points nowhere near the setting
that needs changing. Credentials stay in a file (~/pia-creds, PIA_CREDS to
override) rather than the environment, where docker inspect and every process
in the container would see them.

Tested: install/refresh/remove/no-op paths of install_feature_skill against the
real function; preflight with and without the capability; and a full up --full
/ down round trip, confirming DNS via PIA's resolvers, api.anthropic.com
reachable through the exit, and routes and resolv.conf restored on teardown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:11:23 -07:00
shadow-testandClaude Opus 5 5dd1ab5217 Stop resting the iptables case on a kernel config I cannot verify
Build App (Preview) / compute-version (pull_request) Successful in 6s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build Container / build-container (pull_request) Successful in 29s
Build App (Preview) / build-macos (pull_request) Successful in 2m39s
Build App (Preview) / build-windows (pull_request) Successful in 5m46s
Build App (Preview) / build-linux (pull_request) Successful in 5m59s
Build App (Preview) / prune-previews (pull_request) Successful in 3s
Round 3 argued the macOS rationale is stale: that Docker Desktop no longer
builds from linuxkit/linuxkit and has enabled nft_fib_ipv4 since 4.35. I could
not confirm or refute that from a Linux host — searching turned up no version
matrix either way.

But the decision does not depend on it, and the comment should not have implied
it did. `xt_CONNMARK`, which the iptables path needs, was present in every
kernel config examined. `nft_fib_ipv4`, which the nft path needs, was absent
from the config read here and may be present in current Docker Desktop. That
asymmetry is the actual argument: nftables' viability varies by Docker Desktop
version in a way nobody here can pin down, iptables' requirement did not vary
anywhere it was checked. If nft_fib_ipv4 is present this costs 1.6 MB and
nothing else; if it is absent it is the difference between a working full
tunnel and none.

Rewritten to say that, and to say plainly what is verified versus assumed —
this is the third round in which the previous round's central premise did not
survive, and a confidently-worded paragraph is what the next round inherits.

Also from review:

- CLAUDE.md still said "`iptables` is deliberately absent", the opposite of what
  this PR now does, contradicting the Dockerfile and both other docs.
- The Dockerfile referenced "the pia-vpn skill", which does not exist on this
  branch — the third forward reference of that kind, now gone.
- "full tunnels work on native Linux, Mac and WSL2 6.6" was unconditional and
  contradicted ten lines later by the `DNS =` concession, which stops them on
  every platform. Reordered so the DNS hurdle is named as the first one.
- The WSL2 gap was written as a permanent platform limitation. It is a stale
  install: `wsl --update` moves the host to a current kernel that has the
  symbol. That remedy was missing from the user-facing doc.
- The migration probe label carried an internal comma, which `joinFeatures`
  renders into a comma-joined list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 17:11:17 -07:00
7 changed files with 135 additions and 61 deletions
+2 -1
View File
@@ -315,7 +315,8 @@ container is created once by a very long function where a dropped capability is
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
base-image migration — leaving a project holding the capability with nothing able to exercise it,
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
and no error that points at why. `iptables` is included and `nftables` deliberately is not; see
the Dockerfile comment for why that way round.
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
+5 -4
View File
@@ -548,10 +548,11 @@ Things worth knowing:
at `resolvconf: command not found` and deletes the interface again. This happens before any
routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive
`wg` and `ip route` directly rather than going through `wg-quick`.
- **`wg-quick` full tunnels also need `xt_CONNMARK` from the host kernel.** Native Linux, Docker
Desktop for Mac and WSL2 kernels from 6.6 have it; older WSL2 kernels do not, and a container
cannot load one. There the answer is again to add the routes yourself with `ip route`, which
needs no firewall backend on any platform.
- **`wg-quick` full tunnels additionally need `xt_CONNMARK` from the host kernel.** WSL2 kernels
before 6.6 do not have it and a container cannot load one — on Windows, `wsl --update` moves you
to a current kernel, which does. Failing that, add the routes yourself with `ip route`, which
needs no firewall backend on any platform. Note this is the *second* hurdle: clear the `DNS =`
one above first, or you will not reach this.
> This setting can only be changed when the container is stopped. Capabilities and devices are
> fixed when a container is created, so toggling it recreates the container on the next start.
+2 -2
View File
@@ -135,8 +135,8 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
("/opt/mission-control", "Mission Control (Flight Control)"),
("/usr/bin/wg", "VPN tooling (WireGuard, for the VPN Support toggle)"),
("/opt/triple-c-skills", "Bundled skills (PIA VPN, for the VPN Support toggle)"),
("/usr/bin/wg", "VPN tooling for the VPN Support toggle (WireGuard)"),
("/opt/triple-c-skills", "Bundled skills for the VPN Support toggle (PIA VPN)"),
];
/// Headroom demanded on Docker's storage backend on top of the measured
+17 -10
View File
@@ -87,18 +87,25 @@ RUN for i in 1 2 3 4 5; do \
# Error: Could not process rule: No such file or directory
# ^^^^^^^^^^^^^^ needs nft_fib_ipv4
#
# That matters because of how the two hosts we ship to are configured. From
# LinuxKit's kernel config — Docker Desktop for Mac, identical on both arches:
# The choice therefore turns on which kernel symbol each path needs, and the two
# are not equally safe to bet on. `xt_CONNMARK` (iptables) was present in every
# kernel config examined — LinuxKit's for both arches, and WSL2's from 6.6.
# `nft_fib_ipv4` (nftables) was absent from the LinuxKit config read here, and a
# later review argued Docker Desktop has since enabled it and no longer builds
# from that config at all. That may well be true; it could not be settled from a
# Linux host, and it is the point: nftables' viability varies by Docker Desktop
# version in a way nobody here can pin down, while iptables' requirement did not
# vary anywhere it was checked.
#
# CONFIG_NETFILTER_XT_CONNMARK=y <- the iptables path works
# # CONFIG_NFT_FIB_IPV4 is not set <- the nft path does not
# So `iptables` is chosen for being robust to that uncertainty rather than for
# beating nftables on any particular host. If nft_fib_ipv4 is present, wg-quick
# never reaches the iptables path and this costs 1.6 MB and nothing else; if it
# is absent, this is the difference between a working full tunnel and none.
#
# So shipping `nftables` would forfeit the platform it was meant to fix. With
# `iptables`, full tunnels work on native Linux, on Docker Desktop for Mac, and
# on WSL2 kernels from 6.6 (which added xt_CONNMARK as a module). Only WSL2
# older than that is left out, and nothing installable here changes it — the way
# out there is to add the routes with `ip route` instead of using `wg-quick`,
# which is what the pia-vpn skill does on every platform.
# The residual gap is WSL2 before 6.6, which has neither symbol. Nothing
# installable in the container changes that — but `wsl --update` does, and moves
# the host to a far newer kernel. Add the routes with `ip route` in the meantime;
# that needs no firewall backend on any platform.
#
# ## What this still does not fix
#
+11 -3
View File
@@ -381,10 +381,18 @@ install_feature_skill() {
# Not just $_dest: when Mission Control is off nothing else creates the
# parent, so root would own it and `claude` could not add a skill there.
chown claude:claude /home/claude/.claude/skills
# Stage then swap. Copying over the live path meant a failure (full
# volume, read-only mount) left a truncated SKILL.md and no script
# behind, root-owned, on a persisted volume — which Claude Code then
# discovers and loads.
rm -rf "$_dest.new"
cp -r "$_src" "$_dest.new" || {
rm -rf "$_dest.new"
echo "entrypoint: $_name skill install FAILED (copy from $_src); previous copy left intact"
return 1; }
chown -R claude:claude "$_dest.new"
rm -rf "$_dest"
cp -r "$_src" "$_dest" || {
echo "entrypoint: $_name skill install FAILED (copy from $_src)"; return 1; }
chown -R claude:claude "$_dest"
mv "$_dest.new" "$_dest"
echo "entrypoint: $_name skill installed to ~/.claude/skills/"
elif [ -e "$_dest" ] || [ -L "$_dest" ]; then
# -e/-L rather than -d: a leftover *file* at that path must go too.
+21 -6
View File
@@ -9,7 +9,7 @@ Bring this container's traffic out through Private Internet Access over
WireGuard, using the API PIA documents for headless use.
Run `sudo ~/.claude/skills/pia-vpn/pia-wg.sh` with `up`, `up --full`, `down` or
`status`. Read the rest of this page before the first `up --full`three of the
`status`. Read the rest of this page before the first `up --full`four of the
behaviours below are actively misleading if you meet them without warning, and
each one presents as "the VPN is fine" or "Claude is broken" rather than as
what it is.
@@ -107,7 +107,21 @@ mode: test route only (1.1.1.1 through the tunnel, nothing else)
Two different addresses there is correct and expected in test mode. If you want
the second line to change, you want `up --full`.
## Trap 4: no tunnel survives a restart, and it fails open
## Trap 4: a full tunnel hides the Docker host unless the name is pinned
`host.docker.internal` is answered *only* by the resolver that `up --full`
replaces — it is not in `/etc/hosts`. Triple-C hands that name to the container
for the LiteLLM gateway, and host-side Ollama and custom endpoints default to
it, so losing the name takes the project's model backend down with it.
The nasty part is what a naive check reports. PIA's resolvers answer public
names perfectly well, so a probe of `api.anthropic.com` says everything is fine
while the Docker host has vanished. `pia-wg.sh` pins the address into
`/etc/hosts` before swapping the resolver and restores the file on teardown, and
`status` probes both names — but if you ever rewrite `resolv.conf` by hand, this
is the one that will not announce itself.
## Trap 5: no tunnel survives a restart, and it fails open
The network namespace is rebuilt every time the container starts, and nothing
inside reconnects anything. After a stop/start, Reset or any config change that
@@ -186,10 +200,11 @@ no DNS at all), removes exactly the routes that were added, in reverse order,
and deletes the interface. It is safe to run when nothing is up. Confirm
afterwards that the public address is back to the container's own.
`up` calls it too, but only *after* every network fetch has succeeded, so a
failed `up` leaves an existing tunnel alone rather than tearing it down to
report a bad password. From that point on a rollback is armed: if any step of
the setup fails, the tunnel is torn down rather than left half-configured.
`up` calls it too, but only after the last network fetch — the key registration
— has succeeded, so a failed `up` leaves an existing tunnel alone rather than
tearing it down to report a bad password or an unreachable gateway. From that
point on a rollback is armed: if any step of the setup fails, the tunnel is torn
down rather than left half-configured.
The private key is deleted earlier still — the moment `wg set` has read it,
while the tunnel is being built. That is not housekeeping: `/run` is in the
+77 -35
View File
@@ -121,9 +121,15 @@ up() {
u=$(sed -n 1p "$CREDS"); p=$(sed -n 2p "$CREDS")
[ -n "$u" ] && [ -n "$p" ] || die "$CREDS needs two lines: username, then password"
tok=$(run "PIA rejected the credentials in $CREDS, or could not be reached" \
curl -sf -m 25 -u "$u:$p" \
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
# Via stdin, not `-u`. curl does blank the password in its own argv, but only
# once it is running: sampling /proc/<pid>/cmdline in a tight loop caught the
# plaintext in 3 of 200 tries, in the window between exec and the overwrite.
# Small, but this is the permanent account password, and the mechanism to
# avoid it entirely is already here for the token.
tok=$(printf -- '--user "%s:%s"\n' "$u" "$p" \
| run "PIA rejected the credentials in $CREDS, or could not be reached" \
curl -sf -m 25 -K - \
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA returned no token - check the credentials in $CREDS"
run "could not fetch PIA's server list" \
@@ -133,31 +139,9 @@ up() {
sip=$(echo "$srv" | jq -r .ip); scn=$(echo "$srv" | jq -r .cn)
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region '$REGION'"
# Only now tear down any previous tunnel. Doing it up front (as an earlier
# version did) meant a failed token fetch or an unreachable server list took
# a *working* tunnel down with it and silently reverted the container to its
# real address, while the error talked about credentials. Everything above
# this line can fail; nothing above it has touched the network stack.
#
# It also still does the job it was added for: clearing a stale resolv.conf
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
down >/dev/null 2>&1 || true
# From here on the network stack is being modified, so any failure has to put
# it back rather than exit half-configured. `down` is idempotent and restores
# routes and resolv.conf exactly.
#
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
# ERR trap is not inherited by shell functions without `set -E`, so a failure
# inside add_route would not fire it, and `die` exits explicitly, which is not
# an error and would not fire it either. EXIT catches both.
SETUP_OK=0
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1; }' EXIT
# umask, not a later chmod: the file is created under the inherited 0022
# otherwise, so the key is world-readable for the moment in between.
( umask 077; priv=$(wg genkey); printf '%s' "$priv" > wg.priv )
priv=$(cat wg.priv); pub=$(printf '%s' "$priv" | wg pubkey)
# The key is generated but NOT written yet -- `down` below deletes wg.priv, and
# the teardown has to come after every fetch that can fail.
priv=$(wg genkey); pub=$(printf '%s' "$priv" | wg pubkey)
# The token goes in on stdin as a curl config rather than in the argv, where
# `ps` and /proc/*/cmdline expose it to every process in the container --
@@ -170,6 +154,32 @@ up() {
--cacert ca.rsa.4096.crt "https://$scn:1337/addKey")
[ "$(echo "$resp" | jq -r .status)" = OK ] || die "key registration failed: $resp"
# Only now tear down any previous tunnel. Every network call above this line
# can fail, and an earlier version tore down first -- so a failed token fetch,
# an unreachable server list, or a refused key registration took a *working*
# tunnel with it and silently reverted the container to its real address while
# the error talked about credentials. Nothing above this line has touched the
# network stack. addKey is the most failure-prone of the three: it reaches one
# individual gateway by CN with a pinned certificate.
#
# It also still does the job it was added for: clearing a stale resolv.conf
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
down >/dev/null 2>&1 || true
# From here on the network stack is being modified, so any failure has to put
# it back rather than exit half-configured.
#
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
# ERR trap is not inherited by shell functions without `set -E`, so a failure
# inside add_route would not fire it, and `die` exits explicitly, which is not
# an error and would not fire it either. EXIT catches both.
SETUP_OK=0
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1 || true; }' EXIT
# umask, not a later chmod: created under the inherited 0022 otherwise, so the
# key would be world-readable for the moment in between.
( umask 077; printf '%s' "$priv" > wg.priv )
: > "$STATE/routes"
ip link add "$IFACE" type wireguard 2>/dev/null || \
die "could not create a WireGuard interface." \
@@ -216,6 +226,19 @@ up() {
# PIA's resolvers live inside 10/8, so pin them back through the tunnel with
# /32s -- longer still, so they beat the exclusion just added.
# `host.docker.internal` is answered only by the resolver about to be
# replaced -- it is not in /etc/hosts. Triple-C hands that name to the
# container for the LiteLLM gateway and defaults host-side Ollama and custom
# endpoints to it, so losing it takes the project's model backend with it.
# The *route* to it is already excluded above; only the name needs pinning.
# Resolve it with the old resolver and write it into /etc/hosts first.
local hdi
hdi=$(getent ahostsv4 host.docker.internal 2>/dev/null | awk '{print $1; exit}')
if [ -n "$hdi" ]; then
cp /etc/hosts "$STATE/hosts.bak"
printf '%s host.docker.internal\n' "$hdi" >> /etc/hosts
fi
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
for d in $dns; do add_route "$d/32" dev "$IFACE"; done
# resolv.conf is a bind mount: write through it, never replace it.
@@ -229,10 +252,16 @@ up() {
# A tunnel with no handshake still routes -- into a black hole. Without this
# `up --full` would exit 0 having pointed all traffic *and* resolv.conf at a
# peer that never answered, and `status` would print "mode: full tunnel".
local waited=0
until [ "$(wg show "$IFACE" latest-handshakes | awk '{print $2; exit}')" != 0 ]; do
# Demand a number, not just "different from 0". `wg show` prints nothing at
# all when the interface has no peer, and writes to stderr when the interface
# is gone -- both leave $2 empty, and `[ "" != 0 ]` is true, so the original
# form treated a missing tunnel as a completed handshake and exited 0. `until`
# suspends both `set -e` and `pipefail`, so nothing else was going to catch it.
local waited=0 hs
until hs=$(wg show "$IFACE" latest-handshakes 2>/dev/null | awk 'NR==1{print $2}')
[[ $hs =~ ^[0-9]+$ ]] && [ "$hs" -gt 0 ]; do
waited=$((waited + 1))
[ "$waited" -lt 20 ] || die "no handshake from $REGION after 10s - rolled back"
[ "$waited" -lt 40 ] || die "no handshake from $REGION after 20s"
sleep 0.5
done
@@ -243,6 +272,11 @@ up() {
down() {
[ "$(id -u)" = 0 ] || die "run with sudo"
# Teardown must finish even if a step fails; a half-rollback is the state this
# exists to prevent. Deliberately not inherited from the caller's `set -e`.
set +e
# First, because removing the interface removes every route that points at it.
ip link del "$IFACE" 2>/dev/null
# Only restore something that actually looks like a resolver file. Restoring
# an empty or truncated backup leaves the container with no DNS at all, which
# is worse than leaving the current one alone.
@@ -254,6 +288,10 @@ down() {
fi
rm -f "$STATE/resolv.conf.bak"
fi
if [ -f "$STATE/hosts.bak" ]; then
cat "$STATE/hosts.bak" > /etc/hosts
rm -f "$STATE/hosts.bak"
fi
if [ -f "$STATE/routes" ]; then
# Reverse order: the specific overrides go before the ranges they sit in.
tac "$STATE/routes" | while read -r r; do
@@ -261,7 +299,6 @@ down() {
done
rm -f "$STATE/routes"
fi
ip link del "$IFACE" 2>/dev/null || true
# /run is in the writable layer and `docker commit` bakes it into the
# project's snapshot image, so a key left here rides that image into every
# future container. Verified: a snapshot already carried one.
@@ -287,10 +324,15 @@ status() {
# Resolve a name, not an IP literal. A curl to 1.1.1.1 succeeds while DNS is
# completely broken, which is exactly how a dead resolver goes unnoticed.
printf 'DNS: '
if timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
echo "ok (via $(sed -n 's/^nameserver //p' /etc/resolv.conf | tr '\n' ' '))"
else
if ! timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
echo "BROKEN - cannot resolve api.anthropic.com"
elif ! timeout 10 getent hosts host.docker.internal >/dev/null 2>&1; then
# PIA's resolvers answer public names happily, so probing only
# api.anthropic.com reports "ok" on a container that has just lost the
# Docker host -- and with it the LiteLLM gateway and any host-side Ollama.
echo "public ok, but host.docker.internal is UNRESOLVABLE (gateway/Ollama backends will fail)"
else
echo "ok (via $(sed -n 's/^nameserver //p' /etc/resolv.conf | tr '\n' ' '))"
fi
# Report the exit per mode. In test mode the probe address is itself the one