Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
32149603b2 | ||
|
|
f6e5cf3f05 | ||
|
|
43c7ad1478 | ||
|
|
70c0a8bf7a |
@@ -315,8 +315,7 @@ container is created once by a very long function where a dropped capability is
|
||||
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
|
||||
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
|
||||
base-image migration — leaving a project holding the capability with nothing able to exercise it,
|
||||
and no error that points at why. `iptables` is included and `nftables` deliberately is not; see
|
||||
the Dockerfile comment for why that way round.
|
||||
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
|
||||
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
|
||||
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
|
||||
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
|
||||
|
||||
+4
-5
@@ -548,11 +548,10 @@ Things worth knowing:
|
||||
at `resolvconf: command not found` and deletes the interface again. This happens before any
|
||||
routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive
|
||||
`wg` and `ip route` directly rather than going through `wg-quick`.
|
||||
- **`wg-quick` full tunnels additionally need `xt_CONNMARK` from the host kernel.** WSL2 kernels
|
||||
before 6.6 do not have it and a container cannot load one — on Windows, `wsl --update` moves you
|
||||
to a current kernel, which does. Failing that, add the routes yourself with `ip route`, which
|
||||
needs no firewall backend on any platform. Note this is the *second* hurdle: clear the `DNS =`
|
||||
one above first, or you will not reach this.
|
||||
- **`wg-quick` full tunnels also need `xt_CONNMARK` from the host kernel.** Native Linux, Docker
|
||||
Desktop for Mac and WSL2 kernels from 6.6 have it; older WSL2 kernels do not, and a container
|
||||
cannot load one. There the answer is again to add the routes yourself with `ip route`, which
|
||||
needs no firewall backend on any platform.
|
||||
|
||||
> This setting can only be changed when the container is stopped. Capabilities and devices are
|
||||
> fixed when a container is created, so toggling it recreates the container on the next start.
|
||||
|
||||
@@ -135,8 +135,8 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
|
||||
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
|
||||
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
|
||||
("/opt/mission-control", "Mission Control (Flight Control)"),
|
||||
("/usr/bin/wg", "VPN tooling for the VPN Support toggle (WireGuard)"),
|
||||
("/opt/triple-c-skills", "Bundled skills for the VPN Support toggle (PIA VPN)"),
|
||||
("/usr/bin/wg", "VPN tooling (WireGuard, for the VPN Support toggle)"),
|
||||
("/opt/triple-c-skills", "Bundled skills (PIA VPN, for the VPN Support toggle)"),
|
||||
];
|
||||
|
||||
/// Headroom demanded on Docker's storage backend on top of the measured
|
||||
|
||||
+10
-17
@@ -87,25 +87,18 @@ RUN for i in 1 2 3 4 5; do \
|
||||
# Error: Could not process rule: No such file or directory
|
||||
# ^^^^^^^^^^^^^^ needs nft_fib_ipv4
|
||||
#
|
||||
# The choice therefore turns on which kernel symbol each path needs, and the two
|
||||
# are not equally safe to bet on. `xt_CONNMARK` (iptables) was present in every
|
||||
# kernel config examined — LinuxKit's for both arches, and WSL2's from 6.6.
|
||||
# `nft_fib_ipv4` (nftables) was absent from the LinuxKit config read here, and a
|
||||
# later review argued Docker Desktop has since enabled it and no longer builds
|
||||
# from that config at all. That may well be true; it could not be settled from a
|
||||
# Linux host, and it is the point: nftables' viability varies by Docker Desktop
|
||||
# version in a way nobody here can pin down, while iptables' requirement did not
|
||||
# vary anywhere it was checked.
|
||||
# That matters because of how the two hosts we ship to are configured. From
|
||||
# LinuxKit's kernel config — Docker Desktop for Mac, identical on both arches:
|
||||
#
|
||||
# So `iptables` is chosen for being robust to that uncertainty rather than for
|
||||
# beating nftables on any particular host. If nft_fib_ipv4 is present, wg-quick
|
||||
# never reaches the iptables path and this costs 1.6 MB and nothing else; if it
|
||||
# is absent, this is the difference between a working full tunnel and none.
|
||||
# CONFIG_NETFILTER_XT_CONNMARK=y <- the iptables path works
|
||||
# # CONFIG_NFT_FIB_IPV4 is not set <- the nft path does not
|
||||
#
|
||||
# The residual gap is WSL2 before 6.6, which has neither symbol. Nothing
|
||||
# installable in the container changes that — but `wsl --update` does, and moves
|
||||
# the host to a far newer kernel. Add the routes with `ip route` in the meantime;
|
||||
# that needs no firewall backend on any platform.
|
||||
# So shipping `nftables` would forfeit the platform it was meant to fix. With
|
||||
# `iptables`, full tunnels work on native Linux, on Docker Desktop for Mac, and
|
||||
# on WSL2 kernels from 6.6 (which added xt_CONNMARK as a module). Only WSL2
|
||||
# older than that is left out, and nothing installable here changes it — the way
|
||||
# out there is to add the routes with `ip route` instead of using `wg-quick`,
|
||||
# which is what the pia-vpn skill does on every platform.
|
||||
#
|
||||
# ## What this still does not fix
|
||||
#
|
||||
|
||||
+3
-11
@@ -381,18 +381,10 @@ install_feature_skill() {
|
||||
# Not just $_dest: when Mission Control is off nothing else creates the
|
||||
# parent, so root would own it and `claude` could not add a skill there.
|
||||
chown claude:claude /home/claude/.claude/skills
|
||||
# Stage then swap. Copying over the live path meant a failure (full
|
||||
# volume, read-only mount) left a truncated SKILL.md and no script
|
||||
# behind, root-owned, on a persisted volume — which Claude Code then
|
||||
# discovers and loads.
|
||||
rm -rf "$_dest.new"
|
||||
cp -r "$_src" "$_dest.new" || {
|
||||
rm -rf "$_dest.new"
|
||||
echo "entrypoint: $_name skill install FAILED (copy from $_src); previous copy left intact"
|
||||
return 1; }
|
||||
chown -R claude:claude "$_dest.new"
|
||||
rm -rf "$_dest"
|
||||
mv "$_dest.new" "$_dest"
|
||||
cp -r "$_src" "$_dest" || {
|
||||
echo "entrypoint: $_name skill install FAILED (copy from $_src)"; return 1; }
|
||||
chown -R claude:claude "$_dest"
|
||||
echo "entrypoint: $_name skill installed to ~/.claude/skills/"
|
||||
elif [ -e "$_dest" ] || [ -L "$_dest" ]; then
|
||||
# -e/-L rather than -d: a leftover *file* at that path must go too.
|
||||
|
||||
@@ -9,7 +9,7 @@ Bring this container's traffic out through Private Internet Access over
|
||||
WireGuard, using the API PIA documents for headless use.
|
||||
|
||||
Run `sudo ~/.claude/skills/pia-vpn/pia-wg.sh` with `up`, `up --full`, `down` or
|
||||
`status`. Read the rest of this page before the first `up --full` — four of the
|
||||
`status`. Read the rest of this page before the first `up --full` — three of the
|
||||
behaviours below are actively misleading if you meet them without warning, and
|
||||
each one presents as "the VPN is fine" or "Claude is broken" rather than as
|
||||
what it is.
|
||||
@@ -107,21 +107,7 @@ mode: test route only (1.1.1.1 through the tunnel, nothing else)
|
||||
Two different addresses there is correct and expected in test mode. If you want
|
||||
the second line to change, you want `up --full`.
|
||||
|
||||
## Trap 4: a full tunnel hides the Docker host unless the name is pinned
|
||||
|
||||
`host.docker.internal` is answered *only* by the resolver that `up --full`
|
||||
replaces — it is not in `/etc/hosts`. Triple-C hands that name to the container
|
||||
for the LiteLLM gateway, and host-side Ollama and custom endpoints default to
|
||||
it, so losing the name takes the project's model backend down with it.
|
||||
|
||||
The nasty part is what a naive check reports. PIA's resolvers answer public
|
||||
names perfectly well, so a probe of `api.anthropic.com` says everything is fine
|
||||
while the Docker host has vanished. `pia-wg.sh` pins the address into
|
||||
`/etc/hosts` before swapping the resolver and restores the file on teardown, and
|
||||
`status` probes both names — but if you ever rewrite `resolv.conf` by hand, this
|
||||
is the one that will not announce itself.
|
||||
|
||||
## Trap 5: no tunnel survives a restart, and it fails open
|
||||
## Trap 4: no tunnel survives a restart, and it fails open
|
||||
|
||||
The network namespace is rebuilt every time the container starts, and nothing
|
||||
inside reconnects anything. After a stop/start, Reset or any config change that
|
||||
@@ -200,11 +186,10 @@ no DNS at all), removes exactly the routes that were added, in reverse order,
|
||||
and deletes the interface. It is safe to run when nothing is up. Confirm
|
||||
afterwards that the public address is back to the container's own.
|
||||
|
||||
`up` calls it too, but only after the last network fetch — the key registration
|
||||
— has succeeded, so a failed `up` leaves an existing tunnel alone rather than
|
||||
tearing it down to report a bad password or an unreachable gateway. From that
|
||||
point on a rollback is armed: if any step of the setup fails, the tunnel is torn
|
||||
down rather than left half-configured.
|
||||
`up` calls it too, but only *after* every network fetch has succeeded, so a
|
||||
failed `up` leaves an existing tunnel alone rather than tearing it down to
|
||||
report a bad password. From that point on a rollback is armed: if any step of
|
||||
the setup fails, the tunnel is torn down rather than left half-configured.
|
||||
|
||||
The private key is deleted earlier still — the moment `wg set` has read it,
|
||||
while the tunnel is being built. That is not housekeeping: `/run` is in the
|
||||
|
||||
@@ -121,15 +121,9 @@ up() {
|
||||
u=$(sed -n 1p "$CREDS"); p=$(sed -n 2p "$CREDS")
|
||||
[ -n "$u" ] && [ -n "$p" ] || die "$CREDS needs two lines: username, then password"
|
||||
|
||||
# Via stdin, not `-u`. curl does blank the password in its own argv, but only
|
||||
# once it is running: sampling /proc/<pid>/cmdline in a tight loop caught the
|
||||
# plaintext in 3 of 200 tries, in the window between exec and the overwrite.
|
||||
# Small, but this is the permanent account password, and the mechanism to
|
||||
# avoid it entirely is already here for the token.
|
||||
tok=$(printf -- '--user "%s:%s"\n' "$u" "$p" \
|
||||
| run "PIA rejected the credentials in $CREDS, or could not be reached" \
|
||||
curl -sf -m 25 -K - \
|
||||
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
|
||||
tok=$(run "PIA rejected the credentials in $CREDS, or could not be reached" \
|
||||
curl -sf -m 25 -u "$u:$p" \
|
||||
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
|
||||
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA returned no token - check the credentials in $CREDS"
|
||||
|
||||
run "could not fetch PIA's server list" \
|
||||
@@ -139,9 +133,31 @@ up() {
|
||||
sip=$(echo "$srv" | jq -r .ip); scn=$(echo "$srv" | jq -r .cn)
|
||||
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region '$REGION'"
|
||||
|
||||
# The key is generated but NOT written yet -- `down` below deletes wg.priv, and
|
||||
# the teardown has to come after every fetch that can fail.
|
||||
priv=$(wg genkey); pub=$(printf '%s' "$priv" | wg pubkey)
|
||||
# Only now tear down any previous tunnel. Doing it up front (as an earlier
|
||||
# version did) meant a failed token fetch or an unreachable server list took
|
||||
# a *working* tunnel down with it and silently reverted the container to its
|
||||
# real address, while the error talked about credentials. Everything above
|
||||
# this line can fail; nothing above it has touched the network stack.
|
||||
#
|
||||
# It also still does the job it was added for: clearing a stale resolv.conf
|
||||
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
|
||||
down >/dev/null 2>&1 || true
|
||||
|
||||
# From here on the network stack is being modified, so any failure has to put
|
||||
# it back rather than exit half-configured. `down` is idempotent and restores
|
||||
# routes and resolv.conf exactly.
|
||||
#
|
||||
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
|
||||
# ERR trap is not inherited by shell functions without `set -E`, so a failure
|
||||
# inside add_route would not fire it, and `die` exits explicitly, which is not
|
||||
# an error and would not fire it either. EXIT catches both.
|
||||
SETUP_OK=0
|
||||
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1; }' EXIT
|
||||
|
||||
# umask, not a later chmod: the file is created under the inherited 0022
|
||||
# otherwise, so the key is world-readable for the moment in between.
|
||||
( umask 077; priv=$(wg genkey); printf '%s' "$priv" > wg.priv )
|
||||
priv=$(cat wg.priv); pub=$(printf '%s' "$priv" | wg pubkey)
|
||||
|
||||
# The token goes in on stdin as a curl config rather than in the argv, where
|
||||
# `ps` and /proc/*/cmdline expose it to every process in the container --
|
||||
@@ -154,32 +170,6 @@ up() {
|
||||
--cacert ca.rsa.4096.crt "https://$scn:1337/addKey")
|
||||
[ "$(echo "$resp" | jq -r .status)" = OK ] || die "key registration failed: $resp"
|
||||
|
||||
# Only now tear down any previous tunnel. Every network call above this line
|
||||
# can fail, and an earlier version tore down first -- so a failed token fetch,
|
||||
# an unreachable server list, or a refused key registration took a *working*
|
||||
# tunnel with it and silently reverted the container to its real address while
|
||||
# the error talked about credentials. Nothing above this line has touched the
|
||||
# network stack. addKey is the most failure-prone of the three: it reaches one
|
||||
# individual gateway by CN with a pinned certificate.
|
||||
#
|
||||
# It also still does the job it was added for: clearing a stale resolv.conf
|
||||
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
|
||||
down >/dev/null 2>&1 || true
|
||||
|
||||
# From here on the network stack is being modified, so any failure has to put
|
||||
# it back rather than exit half-configured.
|
||||
#
|
||||
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
|
||||
# ERR trap is not inherited by shell functions without `set -E`, so a failure
|
||||
# inside add_route would not fire it, and `die` exits explicitly, which is not
|
||||
# an error and would not fire it either. EXIT catches both.
|
||||
SETUP_OK=0
|
||||
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1 || true; }' EXIT
|
||||
|
||||
# umask, not a later chmod: created under the inherited 0022 otherwise, so the
|
||||
# key would be world-readable for the moment in between.
|
||||
( umask 077; printf '%s' "$priv" > wg.priv )
|
||||
|
||||
: > "$STATE/routes"
|
||||
ip link add "$IFACE" type wireguard 2>/dev/null || \
|
||||
die "could not create a WireGuard interface." \
|
||||
@@ -226,19 +216,6 @@ up() {
|
||||
|
||||
# PIA's resolvers live inside 10/8, so pin them back through the tunnel with
|
||||
# /32s -- longer still, so they beat the exclusion just added.
|
||||
# `host.docker.internal` is answered only by the resolver about to be
|
||||
# replaced -- it is not in /etc/hosts. Triple-C hands that name to the
|
||||
# container for the LiteLLM gateway and defaults host-side Ollama and custom
|
||||
# endpoints to it, so losing it takes the project's model backend with it.
|
||||
# The *route* to it is already excluded above; only the name needs pinning.
|
||||
# Resolve it with the old resolver and write it into /etc/hosts first.
|
||||
local hdi
|
||||
hdi=$(getent ahostsv4 host.docker.internal 2>/dev/null | awk '{print $1; exit}')
|
||||
if [ -n "$hdi" ]; then
|
||||
cp /etc/hosts "$STATE/hosts.bak"
|
||||
printf '%s host.docker.internal\n' "$hdi" >> /etc/hosts
|
||||
fi
|
||||
|
||||
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
|
||||
for d in $dns; do add_route "$d/32" dev "$IFACE"; done
|
||||
# resolv.conf is a bind mount: write through it, never replace it.
|
||||
@@ -252,16 +229,10 @@ up() {
|
||||
# A tunnel with no handshake still routes -- into a black hole. Without this
|
||||
# `up --full` would exit 0 having pointed all traffic *and* resolv.conf at a
|
||||
# peer that never answered, and `status` would print "mode: full tunnel".
|
||||
# Demand a number, not just "different from 0". `wg show` prints nothing at
|
||||
# all when the interface has no peer, and writes to stderr when the interface
|
||||
# is gone -- both leave $2 empty, and `[ "" != 0 ]` is true, so the original
|
||||
# form treated a missing tunnel as a completed handshake and exited 0. `until`
|
||||
# suspends both `set -e` and `pipefail`, so nothing else was going to catch it.
|
||||
local waited=0 hs
|
||||
until hs=$(wg show "$IFACE" latest-handshakes 2>/dev/null | awk 'NR==1{print $2}')
|
||||
[[ $hs =~ ^[0-9]+$ ]] && [ "$hs" -gt 0 ]; do
|
||||
local waited=0
|
||||
until [ "$(wg show "$IFACE" latest-handshakes | awk '{print $2; exit}')" != 0 ]; do
|
||||
waited=$((waited + 1))
|
||||
[ "$waited" -lt 40 ] || die "no handshake from $REGION after 20s"
|
||||
[ "$waited" -lt 20 ] || die "no handshake from $REGION after 10s - rolled back"
|
||||
sleep 0.5
|
||||
done
|
||||
|
||||
@@ -272,11 +243,6 @@ up() {
|
||||
|
||||
down() {
|
||||
[ "$(id -u)" = 0 ] || die "run with sudo"
|
||||
# Teardown must finish even if a step fails; a half-rollback is the state this
|
||||
# exists to prevent. Deliberately not inherited from the caller's `set -e`.
|
||||
set +e
|
||||
# First, because removing the interface removes every route that points at it.
|
||||
ip link del "$IFACE" 2>/dev/null
|
||||
# Only restore something that actually looks like a resolver file. Restoring
|
||||
# an empty or truncated backup leaves the container with no DNS at all, which
|
||||
# is worse than leaving the current one alone.
|
||||
@@ -288,10 +254,6 @@ down() {
|
||||
fi
|
||||
rm -f "$STATE/resolv.conf.bak"
|
||||
fi
|
||||
if [ -f "$STATE/hosts.bak" ]; then
|
||||
cat "$STATE/hosts.bak" > /etc/hosts
|
||||
rm -f "$STATE/hosts.bak"
|
||||
fi
|
||||
if [ -f "$STATE/routes" ]; then
|
||||
# Reverse order: the specific overrides go before the ranges they sit in.
|
||||
tac "$STATE/routes" | while read -r r; do
|
||||
@@ -299,6 +261,7 @@ down() {
|
||||
done
|
||||
rm -f "$STATE/routes"
|
||||
fi
|
||||
ip link del "$IFACE" 2>/dev/null || true
|
||||
# /run is in the writable layer and `docker commit` bakes it into the
|
||||
# project's snapshot image, so a key left here rides that image into every
|
||||
# future container. Verified: a snapshot already carried one.
|
||||
@@ -324,15 +287,10 @@ status() {
|
||||
# Resolve a name, not an IP literal. A curl to 1.1.1.1 succeeds while DNS is
|
||||
# completely broken, which is exactly how a dead resolver goes unnoticed.
|
||||
printf 'DNS: '
|
||||
if ! timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
|
||||
echo "BROKEN - cannot resolve api.anthropic.com"
|
||||
elif ! timeout 10 getent hosts host.docker.internal >/dev/null 2>&1; then
|
||||
# PIA's resolvers answer public names happily, so probing only
|
||||
# api.anthropic.com reports "ok" on a container that has just lost the
|
||||
# Docker host -- and with it the LiteLLM gateway and any host-side Ollama.
|
||||
echo "public ok, but host.docker.internal is UNRESOLVABLE (gateway/Ollama backends will fail)"
|
||||
else
|
||||
if timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
|
||||
echo "ok (via $(sed -n 's/^nameserver //p' /etc/resolv.conf | tr '\n' ' '))"
|
||||
else
|
||||
echo "BROKEN - cannot resolve api.anthropic.com"
|
||||
fi
|
||||
|
||||
# Report the exit per mode. In test mode the probe address is itself the one
|
||||
|
||||
Reference in New Issue
Block a user