Compare commits
4
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
49ac673045 | ||
|
|
cb848e110e | ||
|
|
87184a4be9 | ||
|
|
ab2c75d0b2 |
@@ -203,7 +203,8 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li
|
||||
### Container (`container/`)
|
||||
|
||||
- **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared
|
||||
libraries a browser links against (see below)
|
||||
libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled`
|
||||
toggle grants capability for (`iproute2`, `wireguard-tools`, `nftables`)
|
||||
- **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs
|
||||
`npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own
|
||||
dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium
|
||||
@@ -316,17 +317,38 @@ container is created once by a very long function where a dropped capability is
|
||||
base-image migration — leaving a project holding the capability with nothing able to exercise it,
|
||||
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
|
||||
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
|
||||
service manager runs inside, so a tunnel never survives stop/start or recreation while `/run`
|
||||
state persists through the snapshot and makes it look as though it did. Traffic silently reverts
|
||||
to the real address. Any future autostart or killswitch work starts here.
|
||||
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
|
||||
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
|
||||
writable layer, so on a stop/start it is simply the same container's files, and on a recreation
|
||||
`docker commit` has carried it into the snapshot. Traffic silently reverts to the real address.
|
||||
Any future autostart or killswitch work starts here.
|
||||
- **`/run` riding the snapshot means a VPN client's key material can end up in an image.** Verified:
|
||||
a fresh container off the whp snapshot already contained the `wg.priv` a previous tunnel left in
|
||||
`/run`. Anything writing key material there inherits the problem — the same `docker commit`
|
||||
hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks
|
||||
ephemeral and is not. A VPN client that does this should delete its key on teardown.
|
||||
- **`wg-quick` full tunnels need `xt_CONNMARK` from the host kernel**, which WSL2 does not have and
|
||||
a container cannot load; `Recommends: nftables | iptables` is also stripped by
|
||||
`--no-install-recommends`, so `nftables` is baked explicitly. See the Dockerfile comment — the
|
||||
short version is that shipping the backend fixes native Linux and Docker Desktop for Mac, nothing
|
||||
fixes Docker Desktop for Windows, and adding the routes directly with `ip route` sidesteps it on
|
||||
all three.
|
||||
- **The `pia-vpn` skill is installed *and removed* from `VPN_SUPPORT_ENABLED`.** `container/skills/`
|
||||
is baked to `/opt/triple-c-skills` and `install_feature_skill()` in `entrypoint.sh` copies it into
|
||||
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches existing projects, and
|
||||
`rm -rf`'d first, so files dropped from a later version do not linger. The removal branch matters
|
||||
as much as the install: `~/.claude` is a persisted volume, so a skill left behind after the toggle
|
||||
goes off would keep instructing an agent to use a capability the container no longer has. Which is
|
||||
also why the variable is sent as `0` rather than omitted, and why it is in `RESERVED_ENV_EXACT` —
|
||||
a custom env var of that name could otherwise claim the skill without the capability behind it.
|
||||
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches any project whose base
|
||||
image has the source, and `rm -rf`'d first, so files dropped from a later version do not linger.
|
||||
The removal branch matters as much as the install: `~/.claude` is a persisted volume, so a skill
|
||||
left behind after the toggle goes off would keep instructing an agent to use a capability the
|
||||
container no longer has. Which is also why the variable is sent as `0` rather than omitted (see
|
||||
`vpn_env_var`, tested), and why it is in `RESERVED_ENV_EXACT` — a custom env var of that name
|
||||
could otherwise claim the skill without the capability behind it.
|
||||
- **Both halves of that live in the base image, so neither reaches an existing project.** A
|
||||
recreation builds from the project's *own snapshot*, which has no `/opt/triple-c-skills` and no
|
||||
updated `entrypoint.sh`; only a migration or a Reset delivers them. The install path says so out
|
||||
loud rather than returning silently, and `/opt/triple-c-skills` is in `FEATURE_PROBES` so the
|
||||
migration pre-flight lists it as missing. Worth knowing before adding anything else behind an
|
||||
existing toggle: the label fingerprints *the setting*, not the set of things the setting drives,
|
||||
so a project already at `true` gets no recreation at all on upgrade.
|
||||
|
||||
### Container Lifecycle
|
||||
|
||||
|
||||
+36
-18
@@ -475,13 +475,18 @@ When enabled, the host Docker socket is mounted into the container so Claude Cod
|
||||
|
||||
When enabled, the container is given the three things a VPN client needs to build a tunnel:
|
||||
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
|
||||
sysctl that WireGuard requires. The `ip` and `wg` commands are always present to use them. This is
|
||||
**off by default**.
|
||||
sysctl that WireGuard requires. This is **off by default**.
|
||||
|
||||
The `ip`, `wg` and `nft` commands ship in the container image so there is something able to use
|
||||
them. If your project's container was created from an older base image it will not have them, and
|
||||
`wg` will simply not be found — **migrating the project onto the current base image** is what picks
|
||||
them up. Installing them by hand with `sudo apt install wireguard-tools` works in the meantime, but
|
||||
lives in the writable layer, so it is undone by a **Reset** and by a migration.
|
||||
|
||||
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
|
||||
redirected, and no client is installed or started on your behalf. Enabling it and expecting the
|
||||
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
|
||||
container's traffic to start leaving through a VPN is the most common misreading of what it does —
|
||||
installing a client and routing traffic into it remains yours to do.
|
||||
configuring a tunnel and routing traffic into it remains yours to do.
|
||||
|
||||
To make that second half easier, enabling this also installs a **`pia-vpn` skill** into the
|
||||
container's `~/.claude/skills/`, so Claude Code can bring up a Private Internet Access tunnel over
|
||||
@@ -490,6 +495,10 @@ easy to get wrong (see the DNS note below), and it is removed again when you tur
|
||||
It needs your PIA credentials in `~/pia-creds`, two lines, username then password. If you use a
|
||||
different provider, ignore it and set up your own client; nothing else depends on it.
|
||||
|
||||
Like the VPN tooling above, the skill ships in the container image, so a project whose container
|
||||
predates it will not get one by toggling the setting — **migrate the project** and it appears. The
|
||||
container says so on start when that is the case, rather than leaving you to wonder where it went.
|
||||
|
||||
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
|
||||
the connection attempt **hangs until it times out** — a default container has no tun device to open
|
||||
and no permission to add an interface or a route, and most clients report that as a generic timeout
|
||||
@@ -511,20 +520,29 @@ Things worth knowing:
|
||||
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
|
||||
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
|
||||
- **No tunnel survives a restart.** The network namespace is built fresh every time the container
|
||||
starts, and there is no service manager inside to reconnect anything. Files under `/run` may
|
||||
persist via the snapshot and make it *look* like the tunnel is still configured, but after any
|
||||
stop/start, Reset or recreation the interface and its routes are gone and traffic goes out your
|
||||
real address again — with no error and nothing visibly different. Re-establish it after every
|
||||
start, and check rather than assume.
|
||||
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Containers
|
||||
resolve through an address on the Docker network (`192.168.65.7` under Docker Desktop) that sits
|
||||
outside the container's own subnet, so a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus
|
||||
`128.0.0.0/1` pair — captures it and sends every lookup into a tunnel that cannot carry it. The
|
||||
symptom is total: Claude Code reports it cannot connect, because it cannot resolve
|
||||
`api.anthropic.com`. Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16`
|
||||
via the original gateway, and use the VPN provider's own resolver for everything else. Note that
|
||||
a health check which fetches an IP literal such as `1.1.1.1` passes cleanly while this is broken —
|
||||
resolve a name instead.
|
||||
starts, and there is no service manager inside to reconnect anything. Leftover state under `/run`
|
||||
makes it *look* like the tunnel is still configured — that directory is in the container's
|
||||
writable layer, so it is simply still there after a stop/start, and `docker commit` carries it
|
||||
into the snapshot that a recreation is built from. Either way the interface and its routes are
|
||||
gone and traffic goes out your real address again, with no error and nothing visibly different.
|
||||
Re-establish it after every start, and check rather than assume.
|
||||
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your
|
||||
resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own
|
||||
subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures
|
||||
it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is
|
||||
`192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`,
|
||||
which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is
|
||||
total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`.
|
||||
Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original
|
||||
gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own,
|
||||
or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own
|
||||
address via the original gateway, or the tunnel's encrypted packets try to route through the
|
||||
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
|
||||
while DNS is broken — resolve a name instead.
|
||||
- **`wg-quick` cannot bring up a full tunnel on Docker Desktop for Windows.** Its `Table=auto` mode
|
||||
routes by firewall mark and needs `xt_CONNMARK` from the host kernel, which WSL2's does not have
|
||||
and a container cannot load. Split tunnels (a specific `AllowedIPs`) work fine, as does adding
|
||||
the routes yourself with `ip route`. Native Linux and Docker Desktop for Mac are unaffected.
|
||||
|
||||
> This setting can only be changed when the container is stopped. Capabilities and devices are
|
||||
> fixed when a container is created, so toggling it recreates the container on the next start.
|
||||
|
||||
@@ -841,6 +841,23 @@ type VpnHostConfigParts = (
|
||||
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
|
||||
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
|
||||
/// Hence opt-in, per project, rather than on for everyone.
|
||||
/// The env var `entrypoint.sh` installs and removes the `pia-vpn` skill from.
|
||||
///
|
||||
/// **Emitted either way, never omitted.** `~/.claude` is a persisted volume, so
|
||||
/// turning the toggle off has to actively tell entrypoint to remove a skill an
|
||||
/// earlier run left there, and an absent variable cannot say that. It is also
|
||||
/// what stops a `=1` baked into a snapshot by `docker commit` from outliving
|
||||
/// the setting — the explicit `=0` overwrites it.
|
||||
///
|
||||
/// Extracted for the same reason as [`vpn_host_config`]: the emitting code sits
|
||||
/// in a very long function where a dropped or inverted value is invisible, and
|
||||
/// `MISSION_CONTROL_ENABLED` twenty lines above shows the failure this avoids —
|
||||
/// it is pushed only when true, so a snapshot's baked `=1` survives the toggle
|
||||
/// going off.
|
||||
fn vpn_env_var(enabled: bool) -> String {
|
||||
format!("VPN_SUPPORT_ENABLED={}", u8::from(enabled))
|
||||
}
|
||||
|
||||
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
|
||||
if !enabled {
|
||||
return (None, None, None);
|
||||
@@ -1276,14 +1293,7 @@ pub async fn create_container(
|
||||
env_vars.push("MISSION_CONTROL_ENABLED=1".to_string());
|
||||
}
|
||||
|
||||
// Drives the pia-vpn skill install in entrypoint.sh. Sent as 0 rather than
|
||||
// omitted when off, because ~/.claude is a persisted volume: entrypoint has
|
||||
// to be told to *remove* a skill left there by an earlier run with the
|
||||
// toggle on, and an absent variable cannot say that.
|
||||
env_vars.push(format!(
|
||||
"VPN_SUPPORT_ENABLED={}",
|
||||
u8::from(project.vpn_support_enabled)
|
||||
));
|
||||
env_vars.push(vpn_env_var(project.vpn_support_enabled));
|
||||
|
||||
// Permission mode — read by triple-c-task-runner for scheduled (headless)
|
||||
// Claude Code runs. Interactive terminals get the flags directly instead.
|
||||
@@ -2799,6 +2809,17 @@ mod tests {
|
||||
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_vpn_skill_flag_is_emitted_either_way_never_omitted() {
|
||||
// The whole removal path depends on this. If `false` ever became "emit
|
||||
// nothing", a project that had the toggle on would keep the skill
|
||||
// forever: the container recreates from a snapshot whose baked
|
||||
// VPN_SUPPORT_ENABLED=1 would then go unchallenged, and entrypoint
|
||||
// would reinstall a skill for a capability the container no longer has.
|
||||
assert_eq!(vpn_env_var(true), "VPN_SUPPORT_ENABLED=1");
|
||||
assert_eq!(vpn_env_var(false), "VPN_SUPPORT_ENABLED=0");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_vpn_skill_flag_is_reserved_from_custom_env() {
|
||||
// entrypoint.sh installs and removes the pia-vpn skill from this
|
||||
|
||||
@@ -135,6 +135,8 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
|
||||
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
|
||||
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
|
||||
("/opt/mission-control", "Mission Control (Flight Control)"),
|
||||
("/usr/bin/wg", "VPN support (WireGuard tools)"),
|
||||
("/opt/triple-c-skills", "Feature skills (PIA VPN)"),
|
||||
];
|
||||
|
||||
/// Headroom demanded on Docker's storage backend on top of the measured
|
||||
|
||||
+46
-10
@@ -36,6 +36,7 @@ RUN for i in 1 2 3 4 5; do \
|
||||
socat \
|
||||
iproute2 \
|
||||
wireguard-tools \
|
||||
nftables \
|
||||
&& rm -rf /var/lib/apt/lists/*
|
||||
|
||||
# `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither
|
||||
@@ -44,22 +45,55 @@ RUN for i in 1 2 3 4 5; do \
|
||||
# corporate CA, no matter what the system trust store says. entrypoint.sh
|
||||
# degrades to a warning if it is ever missing.
|
||||
|
||||
# `iproute2` and `wireguard-tools` above are what the VPN support toggle
|
||||
# (`vpn_support_enabled`) grants capability *for*. That toggle hands a project
|
||||
# CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to add a
|
||||
# route, and without `wg` no way to build the tunnel those two exist to serve —
|
||||
# a capability with nothing able to use it.
|
||||
# `iproute2`, `wireguard-tools` and `nftables` above are what the VPN support
|
||||
# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a
|
||||
# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to
|
||||
# add a route, and without `wg` no way to build the tunnel those two exist to
|
||||
# serve — a capability with nothing able to use it.
|
||||
#
|
||||
# They are baked rather than left to a runtime `apt-get install` for the same
|
||||
# reason as the Playwright libraries below: the writable layer is re-paid after
|
||||
# every Reset and lost on base-image migration. A hand-installed `wg` therefore
|
||||
# works right up until an upgrade, then disappears and takes the tunnel with it
|
||||
# — silently, since a VPN that fails to come up looks exactly like one that was
|
||||
# never started. Together they are ~4.3 MB including dependencies.
|
||||
# never started.
|
||||
#
|
||||
# `iptables` is deliberately NOT here. The only thing that wants it is a desktop
|
||||
# VPN client's killswitch, and those clients need a GUI that a container has no
|
||||
# way to give them; leaving it out keeps the reach of CAP_NET_ADMIN smaller.
|
||||
# Measured against the *current base image*, not a bare ubuntu:24.04 — the base
|
||||
# already ships libelf1t64, so measuring on bare ubuntu over-counts by ~209 kB:
|
||||
# +9 packages, 5,614 kB on amd64 (4,153 kB of that is iproute2+wireguard-tools,
|
||||
# 1,461 kB is nftables). The same set on arm64 is 7,422 kB, measured against
|
||||
# ubuntu:24.04 since the arm64 base is not cached here.
|
||||
#
|
||||
# ## Why `nftables` specifically
|
||||
#
|
||||
# `wireguard-tools` declares `Recommends: nftables | iptables`, which the
|
||||
# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s
|
||||
# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e.
|
||||
# every stock full-tunnel config every provider hands out — and it shells out to
|
||||
# a firewall backend with no `type -p` guard. Measured without one:
|
||||
#
|
||||
# [#] iptables-restore -n
|
||||
# /usr/bin/wg-quick: line 32: iptables-restore: command not found
|
||||
# wg-quick EXIT=127 (interface rolled back, split tunnels unaffected)
|
||||
#
|
||||
# `nftables` rather than `iptables` because `wg-quick` prefers it (`if type -p
|
||||
# nft`, so with both installed iptables is dead weight), it is the first
|
||||
# alternative in the package's own Recommends, and it is roughly half the size.
|
||||
#
|
||||
# This does NOT make `wg-quick`'s full-tunnel mode work everywhere. `Table=auto`
|
||||
# routes by fwmark and needs connection-mark tracking from the *host* kernel:
|
||||
#
|
||||
# Warning: Extension CONNMARK revision 0 not supported, missing kernel module?
|
||||
#
|
||||
# WSL2's kernel has no `xt_CONNMARK` and containers have no /lib/modules to load
|
||||
# one from, so on Docker Desktop for Windows `wg-quick up` on a full tunnel fails
|
||||
# regardless of what is installed here. Native Linux and Docker Desktop for Mac
|
||||
# have it. Shipping the backend is what makes the difference on those two;
|
||||
# nothing shipped here can make the difference on WSL2, where the way out is to
|
||||
# add the routes with `ip route` instead of going through `wg-quick` at all.
|
||||
#
|
||||
# `iptables` is deliberately still NOT here: with `nftables` present `wg-quick`
|
||||
# never reaches for it, so it would add size and firewall surface for nothing.
|
||||
|
||||
# Remove default ubuntu user to free UID 1000 for host-user remapping
|
||||
RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \
|
||||
@@ -340,7 +374,9 @@ COPY mission-control /opt/mission-control
|
||||
# no skill at all. Staged in /opt because ~/.claude is a volume mount: an image
|
||||
# copy underneath it would be masked from the project's first start onward.
|
||||
COPY skills /opt/triple-c-skills
|
||||
RUN chmod +x /opt/triple-c-skills/*/*.sh
|
||||
# `find`, not a `*/*.sh` glob: the glob fails the build the day a skill ships
|
||||
# without a script, which is a legitimate thing for a skill to do.
|
||||
RUN find /opt/triple-c-skills -name '*.sh' -exec chmod +x {} +
|
||||
|
||||
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
|
||||
RUN chmod +x /usr/local/bin/entrypoint.sh
|
||||
|
||||
+25
-6
@@ -347,18 +347,37 @@ fi
|
||||
# Copied on every start rather than only when absent, so a fix to a skill
|
||||
# reaches projects that already have the old copy. Local edits under these
|
||||
# directories do not survive — treat /opt/triple-c-skills as the source.
|
||||
#
|
||||
# The source lives in the *base image*, so a project whose container predates it
|
||||
# recreates from its own snapshot and has no /opt/triple-c-skills to copy from.
|
||||
# That case says so rather than returning silently: the toggle is on, the
|
||||
# capability is there, and the skill simply never appears — which is impossible
|
||||
# to work out from the outside.
|
||||
install_feature_skill() {
|
||||
_name=$1
|
||||
_enabled=$2
|
||||
_dest="/home/claude/.claude/skills/$_name"
|
||||
local _name="$1"
|
||||
local _enabled="$2"
|
||||
local _src="/opt/triple-c-skills/$1"
|
||||
local _dest="/home/claude/.claude/skills/$1"
|
||||
|
||||
# A blank name would make the disabled branch `rm -rf` the whole skills
|
||||
# directory, Mission Control's included, under a persisted volume.
|
||||
[ -n "$_name" ] || { echo "entrypoint: install_feature_skill called with no name"; return 1; }
|
||||
|
||||
if [ "$_enabled" = "1" ]; then
|
||||
[ -d "/opt/triple-c-skills/$_name" ] || return 0
|
||||
if [ ! -d "$_src" ]; then
|
||||
echo "entrypoint: $_name skill unavailable — this container's base image predates it; migrate the project to get it"
|
||||
return 0
|
||||
fi
|
||||
mkdir -p /home/claude/.claude/skills
|
||||
# Not just $_dest: when Mission Control is off nothing else creates the
|
||||
# parent, so root would own it and `claude` could not add a skill there.
|
||||
chown claude:claude /home/claude/.claude/skills
|
||||
rm -rf "$_dest"
|
||||
cp -r "/opt/triple-c-skills/$_name" "$_dest"
|
||||
cp -r "$_src" "$_dest"
|
||||
chown -R claude:claude "$_dest"
|
||||
echo "entrypoint: $_name skill installed to ~/.claude/skills/"
|
||||
elif [ -d "$_dest" ]; then
|
||||
elif [ -e "$_dest" ] || [ -L "$_dest" ]; then
|
||||
# -e/-L rather than -d: a leftover *file* at that path must go too.
|
||||
rm -rf "$_dest"
|
||||
echo "entrypoint: $_name skill removed (feature disabled)"
|
||||
fi
|
||||
|
||||
@@ -65,6 +65,19 @@ PIA's own resolvers through the tunnel with `/32` routes that outrank the
|
||||
`10/8` exclusion. If you ever route traffic by hand, you owe both halves — the
|
||||
exclusions *and* a resolver reachable from wherever you pointed the default.
|
||||
|
||||
The failure has a quiet twin. Do only the first half — exclude the private
|
||||
ranges, leave the resolver alone — and everything *works*, while every DNS
|
||||
query travels outside the tunnel to your ISP. A VPN that leaks the full list of
|
||||
what you looked up is worse than one that is visibly broken, so `up --full`
|
||||
refuses to proceed if PIA does not hand back resolvers rather than carrying on
|
||||
without them.
|
||||
|
||||
The mechanism above is Docker Desktop's. On a user-defined Docker network the
|
||||
resolver is `127.0.0.11`, which is loopback and never captured by a default
|
||||
route — the trap still exists there (that resolver forwards upstream from
|
||||
inside the container's namespace) but arrives by a different path. Check
|
||||
`/etc/resolv.conf` rather than assuming which case you are in.
|
||||
|
||||
## Trap 2: an IP-literal health check cannot see a dead resolver
|
||||
|
||||
`curl https://1.1.1.1/cdn-cgi/trace` needs no DNS, so it returns a cheerful
|
||||
@@ -116,9 +129,20 @@ p1234567
|
||||
your-password
|
||||
```
|
||||
|
||||
Set `PIA_CREDS` to use a different path. Treat the contents as secret: never
|
||||
print the file, never echo the values, and never include them in a commit, a
|
||||
log or a message. The script reads it directly and does not echo it.
|
||||
Treat the contents as secret: never print the file, never echo the values, and
|
||||
never include them in a commit, a log or a message. The script reads it directly
|
||||
and does not echo it, and passes PIA's session token to `curl` on stdin rather
|
||||
than in the argv, where `ps` would expose it to everything in the container.
|
||||
|
||||
`PIA_CREDS` points somewhere else — but `sudo` resets the environment, so it
|
||||
only takes effect **after** the word `sudo`:
|
||||
|
||||
```bash
|
||||
sudo PIA_CREDS=/path/to/creds ~/.claude/skills/pia-vpn/pia-wg.sh up # works
|
||||
PIA_CREDS=/path/to/creds sudo ~/.claude/skills/pia-vpn/pia-wg.sh up # ignored
|
||||
```
|
||||
|
||||
The second form fails silently back to the default path. Same for `PIA_REGION`.
|
||||
|
||||
## Regions
|
||||
|
||||
@@ -156,10 +180,16 @@ true in test mode too, and means much less than it sounds like.
|
||||
|
||||
## Tearing down
|
||||
|
||||
`down` restores `resolv.conf` from its backup and removes exactly the routes
|
||||
that were added, in reverse order, then deletes the interface. It is safe to
|
||||
run when nothing is up. Confirm afterwards that the public address is back to
|
||||
the container's own.
|
||||
`down` restores `resolv.conf` from its backup (only if that backup still looks
|
||||
like a resolver file — restoring a truncated one would leave the container with
|
||||
no DNS at all), removes exactly the routes that were added, in reverse order,
|
||||
deletes the interface, and shreds the WireGuard private key. It is safe to run
|
||||
when nothing is up, and `up` runs it first so a repeat `up` cannot stack state.
|
||||
Confirm afterwards that the public address is back to the container's own.
|
||||
|
||||
The key deletion is not housekeeping: `/run` is in the container's writable
|
||||
layer, so `docker commit` bakes whatever is there into the project's snapshot
|
||||
image. A key left behind rides that image into every future container.
|
||||
|
||||
## What this deliberately does not do
|
||||
|
||||
@@ -173,3 +203,11 @@ the container's own.
|
||||
cannot work headless: the daemon never accepts a client connection without
|
||||
the GUI, and `piactl --help` states that connecting requires it. If you find
|
||||
one installed, it is not a working alternative to this script.
|
||||
- **Not `wg-quick`.** Its `Table=auto` full-tunnel mode routes by firewall mark
|
||||
and needs `xt_CONNMARK` from the host kernel, which Docker Desktop for
|
||||
Windows (WSL2) does not have and a container cannot load. This script adds
|
||||
the routes with `ip route` directly, which works on every host.
|
||||
- **IPv4 only.** The `0.0.0.0/1` + `128.0.0.0/1` pair covers v4. A container
|
||||
with a global IPv6 address and a v6 default route would leak all v6 traffic
|
||||
outside the tunnel; Triple-C's containers do not have one by default, but
|
||||
check `ip -6 route show default` before relying on this where it matters.
|
||||
|
||||
@@ -13,13 +13,20 @@
|
||||
#
|
||||
# Requires the project's "VPN support" setting (Config -> Runtime) to be on.
|
||||
#
|
||||
# Settings are read from the environment, but note that sudo resets it: they
|
||||
# have to be passed *through* sudo, after the word `sudo`, not before it.
|
||||
#
|
||||
# sudo PIA_REGION=uk_london pia-wg.sh up --full # works
|
||||
# PIA_REGION=uk_london sudo pia-wg.sh up --full # silently ignored
|
||||
#
|
||||
# PIA_CREDS credentials file, two lines: username, then password
|
||||
# (default ~/pia-creds; never echoed by this script)
|
||||
# (default /home/claude/pia-creds; never echoed by this script)
|
||||
# PIA_REGION region id (default us_chicago). List them with:
|
||||
# curl -s https://serverlist.piaservers.net/vpninfo/servers/v6 \
|
||||
# | head -1 | jq -r '.regions[].id'
|
||||
set -euo pipefail
|
||||
|
||||
# Not ~/pia-creds: under sudo, HOME is /root.
|
||||
CREDS=${PIA_CREDS:-/home/claude/pia-creds}
|
||||
REGION=${PIA_REGION:-us_chicago}
|
||||
IFACE=pia0
|
||||
@@ -36,11 +43,28 @@ PRIVATE_NETS="10.0.0.0/8 172.16.0.0/12 192.168.0.0/16 169.254.0.0/16"
|
||||
# source lines without the indentation ending up in the output.
|
||||
die() { echo "pia-wg: $*" >&2; exit 1; }
|
||||
|
||||
# `x=$(cmd)` is a plain assignment, so `set -e` kills the script on a non-zero
|
||||
# cmd *before* any `[ -z "$x" ] || die` line can run. Every capture below
|
||||
# therefore goes through `run`; without it a wrong password exits 22 with no
|
||||
# output at all, which is the most likely way this is used wrongly and was the
|
||||
# least explained.
|
||||
#
|
||||
# It takes a description rather than reporting the command it ran: one of these
|
||||
# invocations carries the account password in `-u`, and an error message is
|
||||
# exactly the wrong place for that to surface.
|
||||
run() { local what=$1; shift; "$@" || die "$what (exit $?)"; }
|
||||
|
||||
preflight() {
|
||||
[ "$(id -u)" = 0 ] || die "run with sudo"
|
||||
# CAP_NET_ADMIN is bit 12. Checking it by name gives a usable error; without
|
||||
# it the first `ip` call fails with a bare "Operation not permitted" that
|
||||
# points nowhere near the setting that actually needs changing.
|
||||
#
|
||||
# Deliberately NOT checking /dev/net/tun: kernel WireGuard is a netlink
|
||||
# interface and does not use it (verified -- `ip link add type wireguard`
|
||||
# succeeds with NET_ADMIN and no tun device). It is OpenVPN and userspace
|
||||
# wireguard-go that need it. The real kernel dependency here is the
|
||||
# `wireguard` module, which `ip link add` below reports on directly.
|
||||
local caps
|
||||
caps=$(awk '/^CapEff:/{print $2}' /proc/self/status)
|
||||
if [ $(( 0x$caps & 0x1000 )) -eq 0 ]; then
|
||||
@@ -49,51 +73,88 @@ preflight() {
|
||||
"again. That recreates the container; the home and .claude volumes" \
|
||||
"are preserved, so nothing in them is lost."
|
||||
fi
|
||||
[ -e /dev/net/tun ] || \
|
||||
die "/dev/net/tun is missing." \
|
||||
"Same fix: turn on \"VPN support\" in Config -> Runtime. If it is" \
|
||||
"already on, the Docker host's kernel is missing the tun module."
|
||||
command -v wg >/dev/null || die "wireguard-tools is not installed."
|
||||
command -v wg >/dev/null || \
|
||||
die "wireguard-tools is not installed." \
|
||||
"If this project's container was built from an older base image," \
|
||||
"migrate it onto the current one -- that is what ships \`wg\`."
|
||||
[ -r "$CREDS" ] || \
|
||||
die "no credentials at $CREDS." \
|
||||
"Two lines are expected: username, then password." \
|
||||
"Set PIA_CREDS to read them from somewhere else."
|
||||
"Set PIA_CREDS (after the word \`sudo\`) to read them elsewhere."
|
||||
}
|
||||
|
||||
# Record every route we add so teardown removes exactly those and nothing else.
|
||||
add_route() { ip route add $1 2>/dev/null && echo "$1" >> "$STATE/routes" || true; }
|
||||
# Routes that must work. A silent failure here is the worst state this script
|
||||
# can reach: the two half-routes need no gateway and would succeed, so the
|
||||
# tunnel captures everything while the exclusions that keep DNS and the Docker
|
||||
# host reachable are quietly missing -- and `status` still says "full tunnel".
|
||||
add_route() {
|
||||
ip route add "$@" || die "could not add route '$*'. Run 'down' to undo the partial setup."
|
||||
printf '%s\n' "$*" >> "$STATE/routes"
|
||||
}
|
||||
|
||||
up() {
|
||||
case "${1:-}" in
|
||||
""|--full) ;;
|
||||
*) die "unknown option '$1' (expected --full or nothing)." \
|
||||
"Refusing rather than silently giving you a test route." ;;
|
||||
esac
|
||||
preflight
|
||||
|
||||
# Always start from a known state. Without this a second `up` overwrites the
|
||||
# saved resolv.conf with PIA's own resolvers, so the later `down` "restores"
|
||||
# those and leaves the container with no working DNS and no way back.
|
||||
down >/dev/null 2>&1 || true
|
||||
|
||||
mkdir -p "$STATE"; cd "$STATE"
|
||||
[ -f ca.rsa.4096.crt ] || curl -sf -m 20 -o ca.rsa.4096.crt \
|
||||
https://raw.githubusercontent.com/pia-foss/manual-connections/master/ca.rsa.4096.crt \
|
||||
|| die "could not fetch PIA's CA certificate"
|
||||
|
||||
# `curl -o` creates the file before it knows the request failed, so a plain
|
||||
# `[ -f ]` cache check can pin a truncated cert forever -- and /run rides the
|
||||
# snapshot, so "forever" outlives the container. Fetch to a temp name and
|
||||
# rename only on success.
|
||||
if [ ! -s ca.rsa.4096.crt ]; then
|
||||
run "could not download PIA's CA certificate" \
|
||||
curl -sf -m 20 -o ca.crt.part \
|
||||
https://raw.githubusercontent.com/pia-foss/manual-connections/master/ca.rsa.4096.crt
|
||||
[ -s ca.crt.part ] || die "PIA's CA certificate downloaded empty"
|
||||
mv ca.crt.part ca.rsa.4096.crt
|
||||
fi
|
||||
|
||||
local u p tok srv sip scn priv pub resp ep gw dns
|
||||
u=$(sed -n 1p "$CREDS"); p=$(sed -n 2p "$CREDS")
|
||||
tok=$(curl -sf -m 25 -u "$u:$p" \
|
||||
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
|
||||
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA authentication failed - check $CREDS"
|
||||
[ -n "$u" ] && [ -n "$p" ] || die "$CREDS needs two lines: username, then password"
|
||||
|
||||
curl -sf -m 30 https://serverlist.piaservers.net/vpninfo/servers/v6 | head -1 > servers.json
|
||||
tok=$(run "PIA rejected the credentials in $CREDS, or could not be reached" \
|
||||
curl -sf -m 25 -u "$u:$p" \
|
||||
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
|
||||
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA returned no token - check the credentials in $CREDS"
|
||||
|
||||
run "could not fetch PIA's server list" \
|
||||
curl -sf -m 30 https://serverlist.piaservers.net/vpninfo/servers/v6 \
|
||||
| head -1 > servers.json
|
||||
srv=$(jq -r --arg r "$REGION" '.regions[] | select(.id==$r) | .servers.wg[0]' servers.json)
|
||||
sip=$(echo "$srv" | jq -r .ip); scn=$(echo "$srv" | jq -r .cn)
|
||||
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region $REGION"
|
||||
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region '$REGION'"
|
||||
|
||||
priv=$(wg genkey); pub=$(echo "$priv" | wg pubkey)
|
||||
printf '%s' "$priv" > wg.priv; chmod 600 wg.priv
|
||||
# umask, not a later chmod: the file is created under the inherited 0022
|
||||
# otherwise, so the key is world-readable for the moment in between.
|
||||
( umask 077; priv=$(wg genkey); printf '%s' "$priv" > wg.priv )
|
||||
priv=$(cat wg.priv); pub=$(printf '%s' "$priv" | wg pubkey)
|
||||
|
||||
# The token goes in on stdin as a curl config rather than in the argv, where
|
||||
# `ps` and /proc/*/cmdline expose it to every process in the container --
|
||||
# verified. It is a ~24h bearer credential for the whole PIA account.
|
||||
# PIA pins its certificate to the server's common name, which is why this
|
||||
# connects by CN and lets --connect-to point that name at the real address.
|
||||
resp=$(curl -sf -m 25 -G --connect-to "$scn::$sip:" --cacert ca.rsa.4096.crt \
|
||||
--data-urlencode "pt=$tok" --data-urlencode "pubkey=$pub" \
|
||||
"https://$scn:1337/addKey")
|
||||
resp=$(printf -- '--data-urlencode "pt=%s"\n--data-urlencode "pubkey=%s"\n' "$tok" "$pub" \
|
||||
| run "could not register the key with $scn" \
|
||||
curl -sf -m 25 -G -K - --connect-to "$scn::$sip:" \
|
||||
--cacert ca.rsa.4096.crt "https://$scn:1337/addKey")
|
||||
[ "$(echo "$resp" | jq -r .status)" = OK ] || die "key registration failed: $resp"
|
||||
|
||||
: > "$STATE/routes"
|
||||
ip link del "$IFACE" 2>/dev/null || true
|
||||
ip link add "$IFACE" type wireguard
|
||||
ip link add "$IFACE" type wireguard 2>/dev/null || \
|
||||
die "could not create a WireGuard interface." \
|
||||
"The Docker host's kernel has no 'wireguard' module."
|
||||
wg set "$IFACE" private-key wg.priv \
|
||||
peer "$(echo "$resp" | jq -r .server_key)" \
|
||||
endpoint "$(echo "$resp" | jq -r .server_ip):$(echo "$resp" | jq -r .server_port)" \
|
||||
@@ -102,36 +163,41 @@ up() {
|
||||
ip link set "$IFACE" up
|
||||
|
||||
if [ "${1:-}" = "--full" ]; then
|
||||
ep=$(echo "$resp" | jq -r .server_ip)
|
||||
gw=$(ip route show default | awk '{print $3; exit}')
|
||||
# `default dev eth0` with no `via` yields the literal "eth0" here, which
|
||||
# would make every exclusion below a malformed no-op.
|
||||
[[ $gw =~ ^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+$ ]] || \
|
||||
die "no usable default gateway to pin the tunnel against (got '${gw:-none}')"
|
||||
|
||||
# PIA's resolvers are required in --full. Without them the 10/8 exclusion
|
||||
# below is already in place, so every lookup would go to the container's
|
||||
# own resolver *outside* the tunnel -- a full tunnel leaking all its DNS,
|
||||
# reported by `status` as perfectly healthy.
|
||||
dns=$(echo "$resp" | jq -r '.dns_servers[]? // empty' | head -2)
|
||||
[ -n "$dns" ] || die "PIA returned no DNS servers; refusing a full tunnel that would leak every lookup"
|
||||
|
||||
# Pin the endpoint to the pre-existing gateway first, so the tunnel's own
|
||||
# packets do not try to route through the tunnel. Then beat the default
|
||||
# route with two half-routes rather than replacing it -- nothing to restore
|
||||
# on teardown, and the container keeps working if this script dies midway.
|
||||
ep=$(echo "$resp" | jq -r .server_ip)
|
||||
gw=$(ip route show default | awk '{print $3; exit}')
|
||||
add_route "$ep/32 via $gw"
|
||||
add_route "0.0.0.0/1 dev $IFACE"
|
||||
add_route "128.0.0.0/1 dev $IFACE"
|
||||
add_route "$ep/32" via "$gw"
|
||||
add_route 0.0.0.0/1 dev "$IFACE"
|
||||
add_route 128.0.0.0/1 dev "$IFACE"
|
||||
|
||||
# Keep container, host and LAN traffic off the tunnel. Longer prefixes than
|
||||
# the two halves above, so these win.
|
||||
for n in $PRIVATE_NETS; do add_route "$n via $gw"; done
|
||||
for n in $PRIVATE_NETS; do add_route "$n" via "$gw"; done
|
||||
|
||||
# PIA's resolver lives inside 10/8, so pin it back through the tunnel with a
|
||||
# /32 -- longer still, so it beats the 10.0.0.0/8 exclusion just added.
|
||||
# Using PIA's resolver rather than the container's keeps DNS from leaking,
|
||||
# and the container's own resolver is unreachable from inside the tunnel.
|
||||
dns=$(echo "$resp" | jq -r '.dns_servers[]? // empty' | head -2)
|
||||
if [ -n "$dns" ]; then
|
||||
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
|
||||
for d in $dns; do add_route "$d/32 dev $IFACE"; done
|
||||
# resolv.conf is a bind mount: write through it, never replace it.
|
||||
for d in $dns; do echo "nameserver $d"; done > /etc/resolv.conf
|
||||
else
|
||||
echo "pia-wg: warning - PIA returned no DNS servers; leaving resolv.conf alone" >&2
|
||||
fi
|
||||
# PIA's resolvers live inside 10/8, so pin them back through the tunnel with
|
||||
# /32s -- longer still, so they beat the exclusion just added.
|
||||
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
|
||||
for d in $dns; do add_route "$d/32" dev "$IFACE"; done
|
||||
# resolv.conf is a bind mount: write through it, never replace it.
|
||||
for d in $dns; do echo "nameserver $d"; done > /etc/resolv.conf
|
||||
echo "full tunnel: public traffic exits via PIA; private ranges stay local"
|
||||
else
|
||||
add_route "1.1.1.1/32 dev $IFACE"
|
||||
add_route 1.1.1.1/32 dev "$IFACE"
|
||||
echo "test route only: 1.1.1.1 goes via PIA, everything else unchanged"
|
||||
fi
|
||||
|
||||
@@ -141,8 +207,15 @@ up() {
|
||||
|
||||
down() {
|
||||
[ "$(id -u)" = 0 ] || die "run with sudo"
|
||||
# Only restore something that actually looks like a resolver file. Restoring
|
||||
# an empty or truncated backup leaves the container with no DNS at all, which
|
||||
# is worse than leaving the current one alone.
|
||||
if [ -f "$STATE/resolv.conf.bak" ]; then
|
||||
cat "$STATE/resolv.conf.bak" > /etc/resolv.conf
|
||||
if grep -q '^nameserver' "$STATE/resolv.conf.bak" 2>/dev/null; then
|
||||
cat "$STATE/resolv.conf.bak" > /etc/resolv.conf
|
||||
else
|
||||
echo "pia-wg: warning - saved resolv.conf looks empty; leaving the current one alone" >&2
|
||||
fi
|
||||
rm -f "$STATE/resolv.conf.bak"
|
||||
fi
|
||||
if [ -f "$STATE/routes" ]; then
|
||||
@@ -153,6 +226,10 @@ down() {
|
||||
rm -f "$STATE/routes"
|
||||
fi
|
||||
ip link del "$IFACE" 2>/dev/null || true
|
||||
# /run is in the writable layer and `docker commit` bakes it into the
|
||||
# project's snapshot image, so a key left here rides that image into every
|
||||
# future container. Verified: a snapshot already carried one.
|
||||
rm -f "$STATE/wg.priv"
|
||||
echo "tunnel down"
|
||||
}
|
||||
|
||||
@@ -197,5 +274,5 @@ case "${1:-}" in
|
||||
up) shift; up "${1:-}" ;;
|
||||
down) down ;;
|
||||
status) status ;;
|
||||
*) sed -n '2,20p' "$0" | sed 's/^# \{0,1\}//'; exit 1 ;;
|
||||
*) sed -n '2,27p' "$0" | sed 's/^# \{0,1\}//'; exit 1 ;;
|
||||
esac
|
||||
|
||||
Reference in New Issue
Block a user