Compare commits

..
Author SHA1 Message Date
shadow-testandClaude Opus 5 32149603b2 Stop a failed up from tearing down a working tunnel
The previous commit moved `down` to the top of `up` to fix a resolv.conf
idempotency bug, and in doing so put it *before* every network fetch that can
fail. Re-review caught it and I reproduced it: with a tunnel up, an `up` that
fails on bad credentials left `pia0` gone and traffic silently back on the real
address, while the error talked only about credentials. A privacy regression
introduced by a correctness fix.

`down` now runs after the token, server list and key registration have all
succeeded — nothing above that line touches the network stack — and still
clears the stale backup it was added for.

Everything after it is covered by a rollback. Note this is an EXIT trap with a
flag, not `trap ... ERR`: my first attempt used ERR and did not fire at all,
because ERR is not inherited by shell functions without `set -E`, so a failure
inside add_route missed it, and `die` exits explicitly, which is not an error.
Verified by forcing a route collision — the tunnel is torn down and DNS is
intact, where before the fix it was left half-configured with DNS dead.

Also from the review:

- **The private key lived on disk for the whole life of the tunnel.**
  `commit_container_snapshot` blanks env vars, never files, and nothing tears
  the tunnel down before a recreate or migrate — so the `down`-time cleanup
  never covered the path that put a key in a snapshot in the first place. It is
  now deleted the moment `wg set` has read it; the kernel keeps its own copy,
  verified by checking the interface still works afterwards.
- **`up` claimed success without a handshake.** An unreachable peer still
  routes — into a black hole — so `up --full` could exit 0 having pointed all
  traffic and resolv.conf at a peer that never answered, with `status` printing
  "mode: full tunnel". Now polls for a handshake and rolls back if none arrives.
- **`status` needed root and did not check.** `wg show` fails unprivileged and
  was swallowed, so an unprivileged run printed "no tunnel up" and then "mode:
  full tunnel" in the same breath. An agent reading the first line would re-run
  `up` — which, before the fix above, destroyed the tunnel it failed to see.
- **The killswitch bullet was false.** It said `iptables` is not in the image;
  it is, so a killswitch is buildable. It stays unbuilt because it would cut
  Claude Code's own API traffic — an honest reason, unlike the previous one.
- `install_feature_skill` rejects path-traversal names, not just blank ones —
  verified `../skills` would have deleted the whole skills directory including
  Mission Control's — and reports `mkdir`/`cp` failures instead of printing a
  success line regardless.
- The usage text ended by printing `set -euo pipefail`, off by one line.
- HOW-TO-USE claimed "the container says so on start". entrypoint prints to
  PID 1's stdout, which no terminal or UI surfaces — `docker logs` appears
  nowhere in the repo. Now points at the migration pre-flight, which does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:29:30 -07:00
shadow-testandClaude Opus 5 f6e5cf3f05 Fix what review found in the skill: five real defects
Adversarial review of #29 found bugs I confirmed by reproducing each one.

**Every hand-written error message was unreachable.** `tok=$(curl ...)` is a
plain assignment, so `set -e` acts on the command substitution before the
following `|| die` can run. A wrong password produced exit 22 and no output at
all — the most likely way this gets used wrongly, and the least explained.
All four captures now go through a `run` helper that takes a *description*
rather than echoing the command, because one of them carries the account
password in `-u`.

**`up` was not idempotent, and the second run destroyed DNS.** The resolv.conf
backup was copied unconditionally, so `up --full` twice overwrote the good
backup with PIA's own resolvers; the later `down` then "restored" those and
left the container with no working DNS and no way back. `up` now runs `down`
first. Verified: two `up --full` runs, then `down`, and the backup still holds
the original 192.168.65.7.

**An empty gateway produced total connectivity loss, reported as healthy.**
`$gw` was never validated and `add_route` swallowed every failure to /dev/null.
The two half-routes need no gateway and would succeed, so the tunnel captured
everything while the exclusions keeping DNS and the Docker host reachable
silently did not exist — and `status` still printed "full tunnel". Routes are
now fatal on failure, and a via-less default (`$3` is the literal "eth0") is
rejected.

**The PIA session token was in the process arguments** — confirmed in `ps` and
/proc/*/cmdline, a ~24h bearer credential for the account readable by anything
in the container. It now goes to curl on stdin as a config. Verified: 60 polls
across a full `up`, zero sightings.

**The preflight diagnosed the wrong kernel module.** It checked /dev/net/tun
and blamed the tun module, but kernel WireGuard is a netlink interface and does
not use it — verified by creating one with NET_ADMIN and no tun device. The
check is dropped (the container could not have started without the device
anyway) and `ip link add` now reports the real dependency.

Also: a full tunnel with no DNS servers from PIA used to warn and carry on,
which is a tunnel leaking every lookup while reporting itself healthy — now
fatal. `down` validates the backup before restoring it, so a truncated one
cannot leave the container with no resolver at all. `wg.priv` is shredded on
teardown and created under umask 077, because /run rides `docker commit` into
the snapshot image. A mistyped `up --ful` is rejected instead of silently
giving a test route.

entrypoint: `install_feature_skill` gets `local`, a blank-name guard (the
disabled branch would otherwise `rm -rf` the whole skills directory under a
persisted volume), `-e`/`-L` so a leftover *file* at the destination is cleaned
up, and a chown of the parent so `claude` can still add skills of their own
when Mission Control is off. When the base image predates the skill it now says
so instead of returning silently — and `/opt/triple-c-skills` joins
FEATURE_PROBES so the migration pre-flight reports it. Docs corrected to match:
neither half reaches an existing project without a migration.

`vpn_env_var` extracted and tested, pinning the property the whole removal path
rests on — that the variable is emitted as 0 rather than omitted.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:25:12 -07:00
shadow-testandClaude Opus 5 43c7ad1478 Report which exit is which, instead of one ambiguous "public IP"
`status` probed https://1.1.1.1/cdn-cgi/trace and printed the answer as
"public IP". In test mode 1.1.1.1 is the *only* address routed into the tunnel,
so that line reported a PIA exit while every other packet left directly — a
test tunnel reading exactly like a full one.

Found on a live container: default route still via eth0, one 1.1.1.1/32 route
through pia0, and the old status line claiming a PIA public IP. This is a
plausible route to concluding the VPN is on when it is not, which is close to
the confusion this skill exists to prevent.

Status now names the mode and, in test mode, prints both exits with the real
address called out. 1.0.0.1 serves the same trace endpoint as 1.1.1.1 and is
never routed into the tunnel, so the direct exit can be probed without DNS.

Verified against all four states: no tunnel, test mode on a live tunnel that
was already up, full tunnel, and after teardown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:24:41 -07:00
shadow-testandClaude Opus 5 70c0a8bf7a Ship a pia-vpn skill with the VPN support toggle
The toggle grants CAP_NET_ADMIN and /dev/net/tun and stops there, which users
reasonably read as "turn the VPN on" — the gap between the two is the reported
bug that the default network does not route through a VPN. Close it by giving
the container an agent-usable way to build the tunnel, rather than leaving
each project to rediscover it.

container/skills/ is baked to /opt/triple-c-skills and installed into
~/.claude/skills/ by entrypoint.sh from VPN_SUPPORT_ENABLED, mirroring how
Mission Control installs its own. Staged under /opt because ~/.claude is a
volume mount that would mask an image copy from first start.

Three details that are not incidental:

- The variable is sent as 0 rather than omitted when off, because ~/.claude
  persists: entrypoint has to be *told* to remove a skill left by an earlier
  run with the toggle on, and an absent variable cannot say that. A stale skill
  is worse than none, since it instructs an agent to use a capability the
  container no longer has.
- It is reserved in RESERVED_ENV_EXACT alongside MISSION_CONTROL_ENABLED, or a
  custom env var of the same name could claim the skill without the capability
  behind it. Covered by a test.
- The skill is re-copied on every start, rm -rf'd first, so fixes reach existing
  projects and files dropped from a later version do not linger.

The skill itself carries the three things that are easy to get wrong: that a
full tunnel captures the Docker resolver and takes DNS down with it, that an
IP-literal health check cannot see a dead resolver, and that no tunnel survives
a restart while /run state riding the snapshot makes it look as though one did.
It also states what it deliberately does not do — no killswitch, no autostart —
so an agent proposes those as decisions rather than improvising them.

pia-wg.sh preflights CAP_NET_ADMIN by capability bit rather than letting the
first `ip` call fail with a bare EPERM that points nowhere near the setting
that needs changing. Credentials stay in a file (~/pia-creds, PIA_CREDS to
override) rather than the environment, where docker inspect and every process
in the container would see them.

Tested: install/refresh/remove/no-op paths of install_feature_skill against the
real function; preflight with and without the capability; and a full up --full
/ down round trip, confirming DNS via PIA's resolvers, api.anthropic.com
reachable through the exit, and routes and resolv.conf restored on teardown.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:24:41 -07:00
shadow-testandClaude Opus 5 92d64cf252 Ship iptables, not nftables — nftables forfeits macOS
Build Container / build-container (pull_request) In progress
Build App (Preview) / compute-version (pull_request) Successful in 6s
Build App (Preview) / create-release (pull_request) Successful in 2s
Build App (Preview) / build-macos (pull_request) Successful in 2m36s
Build App (Preview) / build-linux (pull_request) Successful in 6m29s
Build App (Preview) / build-windows (pull_request) Successful in 6m59s
Build App (Preview) / prune-previews (pull_request) Successful in 13s
Re-review overturned the previous commit's package choice, and verifying it
proved the reviewer right.

`wg-quick` picks nft *unconditionally* when it is present (`type -p nft`, line
241), so installing nftables makes the iptables path unreachable. Its nft
ruleset then needs a third expression family the iptables path does not.
Isolating the rules on this host, the two connmark rules install fine and this
is what fails:

    nft add rule ... fib saddr type != local drop
    Error: Could not process rule: No such file or directory

That decides it, because of how the hosts differ. LinuxKit's kernel config —
Docker Desktop for Mac, identical on x86_64 and aarch64:

    CONFIG_NETFILTER_XT_CONNMARK=y      <- the iptables path works
    # CONFIG_NFT_FIB_IPV4 is not set    <- the nft path does not

So nftables would have broken the platform it was added to fix. With iptables,
full tunnels work on native Linux, Docker Desktop for Mac, and WSL2 from 6.6.
Costs 7,203 kB rather than 5,614 kB on amd64.

That also means the mechanism the previous commit documented was wrong: with
nftables installed `xt_CONNMARK` is never consulted, and the real blocker on
that path is `nft_fib_ipv4`. Rewritten around what actually fails.

A second failure neither round had found: `wireguard-tools` only *Suggests*
`openresolv | resolvconf`, so neither is installed, and every provider's stock
config has a `DNS =` line. That fails in `set_dns()` — before any routing — so
it takes split tunnels down too, contradicting what this PR previously claimed:

    [#] resolvconf -a sp -m 0 -x
    /usr/bin/wg-quick: line 32: resolvconf: command not found   EXIT=127

Not fixed, deliberately: `openresolv` has no installation candidate on noble,
and `resolvconf` resolves only by pulling in systemd-resolved — a resolver
daemon and systemd units, into a container with no systemd. Documented instead.

Smaller corrections from the same review:

- the size caveat blamed ~209 kB of libelf1t64; for this package set the real
  over-count is libelf1t64 + netbase. Restated, and arm64 now given against the
  real base rather than left as a bare-ubuntu figure.
- the manual-install fallback omitted `iproute2`, so it left the user without
  `ip` — the command the tunnel needs most.
- "Without it" had been orphaned from its antecedent by inserted paragraphs and
  read as referring to configuring a tunnel.
- the migration probe said "VPN support", presenting VPN as a feature gained to
  users who never enabled it. Now names the tools and the toggle.
- "What's Inside the Container" gains a row; the key-material-in-snapshot
  hazard was in CLAUDE.md only, and is the one genuinely user-facing warning
  here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 16:24:07 -07:00
shadow-testandClaude Opus 5 ab2c75d0b2 Ship a firewall backend, and correct three claims review disproved
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 3s
Build App (Preview) / build-macos (pull_request) Successful in 2m46s
Build App (Preview) / build-linux (pull_request) Successful in 7m22s
Build App (Preview) / build-windows (pull_request) Successful in 7m43s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
Build Container / build-container (pull_request) Successful in 11m20s
Review of #28 found the iptables exclusion was justified by a false premise,
and I confirmed it: `wireguard-tools` declares `Recommends: nftables | iptables`,
`--no-install-recommends` strips it, and `wg-quick`'s add_default() shells out
to a firewall backend with no `type -p` guard. Measured on the image as this PR
shipped it:

    [#] iptables-restore -n
    /usr/bin/wg-quick: line 32: iptables-restore: command not found
    wg-quick EXIT=127

That fires for `AllowedIPs = 0.0.0.0/0` — every stock full-tunnel config from
every provider — not for a desktop client's killswitch as the comment claimed.
Split tunnels are unaffected.

Ship `nftables` rather than `iptables`: wg-quick prefers it (`type -p nft`, so
with both installed iptables is dead weight), it is first in the package's own
Recommends, and it is half the size.

The review's proposed fix stopped there; it does not hold. Adding nftables does
not make wg-quick work on this host, and neither does iptables:

    Warning: Extension CONNMARK revision 0 not supported, missing kernel module?

`Table=auto` routes by fwmark and needs xt_CONNMARK from the *host* kernel.
WSL2 has none and containers have no /lib/modules to load one from. So this
fixes native Linux and Docker Desktop for Mac — which other WHP users are on —
and cannot fix Docker Desktop for Windows, where the answer is to add routes
with `ip route` directly. Documented rather than left to be rediscovered.

Also from review:

- "`ip` and `wg` are always present" was false. A project keeps the base image
  it was first built from, so this reaches new projects only. Reworded to match
  the wording already used for the Playwright libraries, and `/usr/bin/wg` added
  to FEATURE_PROBES so an existing project is *told* it is missing VPN tooling
  and prompted to migrate, rather than finding out via `wg: command not found`.
- "no client is installed" contradicted shipping `wg` four lines earlier. The
  true claim is that no tunnel is configured or started.
- The size figure measured against bare ubuntu:24.04, which over-counts by the
  ~209 kB of libelf1t64 the real base already has, and covered one arch. Now
  measured against the current base on amd64 and stated for arm64 too, per the
  standard CLAUDE.md sets for the Playwright layer.
- `/run` persistence conflated two mechanisms: same-container files on a
  stop/start, `docker commit` on a recreation. Both stated, plus the corollary
  that key material written to /run ends up inside a snapshot image — observed,
  a `wg.priv` was already sitting in one.
- The DNS bullet presented a Docker Desktop address as the general case. Now
  leads with the mechanism, notes 127.0.0.11 on a user-defined network is
  unaffected, and adds the two things the advice omitted: a resolver the tunnel
  can reach (or it leaks every query), and pinning the endpoint via the old
  gateway (or the tunnel routes through itself).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:55:44 -07:00
shadow-testandClaude Opus 5 00937745f7 Ship the tools the VPN toggle grants capability for
Build Container / build-container (pull_request) Successful in 11m28s
`vpn_support_enabled` hands a project CAP_NET_ADMIN and /dev/net/tun, and the
image then contains no `ip` and no `wg` — a capability with nothing able to
exercise it. Bake `iproute2` and `wireguard-tools` (~4.3 MB with deps).

They belong in the image rather than a runtime install for the reason the
Dockerfile already gives for the Playwright libraries: the writable layer is
lost on base-image migration. A hand-installed `wg` works until an upgrade and
then vanishes, which presents as a tunnel that will not come up rather than as
a missing package. One project only had `ip` at all because MariaDB pulled in
iproute2 as a transitive dependency.

`iptables` stays out. Only a desktop client's killswitch wants it, and those
clients need a GUI the container cannot provide.

Also correct three things the docs left users to discover:

- the toggle grants capability and routes nothing, which is being reported as
  the default network "not routing through the VPN automatically"
- no tunnel survives a restart, and `/run` state riding the snapshot makes it
  look as though one did while traffic goes out the real address
- a full tunnel captures the Docker resolver, which sits outside the
  container's subnet, and takes DNS down with it — Claude Code then reports a
  connection failure because it cannot resolve api.anthropic.com, and a health
  check aimed at an IP literal passes throughout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 12:38:37 -07:00
jknapp 01e72e4785 Merge pull request 'Let a project's container run a VPN client' (#27) from feat/vpn-support into main
Build App / compute-version (push) Successful in 3s
Build App / build-macos (push) Successful in 2m37s
Build App / build-linux (push) Successful in 5m33s
Build App / build-windows (push) Successful in 6m11s
Build App / create-tag (push) Successful in 6s
Build App / sync-to-github (push) Successful in 52s
2026-08-14 15:30:03 +00:00
shadow-testandClaude Opus 5 2b35aa8c16 Explain a missing tun device where the failure actually happens
Build App (Preview) / compute-version (pull_request) Successful in 3s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m37s
Build App (Preview) / build-linux (pull_request) Successful in 5m30s
Build App (Preview) / build-windows (pull_request) Successful in 5m55s
Build App (Preview) / prune-previews (pull_request) Successful in 3s
Review caught that the device guard was wired to the wrong call. The
daemon does not resolve `--device` at create: verified against Docker
29.7, `docker create --device /dev/does-not-exist` succeeds and prints an
id, and runc only resolves the device — and validates sysctls — when it
builds the container. So on a host with no tun module the create returns
fine and `start` fails, which means the explanation never ran and the
user saw the raw daemon string naming a path they would go looking for on
the wrong machine. The unit tests fed the create-side string straight in,
so they confirmed a function no real failure could reach.

Move the guard onto `start_container`, covering create as well in case a
future daemon checks earlier. It no longer takes `vpn_support_enabled` —
`start_container` has a container id and no project, and nothing else in
Triple-C ever requests a device, so an error naming /dev/net/tun is
unambiguous on its own. The test now uses the daemon's verbatim message
via bollard's real Display format.

Also from review:

  * Soften the security claim. Docker does not enable user-namespace
    remapping by default, so this is a real CAP_NET_ADMIN in the initial
    user namespace with only the network namespace confining it. It
    cannot touch host interfaces, but "confers no authority outside the
    container" was too strong: within its namespace it can set
    promiscuous mode and add addresses, routes and NAT on the shared
    docker0 segment, which puts sibling containers — the LiteLLM gateway
    among them — within ARP-spoofing reach, and it can flush netfilter
    rules sandbox mode may rely on. Said plainly in the code, CLAUDE.md
    and HOW-TO-USE.
  * Drop Tailscale from the list of clients needing this. Its
    --tun=userspace-networking mode needs neither the capability nor the
    device, and listing it invites granting NET_ADMIN for nothing.
  * Say in the toggle's own hint that changing it recreates the
    container, matching how every other recreation-triggering setting is
    labelled. The tab's generic "stop the container first" chip does not
    tell the user what is about to happen.
  * Add RuntimeSection tests: saves on, saves off explicitly rather than
    dropping the key, reflects state, is disabled while running, and
    carries the recreation warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:15:05 -07:00
shadow-testandClaude Opus 5 65a3d4eb29 Let a project's container run a VPN client
Build App (Preview) / compute-version (pull_request) Successful in 4s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m39s
Build App (Preview) / build-linux (pull_request) Successful in 7m10s
Build App (Preview) / build-windows (pull_request) Successful in 6m33s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
A VPN client installed in a container today starts, runs, and then hangs
until its connection times out. Nothing reports an error: a default
container has no /dev/net/tun to open and no CAP_NET_ADMIN to add an
interface or a route with, and clients surface that as a generic timeout
rather than a permissions failure.

Add an opt-in per-project "VPN support" switch granting the three things
a tunnel needs. They are useless individually, which is why
vpn_host_config() defines the set in one place and the tests assert all
of it:

  * CAP_NET_ADMIN — Docker's default bounding set has net_raw but not
    net_admin, so a client can ping but never connect.
  * /dev/net/tun — passed through from the host so the kernel's tun
    module backs it, rather than mknod-ed inside.
  * net.ipv4.conf.all.src_valid_mark — WireGuard's wg-quick sets this and
    cannot from inside a container, /proc/sys being read-only, so its
    handshakes are dropped by reverse-path filtering.

Off by default and deliberately opt-in: NET_ADMIN lets anything in the
container reconfigure that container's network stack. It is namespaced —
no authority over the host's interfaces or any other container.

Capabilities and devices are fixed when a container is created, so this
is container state and takes the label-and-compare treatment.
triple-c.vpn-support is written unconditionally, false included, for the
usual docker commit reason: a true stamped once would ride the snapshot
image into every future container and make the switch impossible to turn
back off. A missing label reads as false and off is byte-identical to
today, so no existing project is churned.

Requesting the device fails at creation when the host kernel has no tun
module, which would otherwise surface as a project that simply refuses to
start. explain_create_failure() rewrites that one error to name the
switch and the Docker-Desktop-VM-versus-your-machine distinction, and
leaves every other failure untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:06:39 -07:00
jknapp 2b2d9da606 Merge pull request 'Number releases by the highest one already published, not by drift' (#26) from fix/monotonic-release-version into main
Build App / compute-version (push) Successful in 12s
Build App / build-macos (push) Successful in 2m37s
Build App / build-windows (push) Successful in 5m44s
Build App / build-linux (push) Successful in 6m3s
Build App / sync-to-github (push) Successful in 11s
Build App / create-tag (push) Successful in 24s
2026-08-14 06:00:50 +00:00
shadow-testandClaude Opus 5 3741e0fef5 Number releases by the highest one already published, not by drift
The Linux release upload failed with a bare "exitcode '1'" and no output.
The cause was not the upload: compute-version handed it a version that
had already been released three days earlier.

The patch number was `git rev-list --count <highest tag>..HEAD` — how far
HEAD has drifted from whichever tag sorts highest, which resets to zero
every time a tag is cut. It is not a counter, and the published history
is what the old formula returned at each point:

  v0.4.0 -> 3 commits -> v0.4.3    looked fine
  v0.4.3 -> 4 commits -> v0.4.4    fine by luck, 4 > 3
  v0.4.4 -> 2 commits -> v0.4.2    went backwards
  v0.4.4 -> 6 commits -> v0.4.6    jumped, skipping .5
  v0.4.6 -> 3 commits -> v0.4.3    already taken

So the line published 0.4.0, 0.4.3, 0.4.4, 0.4.2, 0.4.6 in that order,
never used 0.4.1 or 0.4.5, and then came back round to 0.4.3.

The patch is now one past the highest already used. Suffixed tags count
towards that: create-tag is skipped whenever a platform job fails, so a
run can publish v0.4.7-mac and never create the plain v0.4.7, and reading
only unsuffixed tags would hand the same number out twice. A commit that
is already tagged reuses its own tag, so re-running a build does not mint
a version.

Reusing a number was doing real damage, not just failing. macOS and
Windows delete-then-upload each asset, so they took the duplicate in
their stride and rewrote v0.4.3-mac and v0.4.3-win — public since
Aug 11 — with today's binaries. Linux is the only platform that failed,
and failing was the correct outcome; its v0.4.3 assets are the only ones
still original.

Linux also gets the idempotent get-or-create the other two already had,
plus `set -euo pipefail` and `-fsS`. Its `curl -s` with no `-f` is why a
409 produced no diagnostic at all: the HTTP error was swallowed, the id
grep came back empty, and the step died without ever printing why.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 22:57:36 -07:00
jknapp 0e6566d903 Merge pull request 'Fix Playwright setup destroying its own install, and the migration notice that stayed silent' (#25) from fix/playwright-container-setup into main
Build Container / build-container (push) Successful in 57s
Build App / compute-version (push) Successful in 5s
Build App / build-macos (push) Successful in 2m37s
Build App / build-windows (push) Successful in 5m45s
Build App / build-linux (push) Failing after 5m18s
Build App / create-tag (push) Skipped
Build App / sync-to-github (push) Skipped
2026-08-14 04:34:20 +00:00
shadow-testandClaude Opus 5 84a67fcd0d Stop an empty base-image label from silencing the migration notice
Build App (Preview) / compute-version (pull_request) Successful in 3s
Build Container / build-container (pull_request) Successful in 1m5s
Build App (Preview) / create-release (pull_request) Successful in 2s
Build App (Preview) / build-macos (pull_request) Successful in 2m41s
Build App (Preview) / build-linux (pull_request) Successful in 5m31s
Build App (Preview) / build-windows (pull_request) Successful in 6m24s
Build App (Preview) / prune-previews (pull_request) Successful in 8s
A project can be out of date and say nothing about it, in two ways that
compound: the lineage lookup treats "unknown" as an answer, and the
fallback that exists for unknown lineage disappears when its probe fails.

`create_container` always writes triple-c.base-image-id, even when the
value is unknown — deliberately, so an inherited image label cannot ride
a snapshot forever. That makes Some("") the ordinary reading from a
container whose lineage was never established. The lookup filtered for
emptiness only on the final result, so that empty string satisfied the
container branch and skipped the snapshot entirely: a snapshot that had
recorded a real lineage was never consulted, and the project reported
"unknown" with the answer one lookup away. Each source is now filtered
before it can answer, in pick_recorded_lineage, which is a plain function
so the case has a test that fails against the old logic.

A genuinely pre-label project stays unknown, and should: its ancestor is
not knowable, and inventing one would make it look permanently current.
The probe is the intended signal for those — but if the probe failed,
get_container_staleness returned early with nothing populated, the banner
found no gaps and rendered null, and the probe_error it already knew how
to display sat behind a gate that returned before reaching it. Silence
there is indistinguishable from "up to date", and it is likeliest for the
oldest and largest projects, whose manifests are the ones apt to exceed
the inspection limit — one real project measured 6.93 MB against an 8 MB
cap. An unknown-lineage container whose probe failed now says the check
could not be completed, with the reason, under the tone that means
unresolved rather than the one that means something is wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 21:24:51 -07:00
shadow-testandClaude Opus 5 f3cc1c4c17 Stop Playwright setup from deleting the package it just installed
Setting up the browser view failed on every container, and re-running it
reproduced the same broken state, because the setup destroyed its own work.

`install_packages` ran two `npm install --no-save` commands into
/workspace, which has no package.json. With no manifest, npm treats the
command line as the whole statement of what the tree should contain and
prunes the rest, so installing `playwright` second removed the
`@playwright/cli` installed first: "removed 3 packages", leaving an empty
node_modules/@playwright/ behind playwright and playwright-core. That
empty directory is exactly what the pane then reported as missing. The
second install now names both specs; the first one is already present, so
it costs nothing and is only there to stop npm pruning it.

Two failures were waiting behind that one:

Nothing in the tree ever configured the browser, so playwright-cli fell
back to channel `chrome` — system Google Chrome — with the Chromium
sandbox on. These containers forbid unprivileged user namespaces, so it
aborted with "Failed to move to new namespace ... Operation not
permitted"; on a base image without Google Chrome the same default failed
as "Chromium distribution 'chrome' is not found". entrypoint.sh now seeds
~/.playwright/cli.config.json on every start, which is the only way to
reach existing projects: ~/.playwright is inside the home volume, so an
image copy would reach new projects only.

The launch check passed for a configuration the viewer never uses. It
launched bundled chromium with no channel, which resolves to
chromium-headless-shell, while the viewer's config pins
chrome-for-testing — the full chromium build, a separate download. A
container could pass every check and still fail in the pane with 'Browser
"chrome-for-testing" is not installed', which is what a stale
chromium-1217 against a wanted chromium-1237 did. Chromium is now
verified on both channels, the sandbox setting is stated rather than
inherited from a default, and a failure names the channel.

triple-c-playwright-heal repairs all of it on a container that is already
broken, including the missing socat that makes the pane report
"127.0.0.1 sent an invalid response" while the container side is
perfectly healthy. It verifies by launching a browser rather than
trusting the preceding steps — which is how the stale-revision case was
found — and lives in /usr/local/bin so a fix to it can still reach an
existing project.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 21:24:37 -07:00
19 changed files with 1869 additions and 59 deletions
+95 -11
View File
@@ -39,13 +39,48 @@ jobs:
MAJOR_MINOR=$(cat VERSION | tr -d '[:space:]')
echo "Major.Minor: ${MAJOR_MINOR}"
# Find the latest tag matching v{MAJOR_MINOR}.N (exclude -mac, -win suffixes)
# `|| true` so an empty grep result doesn't fail the step under pipefail.
LATEST_TAG=$(git tag -l "v${MAJOR_MINOR}.*" --sort=-v:refname | grep -E "^v${MAJOR_MINOR}\.[0-9]+$" | head -1 || true)
# The patch number is **one past the highest patch already used**, and
# never a distance.
#
# It used to be `git rev-list --count <highest tag>..HEAD`, which is
# not a counter at all: it measures how far HEAD has drifted from
# whichever tag sorts highest, and that resets to zero every time a
# tag is cut. The published history is the proof — each of these is
# exactly what the old formula returned at the time:
#
# v0.4.0 -> 3 commits -> v0.4.3 looked fine
# v0.4.3 -> 4 commits -> v0.4.4 fine by luck, 4 > 3
# v0.4.4 -> 2 commits -> v0.4.2 went backwards
# v0.4.4 -> 6 commits -> v0.4.6 jumped, skipping .5
# v0.4.6 -> 3 commits -> v0.4.3 already taken; the upload failed
#
# Reusing a version is worse than failing to publish one: the macOS
# and Windows steps replace assets in place, so a duplicate silently
# rewrote a release that had been public for three days. Monotonic
# numbering is what stops that at the source.
#
# Suffixed tags count too. `create-tag` is skipped when any platform
# job fails, so a run can publish v0.4.7-mac and never create the
# plain v0.4.7 — reading only unsuffixed tags would then hand the
# same number out twice.
HIGHEST=$(git tag -l "v${MAJOR_MINOR}.*" \
| grep -E "^v${MAJOR_MINOR}\.[0-9]+(-mac|-win)?$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+).*/\1/" \
| sort -n | tail -1 || true)
if [ -n "$LATEST_TAG" ]; then
echo "Latest matching tag: ${LATEST_TAG}"
PATCH=$(git rev-list --count "${LATEST_TAG}..HEAD")
# A re-run of a commit that already released must not mint a new
# version just because its own tag now exists.
EXISTING=$(git tag --points-at HEAD \
| grep -E "^v${MAJOR_MINOR}\.[0-9]+$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+)$/\1/" \
| sort -n | tail -1 || true)
if [ -n "$EXISTING" ]; then
echo "HEAD is already tagged v${MAJOR_MINOR}.${EXISTING} — reusing it"
PATCH="${EXISTING}"
elif [ -n "$HIGHEST" ]; then
echo "Highest patch already used on this line: ${HIGHEST}"
PATCH=$((HIGHEST + 1))
else
# A minor line nobody has tagged yet is a *new* line, and a new line
# starts at .0 — that is what "we are moving to 0.4.x" means. The
@@ -165,21 +200,70 @@ jobs:
env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -euo pipefail
TAG="v${{ needs.compute-version.outputs.version }}"
# Create release
curl -s -X POST \
# Idempotent get-or-create, matching build-macos. This step used to
# POST /releases unconditionally: against a tag that already existed
# Gitea answered 409, the grep below found no id, and the run died
# with a bare "exitcode '1'" and not one line of output explaining
# it — `curl -s` with no `-f` swallows the HTTP error, so nothing
# ever said "409" or "duplicate tag". Hence -fsS throughout, and
# pipefail so a failure cannot be stepped over.
HTTP_CODE=$(curl -sS -o release.json -w '%{http_code}' \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/tags/${TAG}")
case "${HTTP_CODE}" in
200)
echo "Release ${TAG} already exists, reusing"
;;
404)
echo "Creating release ${TAG}"
curl -fsS -X POST \
-H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/json" \
-d "{\"tag_name\": \"${TAG}\", \"name\": \"Triple-C ${TAG} (Linux)\", \"body\": \"Automated build from commit ${{ gitea.sha }}\"}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases" > release.json
RELEASE_ID=$(cat release.json | grep -o '"id":[0-9]*' | head -1 | grep -o '[0-9]*')
;;
*)
echo "Unexpected ${HTTP_CODE} looking up release ${TAG}:" >&2
cat release.json >&2
exit 1
;;
esac
RELEASE_ID=$(python3 -c "import json,sys; print(json.load(open('release.json')).get('id',''))")
if [ -z "${RELEASE_ID}" ]; then
echo "No release id for ${TAG}; refusing to upload into nothing:" >&2
cat release.json >&2
exit 1
fi
echo "Release ID: ${RELEASE_ID}"
# Upload each artifact
# Replace-not-conflict, so a retry after a partial upload succeeds.
# Versions are monotonic now (see compute-version), so this can only
# ever be replacing an asset from a failed run of this same commit —
# never one belonging to an already-published version.
for file in artifacts/*; do
[ -f "$file" ] || continue
filename=$(basename "$file")
EXISTING_ID=$(curl -sS \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets" \
| python3 -c "import json,sys; t=sys.argv[1]; print(next((a['id'] for a in json.load(sys.stdin) if a.get('name')==t), ''))" "${filename}" || true)
if [ -n "${EXISTING_ID}" ]; then
echo "Deleting existing asset ${filename} (id ${EXISTING_ID})"
curl -fsS -X DELETE \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets/${EXISTING_ID}"
fi
echo "Uploading ${filename}..."
curl -s -X POST \
curl -fsS --http1.1 \
--retry 5 --retry-all-errors --retry-delay 5 \
--max-time 600 \
-X POST \
-H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/octet-stream" \
--data-binary "@${file}" \
+85 -1
View File
@@ -203,7 +203,8 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li
### Container (`container/`)
- **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared
libraries a browser links against (see below)
libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled`
toggle grants capability for (`iproute2`, `wireguard-tools`, `iptables`)
- **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs
`npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own
dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium
@@ -273,6 +274,89 @@ migration and Reset. Four things here are not obvious:
actively **removes** `triple-c-*.crt` when the setting is cleared — `/usr/local/share` rides the
project's snapshot image, so turning the feature off has to undo, not merely stop.
### VPN support (`vpn_support_enabled`, `docker/container.rs`)
An opt-in per-project switch granting the container what a VPN client needs to build a tunnel.
`vpn_host_config()` is the single definition of what that means, and it is unit-tested because a
container is created once by a very long function where a dropped capability is invisible.
- **All three pieces or none.** `CAP_NET_ADMIN` (Docker's default set has `net_raw` but *not*
`net_admin`, so a client can ping but never connect), the `/dev/net/tun` device (absent
entirely from a default container — nothing to open even with the capability), and
`net.ipv4.conf.all.src_valid_mark=1` (WireGuard's `wg-quick` sets it and cannot from inside a
container, since `/proc/sys` is read-only, so handshake packets die to reverse-path filtering).
Any two without the third still presents as a connection that hangs to a timeout, which is why
the tests assert the whole set.
- **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun`
module has to back it.
- **A missing device fails at `start`, not `create` — verified against Docker 29.7.** `docker
create --device /dev/does-not-exist` succeeds and prints an id; runc resolves the device (and
validates sysctls) only when it builds the container. So the guard belongs on the start path:
`explain_container_failure()` covers both and is called from `start_container`, where it has a
container id and no project — which is why it keys off the error naming `/dev/net/tun` rather
than off `vpn_support_enabled`. Nothing else in Triple-C requests a device, so that is
unambiguous. A version of this check wired to `create` alone is dead code that looks correct.
- **`NET_ADMIN` here is not user-namespaced.** Docker does not enable userns remapping by default,
so only the *network* namespace confines it: no reach onto host interfaces, but promiscuous
mode, arbitrary addresses/routes/NAT on the shared `docker0` segment (sibling containers, the
LiteLLM gateway among them, are ARP-spoofable), netlink-triggered host module auto-load, and
enough authority to flush in-container netfilter rules that sandbox mode may rely on. Keep the
code comments honest about this — an earlier draft claimed it "confers no authority" outside the
container, which is too strong.
- **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual
`docker commit` reason: a `true` stamped once would ride the snapshot image into every future
container and make the switch impossible to turn off.
- Off is byte-identical to a container created before the feature existed, and a missing label
reads as `false`, so no existing project is churned.
- **The toggle grants capability and stops there — it routes nothing.** `vpn_host_config()` returns
a cap, a device and a sysctl; no client is installed, no route is touched, no tunnel is started
or restored. Users read the name as "turn the VPN on" and report the default network not routing
through it as a bug. It isn't, and the docs say so explicitly; keep it that way.
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
base-image migration — leaving a project holding the capability with nothing able to exercise it,
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
writable layer, so on a stop/start it is simply the same container's files, and on a recreation
`docker commit` has carried it into the snapshot. Traffic silently reverts to the real address.
Any future autostart or killswitch work starts here.
- **`/run` riding the snapshot means a VPN client's key material can end up in an image.** Verified:
a fresh container off the whp snapshot already contained the `wg.priv` a previous tunnel left in
`/run`. Anything writing key material there inherits the problem — the same `docker commit`
hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks
ephemeral and is not. A VPN client that does this should delete its key on teardown.
- **`iptables` is baked, and picking `nftables` instead would have been wrong.** `Recommends:
nftables | iptables` is stripped by `--no-install-recommends`, and `wg-quick` needs a backend for
any `AllowedIPs = 0.0.0.0/0`. `nftables` is the tempting choice — preferred by `wg-quick`, half
the size — but `wg-quick` picks nft *unconditionally* when present, and its nft ruleset needs
`nft_fib_ipv4`, which LinuxKit (Docker Desktop for Mac) does not build while it *does* build
`xt_CONNMARK`. Shipping nftables would therefore have forfeited Mac. See the Dockerfile comment;
the kernel-config evidence is quoted there.
- **Two `wg-quick` failures remain, and only one is ours to fix.** Full tunnels still need
`xt_CONNMARK`, which WSL2 before 6.6 lacks — nothing installable changes that. And every
provider's stock config carries a `DNS =` line that fails in `set_dns()` before any routing, so it
breaks split tunnels too; `openresolv` has no candidate on noble and `resolvconf` drags in
systemd-resolved, so that one is documented rather than fixed. Driving `wg` and `ip route`
directly avoids both, which is what the skill does.
- **The `pia-vpn` skill is installed *and removed* from `VPN_SUPPORT_ENABLED`.** `container/skills/`
is baked to `/opt/triple-c-skills` and `install_feature_skill()` in `entrypoint.sh` copies it into
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches any project whose base
image has the source, and `rm -rf`'d first, so files dropped from a later version do not linger.
The removal branch matters as much as the install: `~/.claude` is a persisted volume, so a skill
left behind after the toggle goes off would keep instructing an agent to use a capability the
container no longer has. Which is also why the variable is sent as `0` rather than omitted (see
`vpn_env_var`, tested), and why it is in `RESERVED_ENV_EXACT` — a custom env var of that name
could otherwise claim the skill without the capability behind it.
- **Both halves of that live in the base image, so neither reaches an existing project.** A
recreation builds from the project's *own snapshot*, which has no `/opt/triple-c-skills` and no
updated `entrypoint.sh`; only a migration or a Reset delivers them. The install path says so out
loud rather than returning silently, and `/opt/triple-c-skills` is in `FEATURE_PROBES` so the
migration pre-flight lists it as missing. Worth knowing before adding anything else behind an
existing toggle: the label fingerprints *the setting*, not the set of things the setting drives,
so a project already at `true` gets no recreation at all on upgrade.
### Container Lifecycle
Containers use a **stop/start** model (not create/destroy). Installed packages persist across stops. The `.claude` config dir uses a named Docker volume (`triple-c-claude-config-{projectId}`), nested inside the home volume (`triple-c-home-{projectId}`), so OAuth tokens and Claude Code config survive container stop/start *and* container recreation.
+87
View File
@@ -471,6 +471,92 @@ When enabled, the host Docker socket is mounted into the container so Claude Cod
> Toggling this requires stopping and restarting the container to take effect.
### VPN Support
When enabled, the container is given the three things a VPN client needs to build a tunnel:
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
sysctl that WireGuard requires. This is **off by default**.
The `ip`, `wg` and `iptables` commands ship in the container image so there is something able to use
them. If your project's container was created from an older base image it will not have them, and
`wg` will simply not be found — **migrating the project onto the current base image** is what picks
them up. `sudo apt install iproute2 wireguard-tools iptables` works in the meantime, but lives in
the writable layer, so it is undone by a **Reset** and by a migration.
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
container's traffic to start leaving through a VPN is the most common misreading of what it does —
configuring a tunnel and routing traffic into it remains yours to do.
To make that second half easier, enabling this also installs a **`pia-vpn` skill** into the
container's `~/.claude/skills/`, so Claude Code can bring up a Private Internet Access tunnel over
WireGuard for you — ask it to connect the VPN and it will. The skill carries the parts that are
easy to get wrong (see the DNS note below), and it is removed again when you turn the setting off.
It needs your PIA credentials in `~/pia-creds`, two lines, username then password. If you use a
different provider, ignore it and set up your own client; nothing else depends on it.
Like the VPN tooling above, the skill ships in the container image, so a project whose container
predates it will not get one by toggling the setting — **migrate the project** and it appears; the
migration pre-flight lists it among what you would gain.
With the setting **off**, a client such as PIA or OpenVPN installs and its daemon starts normally,
but the connection attempt **hangs until it times out** — a default container has no tun device to open
and no permission to add an interface or a route, and most clients report that as a generic timeout
rather than a permissions error.
Things worth knowing:
- Tailscale is the exception: in its `--tun=userspace-networking` mode it needs neither the
capability nor the device, so leave this off if that is all you want.
- `NET_ADMIN` applies to the container's **own** network namespace — it cannot touch the host's
interfaces. It is not nothing, though: within that namespace anything in the container can set
promiscuous mode and add arbitrary addresses, routes and firewall rules on the Docker bridge it
shares with your other containers, and it can flush firewall rules that sandbox mode relies on.
Grant it per project, to projects that need it.
- The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is
the Linux VM, not your own machine. If it is missing, the container is created but fails to
**start**, with an error naming `/dev/net/tun` and pointing back at this setting.
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
- **No tunnel survives a restart.** The network namespace is built fresh every time the container
starts, and there is no service manager inside to reconnect anything. Leftover state under `/run`
makes it *look* like the tunnel is still configured — that directory is in the container's
writable layer, so it is simply still there after a stop/start, and `docker commit` carries it
into the snapshot that a recreation is built from. Either way the interface and its routes are
gone and traffic goes out your real address again, with no error and nothing visibly different.
Re-establish it after every start, and check rather than assume.
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your
resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own
subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures
it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is
`192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`,
which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is
total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`.
Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original
gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own,
or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own
address via the original gateway, or the tunnel's encrypted packets try to route through the
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
while DNS is broken — resolve a name instead.
- **Delete a client's key material when you tear a tunnel down.** Anything written under `/run` is
in the container's writable layer, and recreating or migrating the project runs `docker commit`
over it — so a WireGuard private key left there gets baked into the project's snapshot image and
copied forward from then on. This is not hypothetical; it has already happened here.
- **Strip the `DNS =` line from a provider's `.conf` before `wg-quick up`.** Every commercial
provider ships one, and `wg-quick` hands it to `resolvconf`, which is not installed — so it fails
at `resolvconf: command not found` and deletes the interface again. This happens before any
routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive
`wg` and `ip route` directly rather than going through `wg-quick`.
- **`wg-quick` full tunnels also need `xt_CONNMARK` from the host kernel.** Native Linux, Docker
Desktop for Mac and WSL2 kernels from 6.6 have it; older WSL2 kernels do not, and a container
cannot load one. There the answer is again to add the routes yourself with `ip route`, which
needs no firewall backend on any platform.
> This setting can only be changed when the container is stopped. Capabilities and devices are
> fixed when a container is created, so toggling it recreates the container on the next start.
> Recreation preserves the home and `.claude` volumes — it is not a Reset.
### Mission Control
Toggle **Mission Control** to integrate Flight Control — an AI-first development methodology bundled with Triple-C — into the project. When enabled:
@@ -1225,6 +1311,7 @@ The sandbox container (Ubuntu 24.04) comes pre-installed with:
| ruff | Latest | Python linter/formatter |
| Rust | Stable | Rust development (via rustup) |
| Docker CLI | Latest | Container management (when spawning is enabled) |
| iproute2, WireGuard tools, iptables | Latest | Building a tunnel (when VPN Support is enabled) |
| git | Latest | Version control |
| GitHub CLI (gh) | Latest | GitHub integration |
| AWS CLI | v2 | AWS services and Bedrock |
+74 -26
View File
@@ -194,11 +194,24 @@ impl BrowserTarget {
}
}
/// The `channel` a launch check must pass. `None` means the bundled build.
fn channel(self) -> Option<&'static str> {
/// Every `channel` a launch check must pass, comma-separated, where
/// `default` means "no channel — the bundled build".
///
/// Chromium is checked twice because the two consumers of this install do
/// not launch the same binary. A script calling `chromium.launch()` with
/// no channel gets `chromium-headless-shell`; the viewer reads
/// `~/.playwright/cli.config.json`, which pins channel
/// `chrome-for-testing`, and that resolves to the *full* `chromium-<rev>`
/// build — a separate download under the same `install chromium`.
///
/// Checking only the first is how a container reaches "verified" and then
/// fails in the pane with `Browser "chrome-for-testing" is not installed`.
/// Observed on a real project, where a stale `chromium-1217` satisfied the
/// headless-shell launch while the viewer wanted `chromium-1237`.
fn channels(self) -> &'static str {
match self {
Self::Chromium => None,
Self::Chrome => Some("chrome"),
Self::Chromium => "default,chrome-for-testing",
Self::Chrome => "chrome",
}
}
}
@@ -235,7 +248,7 @@ pub async fn install_packages(
&format!("Installing @playwright/cli into {}/node_modules…", INSTALL_DIR),
);
let mut step = npm_install(app, project_id, container_id, VIEWER_PACKAGE).await?;
let mut step = npm_install(app, project_id, container_id, &[VIEWER_PACKAGE]).await?;
if step.exit_code != 0 {
return Err(format!(
"npm couldn't install the viewer package in this container (exit {}).\n\nnpm said:\n{}",
@@ -246,9 +259,15 @@ pub async fn install_packages(
// Second, `playwright` at the version the viewer package pins — see
// `VIEWER_PACKAGE`. Installing it as `@latest` is what splits the tree.
//
// The viewer package is named *again* here. It is already installed, so
// this adds no work, but omitting it is what made npm prune it back out —
// see the note on `npm_install`. The pin can only be read after the first
// install has written the manifest, which is why this stays two commands
// rather than one.
let spec = pinned_playwright_spec(container_id).await;
emit_progress(app, project_id, &format!("Installing {}", spec));
let second = npm_install(app, project_id, container_id, &spec).await?;
let second = npm_install(app, project_id, container_id, &[VIEWER_PACKAGE, &spec]).await?;
if second.exit_code != 0 {
return Err(format!(
"npm couldn't install {} in this container (exit {}).\n\nnpm said:\n{}",
@@ -290,7 +309,7 @@ pub async fn install_packages(
})
}
/// One `npm install` of one spec, into [`INSTALL_DIR`], as `claude`.
/// One `npm install` of one or more specs, into [`INSTALL_DIR`], as `claude`.
///
/// `env VAR=… cmd` rather than an exec env: it keeps the one exec path in
/// `docker/exec.rs` untouched, and `env` is a real binary so no shell is
@@ -298,13 +317,24 @@ pub async fn install_packages(
/// has no postinstall (verified — `playwright@1.62.1` declares no `scripts` at
/// all), but if a future release brings the browser download back, this step
/// must stay small and the download must stay the step the user asked for.
///
/// **Every package that must survive has to appear in `specs`.** `--no-save`
/// in a directory with no `package.json` — which [`INSTALL_DIR`] is — leaves
/// npm with the command line as its only statement of what the tree should
/// contain, and npm ≥7 reconciles the tree against that on every run by
/// removing whatever it now considers extraneous. Installing `@playwright/cli`
/// and then installing `playwright` in a second command therefore *deletes the
/// first one*: verified in a container, `removed 3 packages`, leaving an empty
/// `node_modules/@playwright/` behind `playwright` and `playwright-core`. That
/// empty directory is why a fresh setup could report success and still leave
/// the pane saying `@playwright/cli` was not installed.
async fn npm_install(
app: &AppHandle,
project_id: &str,
container_id: &str,
spec: &str,
specs: &[&str],
) -> Result<StepResult, String> {
let cmd = vec![
let mut cmd = vec![
"env".to_string(),
"PLAYWRIGHT_SKIP_BROWSER_DOWNLOAD=1".to_string(),
"npm".to_string(),
@@ -313,8 +343,8 @@ async fn npm_install(
"--no-save".to_string(),
"--no-fund".to_string(),
"--no-audit".to_string(),
spec.to_string(),
];
cmd.extend(specs.iter().map(|s| s.to_string()));
run_step(
app,
project_id,
@@ -704,7 +734,7 @@ async fn verify_launch(
],
vec![
format!("TRIPLE_C_PW_DIR={}", dir),
format!("TRIPLE_C_PW_CHANNEL={}", target.channel().unwrap_or("")),
format!("TRIPLE_C_PW_CHANNELS={}", target.channels()),
format!("TRIPLE_C_PW_URL={}", REACHABILITY_URL),
],
);
@@ -795,27 +825,43 @@ fn parse_launch_output(output: &str) -> LaunchVerdict {
/// The launch check. One `argv` element, no newlines, same contract as the
/// detection probe.
///
/// Playwright leaves the Chromium sandbox disabled by default, which is what
/// makes this work in a container at all. The timeout exists so a browser that
/// hangs on a missing library still returns a verdict rather than sitting there
/// until the exec is torn down. The navigation is best-effort and never decides
/// `ok` — it exists to tell a TLS-intercepted network apart from a broken
/// install.
/// `chromiumSandbox` is set explicitly rather than left to Playwright's
/// default, so this check states the same thing the seeded
/// `cli.config.json` does instead of agreeing with it by coincidence. The
/// containers forbid unprivileged user namespaces, so a sandboxed Chromium
/// aborts on launch; nothing here should be able to drift back into testing a
/// configuration the viewer will not use.
///
/// Each channel in `TRIPLE_C_PW_CHANNELS` is launched in turn — see
/// [`BrowserTarget::channels`] for why Chromium needs two — and a failure
/// names the channel that failed, because "is not installed" is meaningless
/// without it. Only the last launch loads a page: the navigation is
/// best-effort, never decides `ok`, and exists to tell a TLS-intercepted
/// network apart from a broken install, so doing it once is enough.
///
/// The timeout exists so a browser that hangs on a missing library still
/// returns a verdict rather than sitting there until the exec is torn down.
const LAUNCH_PROBE: &str = concat!(
r#"const d=process.env.TRIPLE_C_PW_DIR,ch=process.env.TRIPLE_C_PW_CHANNEL||undefined,u=process.env.TRIPLE_C_PW_URL;"#,
r#"const d=process.env.TRIPLE_C_PW_DIR,chs=process.env.TRIPLE_C_PW_CHANNELS||"default",u=process.env.TRIPLE_C_PW_URL;"#,
r#"let done=false;const say=(ok,detail,nav)=>{if(done)return;done=true;"#,
r#"process.stdout.write("\n__TRIPLE_C_BROWSER_LAUNCH__"+JSON.stringify({ok,detail,nav:nav||null})+"\n");};"#,
r#"const one=(e)=>String((e&&e.message)||e).split("\n").slice(0,8).join(" | ");"#,
r#"const t=setTimeout(()=>{say(false,"the browser did not finish starting within 90s");process.exit(0);},90000);"#,
r#"(async()=>{let b=null;try{const {chromium}=require(d);b=await chromium.launch(ch?{channel:ch}:{});"#,
r#"let v="";try{v=b.version();}catch(e){}"#,
r#"let nav={ok:true,cert:false,detail:""};"#,
r#"(async()=>{let b=null,cur="";try{const {chromium}=require(d);"#,
r#"const list=chs.split(",").map(s=>s.trim()).filter(Boolean);"#,
r#"let v="",nav={ok:true,cert:false,detail:""};"#,
r#"for(let i=0;i<list.length;i++){cur=list[i];const c=cur==="default"?undefined:cur;"#,
r#"b=await chromium.launch(Object.assign({chromiumSandbox:false},c?{channel:c}:{}));"#,
r#"try{v=b.version();}catch(e){}"#,
r#"if(i===list.length-1){"#,
r#"try{const p=await b.newPage();await p.goto(u,{timeout:20000});}"#,
// A certificate failure is classified here, next to the message, because
// Chromium's wording is the only place the distinction exists.
r#"catch(e){const m=one(e);nav={ok:false,cert:/ERR_CERT|CERT_AUTHORITY|ERR_SSL|SSL_ERROR|self.signed/i.test(m),detail:m};}"#,
r#"await b.close();clearTimeout(t);say(true,v,nav);}"#,
r#"catch(e){clearTimeout(t);try{if(b)await b.close();}catch(e2){}say(false,one(e));}"#,
r#"catch(e){const m=one(e);nav={ok:false,cert:/ERR_CERT|CERT_AUTHORITY|ERR_SSL|SSL_ERROR|self.signed/i.test(m),detail:m};}}"#,
r#"await b.close();b=null;}"#,
r#"clearTimeout(t);say(true,v,nav);}"#,
r#"catch(e){clearTimeout(t);try{if(b)await b.close();}catch(e2){}"#,
r#"say(false,(cur&&cur!=="default"?"channel "+cur+": ":"")+one(e));}"#,
r#"process.exit(0);})();"#,
);
@@ -984,8 +1030,10 @@ mod tests {
// `@playwright/mcp` asks for the chrome channel specifically, so the UI
// must be able to say so.
assert!(BrowserTarget::Chrome.needed_for().contains("@playwright/mcp"));
assert_eq!(BrowserTarget::Chrome.channel(), Some("chrome"));
assert_eq!(BrowserTarget::Chromium.channel(), None);
assert_eq!(BrowserTarget::Chrome.channels(), "chrome");
// Both of Chromium's consumers, or the check passes for a browser the
// viewer cannot open — see `channels`.
assert_eq!(BrowserTarget::Chromium.channels(), "default,chrome-for-testing");
// And a size, before the click, for both.
for t in [BrowserTarget::Chromium, BrowserTarget::Chrome] {
assert!(t.download_note().to_lowercase().contains("mb"), "{:?}", t);
@@ -73,6 +73,25 @@ use crate::AppState;
/// Report how far behind the current base image a project's container is, and
/// what migrating it would actually carry across.
///
/// Choose the recorded lineage from the two places it can be written, most
/// authoritative first: the live container's label, then the snapshot image's.
///
/// **An empty label is absence, not an answer.** `create_container` always
/// writes `triple-c.base-image-id`, even when the value is unknown — that is
/// deliberate, because Docker merges an image's labels into a container's and
/// an inherited value would otherwise ride a snapshot forever. The consequence
/// is that `Some("")` is the *common* reading from a container whose lineage
/// was never established, so treating it as an answer silently skips the
/// snapshot, which may well have recorded a real one.
fn pick_recorded_lineage(
from_container: Option<String>,
from_snapshot: Option<String>,
) -> Option<String> {
from_container
.filter(|v| !v.is_empty())
.or_else(|| from_snapshot.filter(|v| !v.is_empty()))
}
/// Read-only. Runs two filesystem probes (~3 s each) and is therefore meant to
/// be called on demand, not polled.
#[tauri::command]
@@ -98,20 +117,24 @@ pub async fn get_container_staleness(
// Lineage, most authoritative source first: the live container's label,
// then the snapshot image's. Both are written by `create_container` and
// propagated onto the snapshot by `docker commit`.
// Each source is filtered for emptiness *before* it is allowed to satisfy
// the lookup. `create_container` always writes this label, even when the
// value is unknown — deliberately, so an inherited image label cannot ride
// a snapshot forever — which means the container's copy is very often
// `Some("")`. Filtering only the final result let that empty string count
// as an answer and skip the snapshot entirely, so a snapshot that *did*
// record a lineage was never consulted and the project reported "unknown"
// with the information sitting one lookup away.
let container_id = docker::find_existing_container(&project).await.unwrap_or(None);
let recorded = match &container_id {
let from_container = match &container_id {
Some(id) => container_label(id, mig::LABEL_BASE_IMAGE_ID).await,
None => None,
}
.or_else(|| None);
let recorded = match recorded {
Some(v) => Some(v),
None => mig::image_labels(&snapshot_image)
};
let from_snapshot = mig::image_labels(&snapshot_image)
.await
.get(mig::LABEL_BASE_IMAGE_ID)
.cloned(),
}
.filter(|v| !v.is_empty());
.cloned();
let recorded = pick_recorded_lineage(from_container, from_snapshot);
out.base_image_id = recorded.clone();
out.known = recorded.is_some();
@@ -1671,6 +1694,32 @@ fn summarize(
mod tests {
use super::*;
#[test]
fn an_empty_lineage_label_is_absence_and_falls_through_to_the_snapshot() {
let some = |s: &str| Some(s.to_string());
// The regression: the container always carries the label, so an
// unknown lineage reads as `Some("")`. Letting that satisfy the lookup
// skipped a snapshot that had recorded the real thing.
assert_eq!(
pick_recorded_lineage(some(""), some("sha256:base")),
some("sha256:base")
);
// Ordinary precedence still holds: the container wins when it has one.
assert_eq!(
pick_recorded_lineage(some("sha256:container"), some("sha256:snapshot")),
some("sha256:container")
);
assert_eq!(pick_recorded_lineage(None, some("sha256:snap")), some("sha256:snap"));
// Genuinely unknown stays unknown — "probe instead", never a lineage
// invented to make the comparison succeed.
assert_eq!(pick_recorded_lineage(None, None), None);
assert_eq!(pick_recorded_lineage(some(""), some("")), None);
assert_eq!(pick_recorded_lineage(some(""), None), None);
}
#[test]
fn byte_sizes_read_the_way_a_disk_warning_should() {
assert_eq!(human_bytes(512), "512 B");
+272 -2
View File
@@ -233,6 +233,7 @@ const RESERVED_ENV_EXACT: &[&str] = &[
"MCP_SERVERS_JSON",
"CLAUDE_CODE_SETTINGS_JSON",
"MISSION_CONTROL_ENABLED",
"VPN_SUPPORT_ENABLED",
"TRIPLE_C_PERMISSION_MODE",
CLAUDE_OAUTH_TOKEN_ENV,
// The model-alias vars are already covered by the `ANTHROPIC_` prefix
@@ -798,6 +799,128 @@ async fn resolve_base_image_id(image_name: &str, base_image_name: &str) -> Strin
.unwrap_or_default()
}
/// The `/dev/net/tun` character device, as it is named on both sides.
const TUN_DEVICE: &str = "/dev/net/tun";
/// The `HostConfig` fields "VPN support" contributes: `CapAdd`, `Devices`,
/// `Sysctls` — in that order.
type VpnHostConfigParts = (
Option<Vec<String>>,
Option<Vec<bollard::models::DeviceMapping>>,
Option<HashMap<String, String>>,
);
/// The three host-config pieces a VPN client needs, or all-`None` when the
/// project has not opted in.
///
/// Returned as a triple rather than set inline so the exact shape is unit
/// testable — a container is created once, by a very long async function, and a
/// silently-dropped capability looks identical to a VPN server that is simply
/// unreachable.
///
/// All three are required together and each fails differently on its own:
/// * **`CAP_NET_ADMIN`** — without it the client cannot create an interface or
/// write a route. Docker's default bounding set grants `net_raw` but not
/// `net_admin`, which is why a client can ping but never connect.
/// * **`/dev/net/tun`** — the device is absent from a default container, so
/// there is nothing to open even with the capability. It is passed through
/// from the host rather than `mknod`-ed inside, so the kernel's `tun` module
/// backs it.
/// * **`net.ipv4.conf.all.src_valid_mark`** — WireGuard's own `wg-quick` sets
/// this, and cannot from inside a container (`/proc/sys` is read-only), so
/// its handshake packets are dropped by reverse-path filtering. Harmless for
/// OpenVPN-based clients, so it is set unconditionally with the rest.
///
/// What it costs, stated accurately: Docker does not enable user-namespace
/// remapping by default, so this is a real `CAP_NET_ADMIN` in the *initial*
/// user namespace and only the **network** namespace confines it. It cannot
/// touch the host's interfaces, but within its own namespace it can set
/// promiscuous mode and add arbitrary addresses, routes and NAT rules on the
/// shared `docker0` L2 segment — which puts sibling containers (the LiteLLM
/// gateway among them) within reach of ARP spoofing, and lets netlink trigger
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
/// Hence opt-in, per project, rather than on for everyone.
/// The env var `entrypoint.sh` installs and removes the `pia-vpn` skill from.
///
/// **Emitted either way, never omitted.** `~/.claude` is a persisted volume, so
/// turning the toggle off has to actively tell entrypoint to remove a skill an
/// earlier run left there, and an absent variable cannot say that. It is also
/// what stops a `=1` baked into a snapshot by `docker commit` from outliving
/// the setting — the explicit `=0` overwrites it.
///
/// Extracted for the same reason as [`vpn_host_config`]: the emitting code sits
/// in a very long function where a dropped or inverted value is invisible, and
/// `MISSION_CONTROL_ENABLED` twenty lines above shows the failure this avoids —
/// it is pushed only when true, so a snapshot's baked `=1` survives the toggle
/// going off.
fn vpn_env_var(enabled: bool) -> String {
format!("VPN_SUPPORT_ENABLED={}", u8::from(enabled))
}
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
if !enabled {
return (None, None, None);
}
let devices = vec![bollard::models::DeviceMapping {
path_on_host: Some(TUN_DEVICE.to_string()),
path_in_container: Some(TUN_DEVICE.to_string()),
cgroup_permissions: Some("rwm".to_string()),
}];
let sysctls = HashMap::from([(
"net.ipv4.conf.all.src_valid_mark".to_string(),
"1".to_string(),
)]);
(
Some(vec!["NET_ADMIN".to_string()]),
Some(devices),
Some(sysctls),
)
}
/// Turn the daemon's device-passthrough failure into an explanation.
///
/// **This fires on `start`, not `create`.** Verified against Docker 29.7:
/// `docker create --device /dev/does-not-exist` succeeds and prints an id; the
/// device is only resolved when runc builds the container, so the failure lands
/// on the *next* call. Sysctls validate at the same point. Anything that
/// inspects only the create path will never see it — which is why both paths
/// route through here and the tests exercise the start-side string.
///
/// Unmapped, this reads as `Failed to start container: Docker responded with
/// status code 500: error gathering device information while adding custom
/// device "/dev/net/tun": no such file or directory` — a path the user will go
/// looking for on the wrong machine, since with Docker Desktop the relevant
/// host is the Linux VM rather than their own, and with nothing pointing back
/// at the switch that caused it.
///
/// Deliberately not gated on `vpn_support_enabled`: nothing else in Triple-C
/// ever asks for a device, so an error naming `/dev/net/tun` can only have come
/// from a container created with the switch on. That keeps the check usable
/// from [`start_container`], which has a container id and no project.
fn explain_container_failure(action: &str, err: &str) -> String {
let device_missing = err.contains(TUN_DEVICE)
&& (err.contains("no such file or directory")
|| err.contains("No such file or directory")
|| err.contains("error gathering device information"));
if device_missing {
return format!(
"Failed to {} container: the Docker host has no {} device, which \
\"VPN support\" requires. The host kernel needs the `tun` module \
loaded (on Docker Desktop that is the Linux VM, not your own \
machine). Turn VPN support off in Config → Runtime to start this \
project without it. Original error: {}",
action, TUN_DEVICE, err
);
}
format!("Failed to {} container: {}", action, err)
}
pub async fn create_container(
project: &Project,
docker_socket_path: &str,
@@ -1170,6 +1293,8 @@ pub async fn create_container(
env_vars.push("MISSION_CONTROL_ENABLED=1".to_string());
}
env_vars.push(vpn_env_var(project.vpn_support_enabled));
// Permission mode — read by triple-c-task-runner for scheduled (headless)
// Claude Code runs. Interactive terminals get the flags directly instead.
env_vars.push(format!(
@@ -1375,6 +1500,13 @@ pub async fn create_container(
labels.insert("triple-c.image".to_string(), image_name.to_string());
labels.insert("triple-c.timezone".to_string(), timezone.unwrap_or("").to_string());
labels.insert("triple-c.mission-control".to_string(), project.mission_control_enabled.to_string());
// Capabilities, devices and sysctls are fixed at creation, so this is
// container state and gets the label-and-compare treatment. Written
// unconditionally (`false`, not omitted) because `docker commit` copies
// container labels onto the snapshot image: a `true` stamped once would
// otherwise ride that snapshot into every future container and make the
// switch impossible to turn back off.
labels.insert("triple-c.vpn-support".to_string(), project.vpn_support_enabled.to_string());
labels.insert("triple-c.permission-mode".to_string(),
project.effective_permission_mode().as_env_value().to_string());
labels.insert("triple-c.custom-env-fingerprint".to_string(), custom_env_fingerprint.clone());
@@ -1443,10 +1575,15 @@ pub async fn create_container(
labels.insert((*key).to_string(), (*value).to_string());
}
let (cap_add, devices, sysctls) = vpn_host_config(project.vpn_support_enabled);
let host_config = HostConfig {
mounts: Some(mounts),
port_bindings: if port_bindings.is_empty() { None } else { Some(port_bindings) },
init: Some(true),
cap_add,
devices,
sysctls,
..Default::default()
};
@@ -1476,7 +1613,7 @@ pub async fn create_container(
let response = docker
.create_container(Some(options), config)
.await
.map_err(|e| format!("Failed to create container: {}", e))?;
.map_err(|e| explain_container_failure("create", &e.to_string()))?;
Ok(response.id)
}
@@ -1486,7 +1623,7 @@ pub async fn start_container(container_id: &str) -> Result<(), String> {
docker
.start_container(container_id, None::<StartContainerOptions<String>>)
.await
.map_err(|e| format!("Failed to start container: {}", e))
.map_err(|e| explain_container_failure("start", &e.to_string()))
}
pub async fn stop_container(container_id: &str) -> Result<(), String> {
@@ -2367,6 +2504,19 @@ pub async fn container_needs_recreation(
return Ok(true);
}
// ── VPN support (NET_ADMIN + /dev/net/tun + sysctl) ───────────────────
// A container's capabilities, devices and sysctls are set at creation and
// cannot be changed on a running or stopped container, so recreation is the
// only way a toggle here takes effect. A missing label means the container
// predates the feature, which is the same thing as having it off — so
// existing projects are not churned until someone actually turns it on.
let expected_vpn = project.vpn_support_enabled.to_string();
let container_vpn = get_label("triple-c.vpn-support").unwrap_or_else(|| "false".to_string());
if container_vpn != expected_vpn {
log::info!("VPN support mismatch (container={:?}, expected={:?})", container_vpn, expected_vpn);
return Ok(true);
}
// ── Permission mode ────────────────────────────────────────────────────
// The mode is injected as the TRIPLE_C_PERMISSION_MODE env var, and
// container env can only change by recreating the container. A missing
@@ -2616,6 +2766,126 @@ mod tests {
assert_eq!(fp, "");
}
#[test]
fn vpn_support_off_touches_nothing_in_the_host_config() {
// The default must stay byte-identical to a container created before the
// feature existed, or every project recreates on the next start.
let (cap_add, devices, sysctls) = vpn_host_config(false);
assert_eq!(cap_add, None);
assert_eq!(devices, None);
assert_eq!(sysctls, None);
}
#[test]
fn vpn_support_on_grants_all_three_pieces() {
// Each is useless without the others — a client with the capability but
// no device, or the device but no capability, still times out — so this
// asserts the whole set rather than any one of them.
let (cap_add, devices, sysctls) = vpn_host_config(true);
assert_eq!(cap_add, Some(vec!["NET_ADMIN".to_string()]));
let devices = devices.expect("the tun device must be passed through");
assert_eq!(devices.len(), 1);
assert_eq!(devices[0].path_on_host.as_deref(), Some(TUN_DEVICE));
assert_eq!(devices[0].path_in_container.as_deref(), Some(TUN_DEVICE));
assert_eq!(devices[0].cgroup_permissions.as_deref(), Some("rwm"));
assert_eq!(
sysctls
.expect("wireguard needs src_valid_mark")
.get("net.ipv4.conf.all.src_valid_mark")
.map(String::as_str),
Some("1")
);
}
#[test]
fn vpn_support_never_grants_more_than_net_admin() {
// NET_ADMIN is already a step out of the sandbox. Anything else added
// here (SYS_ADMIN, or a blanket privileged flag) would be a much larger
// one, so pin the set.
let (cap_add, _, _) = vpn_host_config(true);
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
}
#[test]
fn the_vpn_skill_flag_is_emitted_either_way_never_omitted() {
// The whole removal path depends on this. If `false` ever became "emit
// nothing", a project that had the toggle on would keep the skill
// forever: the container recreates from a snapshot whose baked
// VPN_SUPPORT_ENABLED=1 would then go unchallenged, and entrypoint
// would reinstall a skill for a capability the container no longer has.
assert_eq!(vpn_env_var(true), "VPN_SUPPORT_ENABLED=1");
assert_eq!(vpn_env_var(false), "VPN_SUPPORT_ENABLED=0");
}
#[test]
fn the_vpn_skill_flag_is_reserved_from_custom_env() {
// entrypoint.sh installs and removes the pia-vpn skill from this
// variable. A custom env var of the same name would let a project claim
// the skill without the capability behind it — or keep it after the
// toggle is off — so it has to be unsettable like the others.
assert!(is_reserved_env_key("VPN_SUPPORT_ENABLED"));
assert!(is_reserved_env_key("vpn_support_enabled"));
assert_eq!(
compute_env_fingerprint(&[EnvVar {
key: "VPN_SUPPORT_ENABLED".to_string(),
value: "1".to_string(),
}]),
""
);
}
/// What bollard actually hands us when a tun-less host rejects the device.
///
/// Captured verbatim from Docker 29.7: `docker create` with a missing
/// device **succeeds**, and this arrives from the subsequent `start`.
/// `DockerResponseServerError`'s Display is
/// `"Docker responded with status code {code}: {message}"` with the
/// daemon's message unaltered.
const REAL_TUN_ERROR: &str = "Docker responded with status code 500: error \
gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
#[test]
fn a_missing_tun_device_is_explained_on_the_path_that_actually_fails() {
// The start path is the one that matters: the daemon defers device
// resolution to runc, so create returns an id on a host with no tun
// module and only start fails. A version of this that checked create
// alone would be dead code.
let msg = explain_container_failure("start", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to start container:"), "{}", msg);
assert!(msg.contains("VPN support"), "should name the switch: {}", msg);
assert!(msg.contains("tun` module"), "should name the cause: {}", msg);
assert!(msg.contains("Config → Runtime"), "should say where to fix it: {}", msg);
assert!(msg.contains(REAL_TUN_ERROR), "should keep the original: {}", msg);
}
#[test]
fn the_same_explanation_covers_create_if_the_daemon_ever_checks_earlier() {
// Belt and braces — older and future daemons may validate at create.
let msg = explain_container_failure("create", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to create container:"), "{}", msg);
assert!(msg.contains("VPN support"), "{}", msg);
}
#[test]
fn unrelated_failures_are_left_alone() {
for (action, err) in [
("create", "Conflict. The container name \"/triple-c-x\" is already in use"),
("start", "Docker responded with status code 404: No such container"),
("start", "error gathering device information while adding custom device \"/dev/dri/card0\""),
] {
assert_eq!(
explain_container_failure(action, err),
format!("Failed to {} container: {}", action, err),
"{} should pass through untouched",
err
);
}
}
#[test]
fn the_orphan_sweep_only_ever_looks_at_our_own_untagged_images() {
// Both conditions are load-bearing. Without `dangling` the sweep would
+2
View File
@@ -135,6 +135,8 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
("/opt/mission-control", "Mission Control (Flight Control)"),
("/usr/bin/wg", "VPN tooling (WireGuard, for the VPN Support toggle)"),
("/opt/triple-c-skills", "Bundled skills (PIA VPN, for the VPN Support toggle)"),
];
/// Headroom demanded on Docker's storage backend on top of the measured
+17
View File
@@ -145,6 +145,22 @@ pub struct Project {
/// container-recreation label.
#[serde(default)]
pub browser_view_enabled: bool,
/// Grant the container what a VPN client needs to build a tunnel:
/// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard
/// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard,
/// OpenVPN) installs and runs but its connection attempt hangs until it
/// times out, because it cannot create the tunnel interface or touch the
/// routing table.
///
/// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the
/// container reconfigure its own network stack, which reaches further than
/// it sounds — see `vpn_host_config` for what it does and does not confer.
/// Unlike `auth_bridge_enabled` this *is*
/// container state, so it carries a `triple-c.vpn-support` label and is
/// compared in `container_needs_recreation` — capabilities and devices are
/// fixed at creation and can only change by recreating the container.
#[serde(default)]
pub vpn_support_enabled: bool,
/// Use the shared, long-lived Claude Code OAuth token (from
/// `claude setup-token`, held in the OS keychain) for this project instead
/// of requiring its own `claude login`. Only consulted when `backend` is
@@ -366,6 +382,7 @@ impl Project {
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: default_use_shared_auth_token(),
full_permissions: false,
permission_mode: None,
@@ -127,6 +127,46 @@ describe("ContainerMigrationBanner", () => {
expect(container).toBeEmptyDOMElement();
});
it("speaks up when an unlabelled container could not be probed at all", () => {
// The probe is the only signal a container with no lineage label has. If
// it fails and the banner stays silent, that is indistinguishable from
// "up to date" — the exact reading that let an out-of-date project go
// unnoticed indefinitely.
renderBanner(
migration({
staleness: {
...FRESH,
known: false,
stale: false,
probe_error: "output exceeded the inspection limit",
},
probeSettled: false,
}),
);
expect(
screen.getByText(/Container base could not be checked/i),
).toBeInTheDocument();
expect(
screen.getByText(/output exceeded the inspection limit/i),
).toBeInTheDocument();
// And it must not pose as a finding about the container itself.
expect(
screen.queryByText(/Container is missing things/i),
).not.toBeInTheDocument();
});
it("stays quiet when a labelled container's probe fails but its lineage is current", () => {
// `known` means the version comparison already answered the question, so
// a failed probe is not grounds to raise anything.
const { container } = renderBanner(
migration({
staleness: { ...FRESH, probe_error: "could not exec in the container" },
probeSettled: false,
}),
);
expect(container).toBeEmptyDOMElement();
});
it("disables the action and explains why while the container is running", () => {
renderBanner(migration({ staleness: STALE }), false);
expect(
@@ -131,7 +131,15 @@ export default function ContainerMigrationBanner({
const probeFoundGaps =
!staleness.known &&
(staleness.missing_features.length > 0 || staleness.missing_paths.length > 0);
if (!staleness.stale && !probeFoundGaps) return null;
// The probe is the *only* signal a container with no lineage label has, so
// when it fails there is nothing left to be quiet about. Staying silent here
// is indistinguishable from "everything is fine" — and it is the likeliest
// outcome for the oldest, largest projects, whose manifests are the ones apt
// to exceed the inspection limit. Say that the check did not run instead.
const probeUnavailable = !staleness.known && !!staleness.probe_error;
if (!staleness.stale && !probeFoundGaps && !probeUnavailable) return null;
const snapshot = formatSnapshotDate(staleness.snapshot_created_at);
const features = joinFeatures(staleness.missing_features);
@@ -139,15 +147,24 @@ export default function ContainerMigrationBanner({
return (
<section
className={`${SHELL} border-[var(--warning)]/40 bg-[var(--warning-muted)]`}
aria-label="Container base is out of date"
aria-label={
probeUnavailable
? "Container base could not be checked"
: "Container base is out of date"
}
>
<div className="flex items-start justify-between gap-3">
<div className="min-w-0 space-y-1">
<StatusIndicator
tone="error"
// A check that could not run is not a finding: it gets the
// "unresolved" tone rather than the one that says something is
// wrong with the container.
tone={probeUnavailable ? "unknown" : "error"}
label={
staleness.known
? "Container base is out of date"
: probeUnavailable
? "Container base could not be checked"
: "Container is missing things the current base ships"
}
className="text-[13px] font-semibold"
@@ -158,6 +175,8 @@ export default function ContainerMigrationBanner({
? snapshot
? `Running on a saved image from ${snapshot}.`
: "Running on a saved image older than the current base."
: probeUnavailable
? "This container predates base-image tracking, so probing it is the only way to tell whether it is behind — and that did not complete."
: "This container predates base-image tracking, so it was probed directly."}
</p>
@@ -120,6 +120,15 @@ export default function OverviewTab({
{project.mission_control_enabled ? "ON" : "OFF"}
</span>
</span>
{/* Only when granted. It is off for nearly every project and an
always-present "VPN OFF" would be noise, but where it *is* on the
container holds NET_ADMIN, which is worth seeing at a glance. */}
{project.vpn_support_enabled && (
<span className="text-[var(--text-secondary)]">
VPN support{" "}
<span className="text-[var(--text-primary)] font-medium">ON</span>
</span>
)}
<button
type="button"
onClick={() => onOpenTab("config")}
@@ -0,0 +1,94 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
import { render, screen, fireEvent } from "@testing-library/react";
import RuntimeSection from "./RuntimeSection";
import type { Project } from "../../../../lib/types";
const baseProject: Project = {
id: "p1",
name: "api-server",
paths: [{ host_path: "/src/api", mount_name: "api" }],
container_id: null,
status: "stopped",
backend: "anthropic",
bedrock_config: null,
ollama_config: null,
llamacpp_config: null,
openai_compatible_config: null,
allow_docker_access: false,
sandbox_mode_enabled: true,
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: true,
full_permissions: false,
permission_mode: null,
ssh_key_path: null,
ca_cert_path: null,
git_token: null,
git_user_name: null,
git_user_email: null,
custom_env_vars: [],
port_mappings: [],
claude_instructions: null,
claude_code_settings: null,
renamed_session_names: {},
created_at: "2026-01-01T00:00:00Z",
updated_at: "2026-01-01T00:00:00Z",
};
const VPN = "VPN support";
const save = vi.fn().mockResolvedValue(true);
function renderSection(over: Partial<Project> = {}, disabled = false) {
return render(
<RuntimeSection
project={{ ...baseProject, ...over }}
save={save}
disabled={disabled}
disabledReason="Container must be stopped to change this setting."
/>,
);
}
describe("RuntimeSection — VPN support toggle", () => {
beforeEach(() => vi.clearAllMocks());
it("saves only the VPN flag when switched on", () => {
renderSection();
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: true });
});
it("saves the flag off again, rather than dropping the key", () => {
// Off has to be written explicitly: the container carries a
// `triple-c.vpn-support` label either way, and an absent value would leave
// the capability granted.
renderSection({ vpn_support_enabled: true });
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: false });
});
it("reflects the project's current state", () => {
renderSection({ vpn_support_enabled: true });
expect(screen.getByRole("switch", { name: VPN })).toBeChecked();
});
it("cannot be changed while the container is running", () => {
// Capabilities and devices are fixed at creation, so this setting is gated
// on the container being stopped along with the rest of the tab.
renderSection({}, true);
const toggle = screen.getByRole("switch", { name: VPN });
expect(toggle).toBeDisabled();
fireEvent.click(toggle);
expect(save).not.toHaveBeenCalled();
});
it("warns that the change recreates the container", () => {
renderSection();
expect(
screen.getByText(/recreates the container on its next start/i),
).toBeInTheDocument();
});
});
@@ -57,6 +57,19 @@ export default function RuntimeSection({
}
/>
<SwitchRow
label="VPN support"
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched. Changing this recreates the container on its next start — the home and .claude volumes are preserved."
control={
<Toggle
label="VPN support"
checked={project.vpn_support_enabled}
disabled={disabled}
onChange={(v) => save({ vpn_support_enabled: v })}
/>
}
/>
<SwitchRow
label="Mission Control"
hint="A web dashboard for monitoring and managing Claude sessions remotely."
+4
View File
@@ -33,6 +33,10 @@ export interface Project {
auth_bridge_enabled: boolean;
/** Opt in to the browser-view pane. Host-side only, like `auth_bridge_enabled`. */
browser_view_enabled: boolean;
/** Grant NET_ADMIN, /dev/net/tun and the WireGuard `src_valid_mark` sysctl so
* a VPN client inside the container can build a tunnel. Unlike the two flags
* above this is container state changing it recreates the container. */
vpn_support_enabled: boolean;
/** Use the shared long-lived Claude Code token (from `claude setup-token`,
* held in the OS keychain) instead of this project's own `claude login`.
* Defaults to true; only applies when `backend` is "anthropic" and a token
+91
View File
@@ -34,6 +34,9 @@ RUN for i in 1 2 3 4 5; do \
cron \
bubblewrap \
socat \
iproute2 \
wireguard-tools \
iptables \
&& rm -rf /var/lib/apt/lists/*
# `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither
@@ -42,6 +45,76 @@ RUN for i in 1 2 3 4 5; do \
# corporate CA, no matter what the system trust store says. entrypoint.sh
# degrades to a warning if it is ever missing.
# `iproute2`, `wireguard-tools` and `iptables` above are what the VPN support
# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a
# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to
# add a route, and without `wg` no way to build the tunnel those two exist to
# serve — a capability with nothing able to use it.
#
# They are baked rather than left to a runtime `apt-get install` for the same
# reason as the Playwright libraries below: the writable layer is re-paid after
# every Reset and lost on base-image migration. A hand-installed `wg` therefore
# works right up until an upgrade, then disappears and takes the tunnel with it
# — silently, since a VPN that fails to come up looks exactly like one that was
# never started.
#
# Measured against the *current base image*, since a bare ubuntu:24.04 also
# pulls libelf1t64 and netbase, which this base already has, and so over-reports
# by ~258 kB: **+12 packages, 7,203 kB on amd64**. The same set on arm64 is
# ~14.4 MB — the package list is identical on both arches, the binaries are
# simply larger (measured as 14.7 MB on arm64 ubuntu:24.04, less that 258 kB).
#
# ## Why `iptables`, and not `nftables`
#
# `wireguard-tools` declares `Recommends: nftables | iptables`, which the
# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s
# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e.
# every stock full-tunnel config every provider hands out — and it shells out to
# a firewall backend with no `type -p` guard. Measured with neither installed:
#
# [#] iptables-restore -n
# /usr/bin/wg-quick: line 32: iptables-restore: command not found
# wg-quick EXIT=127
#
# `nftables` looks like the better pick — wg-quick prefers it, it is first in
# that Recommends, it is half the size — and it is the wrong one. wg-quick picks
# nft *unconditionally* when present (`if type -p nft`, line 241), so installing
# it makes the iptables path unreachable; and its nft ruleset needs three
# expression families where the iptables path needs one. Isolating them on a
# WSL2 host, the two connmark rules install fine and this is what fails:
#
# nft add rule ... fib saddr type != local drop
# Error: Could not process rule: No such file or directory
# ^^^^^^^^^^^^^^ needs nft_fib_ipv4
#
# That matters because of how the two hosts we ship to are configured. From
# LinuxKit's kernel config — Docker Desktop for Mac, identical on both arches:
#
# CONFIG_NETFILTER_XT_CONNMARK=y <- the iptables path works
# # CONFIG_NFT_FIB_IPV4 is not set <- the nft path does not
#
# So shipping `nftables` would forfeit the platform it was meant to fix. With
# `iptables`, full tunnels work on native Linux, on Docker Desktop for Mac, and
# on WSL2 kernels from 6.6 (which added xt_CONNMARK as a module). Only WSL2
# older than that is left out, and nothing installable here changes it — the way
# out there is to add the routes with `ip route` instead of using `wg-quick`,
# which is what the pia-vpn skill does on every platform.
#
# ## What this still does not fix
#
# `wireguard-tools` only *Suggests* `openresolv | resolvconf`, so neither is
# installed, and every provider's stock config carries a `DNS =` line. That
# fails in `set_dns()`, *before* the firewall step, so it takes split tunnels
# down too:
#
# [#] resolvconf -a wg0 -m 0 -x
# /usr/bin/wg-quick: line 32: resolvconf: command not found
#
# Deliberately not fixed here: `openresolv` has no installation candidate on
# noble, and `resolvconf` resolves only by pulling in systemd-resolved — a
# resolver daemon and systemd units, into a container with no systemd. Strip the
# `DNS =` line and set the resolver another way. Documented in HOW-TO-USE.md.
# Remove default ubuntu user to free UID 1000 for host-user remapping
RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \
&& if getent group ubuntu >/dev/null 2>&1; then groupdel ubuntu 2>/dev/null || true; fi
@@ -314,10 +387,28 @@ RUN chmod +x /usr/local/bin/triple-c-sso-refresh
COPY mission-control /opt/mission-control
# Skills that ship with a Triple-C feature rather than with Mission Control.
# entrypoint.sh installs them into ~/.claude/skills/ when the feature that owns
# them is enabled, and removes them when it is not — a skill telling an agent to
# build a tunnel in a container that no longer has CAP_NET_ADMIN is worse than
# no skill at all. Staged in /opt because ~/.claude is a volume mount: an image
# copy underneath it would be masked from the project's first start onward.
COPY skills /opt/triple-c-skills
# `find`, not a `*/*.sh` glob: the glob fails the build the day a skill ships
# without a script, which is a legitimate thing for a skill to do.
RUN find /opt/triple-c-skills -name '*.sh' -exec chmod +x {} +
COPY entrypoint.sh /usr/local/bin/entrypoint.sh
RUN chmod +x /usr/local/bin/entrypoint.sh
COPY triple-c-scheduler /usr/local/bin/triple-c-scheduler
RUN chmod +x /usr/local/bin/triple-c-scheduler
# Lives in /usr/local/bin rather than under /home/claude on purpose: the home
# directory is the mount point of the project's home volume, so an image copy
# of it is masked after the project's first start and can never be updated
# again. /usr/local/bin rides the snapshot and is replaced on migration, which
# is what lets a fix to this script reach an existing project at all.
COPY triple-c-playwright-heal /usr/local/bin/triple-c-playwright-heal
RUN chmod +x /usr/local/bin/triple-c-playwright-heal
COPY triple-c-task-runner /usr/local/bin/triple-c-task-runner
RUN chmod +x /usr/local/bin/triple-c-task-runner
+84
View File
@@ -338,6 +338,64 @@ if [ "$MISSION_CONTROL_ENABLED" = "1" ]; then
unset MISSION_CONTROL_ENABLED
fi
# ── Feature skills ──────────────────────────────────────────────────────────
# Skills owned by a Triple-C feature rather than by Mission Control. Installed
# when the feature is on, removed when it is off: ~/.claude is a persisted
# volume, so a skill left behind after its feature is disabled would keep
# telling an agent to use a capability the container no longer has.
#
# Copied on every start rather than only when absent, so a fix to a skill
# reaches projects that already have the old copy. Local edits under these
# directories do not survive — treat /opt/triple-c-skills as the source.
#
# The source lives in the *base image*, so a project whose container predates it
# recreates from its own snapshot and has no /opt/triple-c-skills to copy from.
# That case says so rather than returning silently: the toggle is on, the
# capability is there, and the skill simply never appears — which is impossible
# to work out from the outside.
install_feature_skill() {
local _name="$1"
local _enabled="$2"
local _src="/opt/triple-c-skills/$1"
local _dest="/home/claude/.claude/skills/$1"
# Reject anything that is not a plain directory name. The disabled branch
# `rm -rf`s $_dest under a *persisted volume*, so a blank name would take the
# whole skills directory (Mission Control's included) and `../x` would escape
# it entirely. Only the literal `pia-vpn` is passed today; this is so that
# stays true.
case "$_name" in
''|*/*|.*) echo "entrypoint: install_feature_skill: bad skill name '$_name'"; return 1 ;;
esac
if [ "$_enabled" = "1" ]; then
if [ ! -d "$_src" ]; then
echo "entrypoint: $_name skill unavailable — this container's base image predates it; migrate the project to get it"
return 0
fi
# Checked, not assumed: with no `set -e` in this script every step here
# can fail (full volume, read-only mount, a file where the directory
# should be) and the success line would still print.
mkdir -p /home/claude/.claude/skills || {
echo "entrypoint: $_name skill install FAILED (cannot create ~/.claude/skills)"; return 1; }
# Not just $_dest: when Mission Control is off nothing else creates the
# parent, so root would own it and `claude` could not add a skill there.
chown claude:claude /home/claude/.claude/skills
rm -rf "$_dest"
cp -r "$_src" "$_dest" || {
echo "entrypoint: $_name skill install FAILED (copy from $_src)"; return 1; }
chown -R claude:claude "$_dest"
echo "entrypoint: $_name skill installed to ~/.claude/skills/"
elif [ -e "$_dest" ] || [ -L "$_dest" ]; then
# -e/-L rather than -d: a leftover *file* at that path must go too.
rm -rf "$_dest"
echo "entrypoint: $_name skill removed (feature disabled)"
fi
}
install_feature_skill pia-vpn "${VPN_SUPPORT_ENABLED:-0}"
unset VPN_SUPPORT_ENABLED
# ── Claude Code settings ────────────────────────────────────────────────────
# Merge Claude Code settings into ~/.claude/settings.json (preserves existing
# keys). Creates the file if it doesn't exist. These control TUI mode, effort
@@ -425,6 +483,32 @@ if [ -x /usr/local/bin/triple-c-open ]; then
export BROWSER=/usr/local/bin/triple-c-open
fi
# ── Playwright browser config ───────────────────────────────────────────────
# Seed ~/.playwright/cli.config.json on every start.
#
# Without it `playwright-cli` resolves to channel `chrome` — system Google
# Chrome — with the Chromium sandbox ON, and these containers do not permit
# unprivileged user namespaces, so the browser aborts with "Failed to move to
# new namespace ... Operation not permitted". On a base image that no longer
# ships Google Chrome the same default fails the other way, with "Chromium
# distribution 'chrome' is not found". One cause, two error messages, and
# neither of them looks like a configuration problem.
#
# Seeded here rather than baked into the image because ~/.playwright is inside
# the home volume: an image copy would reach new projects only, and every
# existing project would stay broken forever. Written on every start from a
# source outside the volume, the way CLAUDE_INSTRUCTIONS and the Mission
# Control skills already are.
#
# --seed-config-only is the cheap path: no npm install, no browser download, no
# apt, no verify launch, nothing over the network. It writes one small file if
# it is absent and returns. Measured at ~2 ms. The heavier repairs stay
# on-demand — run `triple-c-playwright-heal` with no arguments for those.
if [ -x /usr/local/bin/triple-c-playwright-heal ]; then
/usr/local/bin/triple-c-playwright-heal --seed-config-only --quiet || \
echo "entrypoint: warning — playwright config seeding failed (browser view may not launch)"
fi
# ── Scheduler setup ─────────────────────────────────────────────────────────
SCHEDULER_DIR="/home/claude/.claude/scheduler"
mkdir -p "$SCHEDULER_DIR/tasks" "$SCHEDULER_DIR/logs" "$SCHEDULER_DIR/notifications"
+225
View File
@@ -0,0 +1,225 @@
---
name: pia-vpn
description: Connect this container's traffic through a PIA VPN tunnel over WireGuard, or diagnose one that is not working. Use when asked to enable, route through, check, or tear down a VPN, when traffic needs to leave from a different location, or when DNS or connectivity broke after a VPN was brought up.
---
# PIA VPN
Bring this container's traffic out through Private Internet Access over
WireGuard, using the API PIA documents for headless use.
Run `sudo ~/.claude/skills/pia-vpn/pia-wg.sh` with `up`, `up --full`, `down` or
`status`. Read the rest of this page before the first `up --full` — three of the
behaviours below are actively misleading if you meet them without warning, and
each one presents as "the VPN is fine" or "Claude is broken" rather than as
what it is.
## Before anything else: what the toggle does not do
Triple-C's **VPN support** setting grants three things — `CAP_NET_ADMIN`, the
`/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark` sysctl — and
stops there. It starts no client, builds no tunnel and changes no route.
So "the VPN is enabled but traffic isn't going through it" is normally not a
fault. It means the capability is present and nothing has used it yet. Check
with `status` before assuming something is broken.
If the toggle is off, the script says so and names the setting. It cannot be
turned on from inside the container; the user changes it in Config → Runtime,
and it recreates the container on the next start (home and `.claude` volumes
are preserved — it is not a Reset).
## Two modes
| | routes | use when |
|---|---|---|
| `up` | only `1.1.1.1/32` | verifying the tunnel works without disturbing anything |
| `up --full` | all public traffic | you actually want traffic leaving via PIA |
Prefer `up` first. It proves the handshake, credentials and region are good
while your own connectivity is untouched, so a failure is cheap.
**`up --full` routes Claude Code's own API traffic through PIA.** If the tunnel
drops, that traffic stops until it recovers or you run `down`. Say so before
running it — the user may be mid-session, and they will experience the failure
as Claude going away, not as a VPN problem.
## Trap 1: a full tunnel takes DNS with it
The container resolves through an address on the Docker network — under Docker
Desktop, `192.168.65.7` — which sits **outside** the container's own subnet. A
default route of `0.0.0.0/0`, or the `0.0.0.0/1` + `128.0.0.0/1` pair, captures
it and posts every lookup into a tunnel that cannot carry private traffic.
Nothing resolves after that. The visible symptom is Claude Code reporting it
cannot connect, because `api.anthropic.com` no longer resolves:
```
$ curl https://api.anthropic.com/v1/messages
* Could not resolve host: api.anthropic.com (rc=6)
```
`pia-wg.sh` already handles this: it routes `10.0.0.0/8`, `172.16.0.0/12`,
`192.168.0.0/16` and `169.254.0.0/16` back via the original gateway, then pins
PIA's own resolvers through the tunnel with `/32` routes that outrank the
`10/8` exclusion. If you ever route traffic by hand, you owe both halves — the
exclusions *and* a resolver reachable from wherever you pointed the default.
The failure has a quiet twin. Do only the first half — exclude the private
ranges, leave the resolver alone — and everything *works*, while every DNS
query travels outside the tunnel to your ISP. A VPN that leaks the full list of
what you looked up is worse than one that is visibly broken, so `up --full`
refuses to proceed if PIA does not hand back resolvers rather than carrying on
without them.
The mechanism above is Docker Desktop's. On a user-defined Docker network the
resolver is `127.0.0.11`, which is loopback and never captured by a default
route — the trap still exists there (that resolver forwards upstream from
inside the container's namespace) but arrives by a different path. Check
`/etc/resolv.conf` rather than assuming which case you are in.
## Trap 2: an IP-literal health check cannot see a dead resolver
`curl https://1.1.1.1/cdn-cgi/trace` needs no DNS, so it returns a cheerful
PIA exit address while name resolution is entirely broken. A tunnel verified
that way looks perfect and works for nothing.
`status` resolves a real name for this reason. Trust its `DNS:` line, and if
you check by hand, resolve a name rather than fetching an address.
## Trap 3: in test mode, the obvious probe is the one thing tunnelled
`up` routes `1.1.1.1` and nothing else. So checking your address by fetching
`https://1.1.1.1/cdn-cgi/trace` reports a **PIA** address — not because your
traffic is going through PIA, but because that single probe is. Everything else
still leaves directly.
This reads exactly like a working full tunnel, and it is the likeliest reason
someone concludes the VPN is on when it is not. `status` prints both exits in
test mode for this reason:
```
mode: test route only (1.1.1.1 through the tunnel, nothing else)
through the tunnel: 64.113.5.73
everything else: 172.116.197.166 <- your real address
```
Two different addresses there is correct and expected in test mode. If you want
the second line to change, you want `up --full`.
## Trap 4: no tunnel survives a restart, and it fails open
The network namespace is rebuilt every time the container starts, and nothing
inside reconnects anything. After a stop/start, Reset or any config change that
recreates the container, the interface and its routes are gone.
State under `/run/pia-wg` rides the snapshot and persists, so leftover files
make it look as though the tunnel is still configured. It is not. Traffic goes
out the real address with no error and nothing visibly different.
Never infer from `/run/pia-wg` that a tunnel is up. Run `status` — if the
handshake line is missing, there is no tunnel. Re-run `up` after every start.
## Credentials
Two lines in `~/pia-creds` — username, then password:
```
p1234567
your-password
```
Treat the contents as secret: never print the file, never echo the values, and
never include them in a commit, a log or a message. The script reads it directly
and does not echo it, and passes PIA's session token to `curl` on stdin rather
than in the argv, where `ps` would expose it to everything in the container.
`PIA_CREDS` points somewhere else — but `sudo` resets the environment, so it
only takes effect **after** the word `sudo`:
```bash
sudo PIA_CREDS=/path/to/creds ~/.claude/skills/pia-vpn/pia-wg.sh up # works
PIA_CREDS=/path/to/creds sudo ~/.claude/skills/pia-vpn/pia-wg.sh up # ignored
```
The second form fails silently back to the default path. Same for `PIA_REGION`.
## Regions
Defaults to `us_chicago`. Override with `PIA_REGION`:
```bash
sudo PIA_REGION=uk_london ~/.claude/skills/pia-vpn/pia-wg.sh up --full
```
List the ids:
```bash
curl -s https://serverlist.piaservers.net/vpninfo/servers/v6 \
| head -1 | jq -r '.regions[].id'
```
## Verifying
`status` prints the handshake, DNS, and which address traffic actually leaves
from — labelled by mode, so the answer cannot be misread:
```
latest handshake: 2 seconds ago
transfer: 92 B received, 180 B sent
DNS: ok (via 10.0.0.243 10.0.0.242)
mode: full tunnel
all traffic exits: 64.113.5.244
```
All of it matters. A handshake with `DNS: BROKEN` is trap 1. `mode: test route
only` with two different addresses is trap 3, and is correct — it means the
tunnel works and you have not asked for it to carry anything yet. Report the
mode line when telling someone the VPN is on; "the public IP is a PIA one" is
true in test mode too, and means much less than it sounds like.
## Tearing down
`down` restores `resolv.conf` from its backup (only if that backup still looks
like a resolver file — restoring a truncated one would leave the container with
no DNS at all), removes exactly the routes that were added, in reverse order,
and deletes the interface. It is safe to run when nothing is up. Confirm
afterwards that the public address is back to the container's own.
`up` calls it too, but only *after* every network fetch has succeeded, so a
failed `up` leaves an existing tunnel alone rather than tearing it down to
report a bad password. From that point on a rollback is armed: if any step of
the setup fails, the tunnel is torn down rather than left half-configured.
The private key is deleted earlier still — the moment `wg set` has read it,
while the tunnel is being built. That is not housekeeping: `/run` is in the
container's writable layer, and recreating or migrating the project runs
`docker commit` over it *without* tearing the tunnel down first. A key that
lived for the tunnel's lifetime would be baked into the snapshot image and
copied forward from then on. The kernel keeps its own copy, so nothing is lost.
## What this deliberately does not do
- **No killswitch.** `iptables` *is* in the image, so one is buildable — this
is a deliberate omission, not a missing dependency. Blocking non-tunnel egress
cuts Claude Code's own API traffic the moment the tunnel drops, which ends the
session that would otherwise fix it. If the user needs guaranteed egress
rather than convenient egress, say so plainly and let them decide, rather than
improvising one.
- **No autostart.** There is no service manager in the container and Triple-C
has no start hook, so nothing re-establishes the tunnel on its own. `cron` is
in the image and `triple-c-scheduler` runs on it, so a scheduled reconnect is
possible if the user wants one — it is just not set up, and a tunnel that
reconnects unattended deserves an explicit decision.
- **Not PIA's desktop client.** `pia-daemon` and `piactl` are installable but
cannot work headless: the daemon never accepts a client connection without
the GUI, and `piactl --help` states that connecting requires it. If you find
one installed, it is not a working alternative to this script.
- **Not `wg-quick`.** Its `Table=auto` full-tunnel mode routes by firewall mark
and needs `xt_CONNMARK` from the host kernel, which Docker Desktop for
Windows (WSL2) does not have and a container cannot load. This script adds
the routes with `ip route` directly, which works on every host.
- **IPv4 only.** The `0.0.0.0/1` + `128.0.0.0/1` pair covers v4. A container
with a global IPv6 address and a v6 default route would leak all v6 traffic
outside the tunnel; Triple-C's containers do not have one by default, but
check `ip -6 route show default` before relying on this where it matters.
+318
View File
@@ -0,0 +1,318 @@
#!/usr/bin/env bash
# PIA over WireGuard, headless.
#
# PIA's desktop client (pia-daemon + piactl) cannot work here: its daemon never
# accepts a client connection without the GUI running, and `piactl --help` says
# as much. This talks to PIA's public API directly instead, which is the path
# PIA themselves document for headless use.
#
# sudo pia-wg.sh up tunnel up, only 1.1.1.1 routed through it (safe test)
# sudo pia-wg.sh up --full tunnel up, all *public* traffic exits via PIA
# sudo pia-wg.sh down tear down, restoring DNS and routes
# sudo pia-wg.sh status handshake, DNS and current public IP
#
# Requires the project's "VPN support" setting (Config -> Runtime) to be on.
#
# Settings are read from the environment, but note that sudo resets it: they
# have to be passed *through* sudo, after the word `sudo`, not before it.
#
# sudo PIA_REGION=uk_london pia-wg.sh up --full # works
# PIA_REGION=uk_london sudo pia-wg.sh up --full # silently ignored
#
# PIA_CREDS credentials file, two lines: username, then password
# (default /home/claude/pia-creds; never echoed by this script)
# PIA_REGION region id (default us_chicago). List them with:
# curl -s https://serverlist.piaservers.net/vpninfo/servers/v6 \
# | head -1 | jq -r '.regions[].id'
set -euo pipefail
# Not ~/pia-creds: under sudo, HOME is /root.
# Read by up()'s EXIT trap, which runs after the function's locals are gone.
SETUP_OK=0
CREDS=${PIA_CREDS:-/home/claude/pia-creds}
REGION=${PIA_REGION:-us_chicago}
IFACE=pia0
STATE=/run/pia-wg
# Kept off the tunnel in --full mode. The container's DNS resolver, the Docker
# host network (host.docker.internal, any host-side Ollama), sibling containers
# and the LAN all live in here. PIA cannot route any of it, so without these
# exclusions the container reaches the public internet and nothing else --
# including, fatally, its own resolver.
PRIVATE_NETS="10.0.0.0/8 172.16.0.0/12 192.168.0.0/16 169.254.0.0/16"
# Args are joined with spaces so a long message can be written as several
# source lines without the indentation ending up in the output.
die() { echo "pia-wg: $*" >&2; exit 1; }
# `x=$(cmd)` is a plain assignment, so `set -e` kills the script on a non-zero
# cmd *before* any `[ -z "$x" ] || die` line can run. Every capture below
# therefore goes through `run`; without it a wrong password exits 22 with no
# output at all, which is the most likely way this is used wrongly and was the
# least explained.
#
# It takes a description rather than reporting the command it ran: one of these
# invocations carries the account password in `-u`, and an error message is
# exactly the wrong place for that to surface.
run() { local what=$1; shift; "$@" || die "$what (exit $?)"; }
preflight() {
[ "$(id -u)" = 0 ] || die "run with sudo"
# CAP_NET_ADMIN is bit 12. Checking it by name gives a usable error; without
# it the first `ip` call fails with a bare "Operation not permitted" that
# points nowhere near the setting that actually needs changing.
#
# Deliberately NOT checking /dev/net/tun: kernel WireGuard is a netlink
# interface and does not use it (verified -- `ip link add type wireguard`
# succeeds with NET_ADMIN and no tun device). It is OpenVPN and userspace
# wireguard-go that need it. The real kernel dependency here is the
# `wireguard` module, which `ip link add` below reports on directly.
local caps
caps=$(awk '/^CapEff:/{print $2}' /proc/self/status)
if [ $(( 0x$caps & 0x1000 )) -eq 0 ]; then
die "this container has no CAP_NET_ADMIN." \
"Turn on \"VPN support\" in Config -> Runtime and start the project" \
"again. That recreates the container; the home and .claude volumes" \
"are preserved, so nothing in them is lost."
fi
command -v wg >/dev/null || \
die "wireguard-tools is not installed." \
"If this project's container was built from an older base image," \
"migrate it onto the current one -- that is what ships \`wg\`."
[ -r "$CREDS" ] || \
die "no credentials at $CREDS." \
"Two lines are expected: username, then password." \
"Set PIA_CREDS (after the word \`sudo\`) to read them elsewhere."
}
# Routes that must work. A silent failure here is the worst state this script
# can reach: the two half-routes need no gateway and would succeed, so the
# tunnel captures everything while the exclusions that keep DNS and the Docker
# host reachable are quietly missing -- and `status` still says "full tunnel".
add_route() {
ip route add "$@" || die "could not add route '$*'"
printf '%s\n' "$*" >> "$STATE/routes"
}
up() {
case "${1:-}" in
""|--full) ;;
*) die "unknown option '$1' (expected --full or nothing)." \
"Refusing rather than silently giving you a test route." ;;
esac
preflight
mkdir -p "$STATE"; cd "$STATE"
# `curl -o` creates the file before it knows the request failed, so a plain
# `[ -f ]` cache check can pin a truncated cert forever -- and /run rides the
# snapshot, so "forever" outlives the container. Fetch to a temp name and
# rename only on success.
if [ ! -s ca.rsa.4096.crt ]; then
run "could not download PIA's CA certificate" \
curl -sf -m 20 -o ca.crt.part \
https://raw.githubusercontent.com/pia-foss/manual-connections/master/ca.rsa.4096.crt
[ -s ca.crt.part ] || die "PIA's CA certificate downloaded empty"
mv ca.crt.part ca.rsa.4096.crt
fi
local u p tok srv sip scn priv pub resp ep gw dns
u=$(sed -n 1p "$CREDS"); p=$(sed -n 2p "$CREDS")
[ -n "$u" ] && [ -n "$p" ] || die "$CREDS needs two lines: username, then password"
tok=$(run "PIA rejected the credentials in $CREDS, or could not be reached" \
curl -sf -m 25 -u "$u:$p" \
https://www.privateinternetaccess.com/gtoken/generateToken | jq -r .token)
[ -n "$tok" ] && [ "$tok" != null ] || die "PIA returned no token - check the credentials in $CREDS"
run "could not fetch PIA's server list" \
curl -sf -m 30 https://serverlist.piaservers.net/vpninfo/servers/v6 \
| head -1 > servers.json
srv=$(jq -r --arg r "$REGION" '.regions[] | select(.id==$r) | .servers.wg[0]' servers.json)
sip=$(echo "$srv" | jq -r .ip); scn=$(echo "$srv" | jq -r .cn)
[ -n "$sip" ] && [ "$sip" != null ] || die "no WireGuard server for region '$REGION'"
# Only now tear down any previous tunnel. Doing it up front (as an earlier
# version did) meant a failed token fetch or an unreachable server list took
# a *working* tunnel down with it and silently reverted the container to its
# real address, while the error talked about credentials. Everything above
# this line can fail; nothing above it has touched the network stack.
#
# It also still does the job it was added for: clearing a stale resolv.conf
# backup so a second `up` cannot save PIA's own resolvers over the real ones.
down >/dev/null 2>&1 || true
# From here on the network stack is being modified, so any failure has to put
# it back rather than exit half-configured. `down` is idempotent and restores
# routes and resolv.conf exactly.
#
# EXIT rather than ERR, and a flag rather than the trap's own exit status: an
# ERR trap is not inherited by shell functions without `set -E`, so a failure
# inside add_route would not fire it, and `die` exits explicitly, which is not
# an error and would not fire it either. EXIT catches both.
SETUP_OK=0
trap '[ "$SETUP_OK" = 1 ] || { echo "pia-wg: setup failed - rolling back" >&2; down >/dev/null 2>&1; }' EXIT
# umask, not a later chmod: the file is created under the inherited 0022
# otherwise, so the key is world-readable for the moment in between.
( umask 077; priv=$(wg genkey); printf '%s' "$priv" > wg.priv )
priv=$(cat wg.priv); pub=$(printf '%s' "$priv" | wg pubkey)
# The token goes in on stdin as a curl config rather than in the argv, where
# `ps` and /proc/*/cmdline expose it to every process in the container --
# verified. It is a ~24h bearer credential for the whole PIA account.
# PIA pins its certificate to the server's common name, which is why this
# connects by CN and lets --connect-to point that name at the real address.
resp=$(printf -- '--data-urlencode "pt=%s"\n--data-urlencode "pubkey=%s"\n' "$tok" "$pub" \
| run "could not register the key with $scn" \
curl -sf -m 25 -G -K - --connect-to "$scn::$sip:" \
--cacert ca.rsa.4096.crt "https://$scn:1337/addKey")
[ "$(echo "$resp" | jq -r .status)" = OK ] || die "key registration failed: $resp"
: > "$STATE/routes"
ip link add "$IFACE" type wireguard 2>/dev/null || \
die "could not create a WireGuard interface." \
"The Docker host's kernel has no 'wireguard' module."
wg set "$IFACE" private-key wg.priv \
peer "$(echo "$resp" | jq -r .server_key)" \
endpoint "$(echo "$resp" | jq -r .server_ip):$(echo "$resp" | jq -r .server_port)" \
allowed-ips 0.0.0.0/0 persistent-keepalive 25
# The kernel holds the key from here, so the file has no reason to outlive
# this line -- and every reason not to: /run is in the writable layer, and a
# recreate or migrate runs `docker commit` over it without tearing the tunnel
# down first, baking the key into the project's snapshot image. `down` also
# removes it, for the case where `up` never got this far.
rm -f wg.priv
ip addr add "$(echo "$resp" | jq -r .peer_ip)/32" dev "$IFACE"
ip link set "$IFACE" up
if [ "${1:-}" = "--full" ]; then
ep=$(echo "$resp" | jq -r .server_ip)
gw=$(ip route show default | awk '{print $3; exit}')
# `default dev eth0` with no `via` yields the literal "eth0" here, which
# would make every exclusion below a malformed no-op.
[[ $gw =~ ^[0-9]+\.[0-9]+\.[0-9]+\.[0-9]+$ ]] || \
die "no usable default gateway to pin the tunnel against (got '${gw:-none}')"
# PIA's resolvers are required in --full. Without them the 10/8 exclusion
# below is already in place, so every lookup would go to the container's
# own resolver *outside* the tunnel -- a full tunnel leaking all its DNS,
# reported by `status` as perfectly healthy.
dns=$(echo "$resp" | jq -r '.dns_servers[]? // empty' | head -2)
[ -n "$dns" ] || die "PIA returned no DNS servers; refusing a full tunnel that would leak every lookup"
# Pin the endpoint to the pre-existing gateway first, so the tunnel's own
# packets do not try to route through the tunnel. Then beat the default
# route with two half-routes rather than replacing it -- nothing to restore
# on teardown, and the container keeps working if this script dies midway.
add_route "$ep/32" via "$gw"
add_route 0.0.0.0/1 dev "$IFACE"
add_route 128.0.0.0/1 dev "$IFACE"
# Keep container, host and LAN traffic off the tunnel. Longer prefixes than
# the two halves above, so these win.
for n in $PRIVATE_NETS; do add_route "$n" via "$gw"; done
# PIA's resolvers live inside 10/8, so pin them back through the tunnel with
# /32s -- longer still, so they beat the exclusion just added.
cp /etc/resolv.conf "$STATE/resolv.conf.bak"
for d in $dns; do add_route "$d/32" dev "$IFACE"; done
# resolv.conf is a bind mount: write through it, never replace it.
for d in $dns; do echo "nameserver $d"; done > /etc/resolv.conf
echo "full tunnel: public traffic exits via PIA; private ranges stay local"
else
add_route 1.1.1.1/32 dev "$IFACE"
echo "test route only: 1.1.1.1 goes via PIA, everything else unchanged"
fi
# A tunnel with no handshake still routes -- into a black hole. Without this
# `up --full` would exit 0 having pointed all traffic *and* resolv.conf at a
# peer that never answered, and `status` would print "mode: full tunnel".
local waited=0
until [ "$(wg show "$IFACE" latest-handshakes | awk '{print $2; exit}')" != 0 ]; do
waited=$((waited + 1))
[ "$waited" -lt 20 ] || die "no handshake from $REGION after 10s - rolled back"
sleep 0.5
done
SETUP_OK=1
trap - EXIT
status
}
down() {
[ "$(id -u)" = 0 ] || die "run with sudo"
# Only restore something that actually looks like a resolver file. Restoring
# an empty or truncated backup leaves the container with no DNS at all, which
# is worse than leaving the current one alone.
if [ -f "$STATE/resolv.conf.bak" ]; then
if grep -q '^nameserver' "$STATE/resolv.conf.bak" 2>/dev/null; then
cat "$STATE/resolv.conf.bak" > /etc/resolv.conf
else
echo "pia-wg: warning - saved resolv.conf looks empty; leaving the current one alone" >&2
fi
rm -f "$STATE/resolv.conf.bak"
fi
if [ -f "$STATE/routes" ]; then
# Reverse order: the specific overrides go before the ranges they sit in.
tac "$STATE/routes" | while read -r r; do
[ -n "$r" ] && ip route del $r 2>/dev/null || true
done
rm -f "$STATE/routes"
fi
ip link del "$IFACE" 2>/dev/null || true
# /run is in the writable layer and `docker commit` bakes it into the
# project's snapshot image, so a key left here rides that image into every
# future container. Verified: a snapshot already carried one.
rm -f "$STATE/wg.priv"
echo "tunnel down"
}
# Both are Cloudflare and both answer /cdn-cgi/trace over their bare address, so
# neither needs DNS. Only 1.1.1.1 is ever routed into the tunnel, which is what
# lets status tell the two exits apart.
TRACE_TUNNELLED=https://1.1.1.1/cdn-cgi/trace
TRACE_DIRECT=https://1.0.0.1/cdn-cgi/trace
exit_ip() { curl -s -m 20 "$1" | sed -n 's/^ip=//p'; }
status() {
# `wg show` needs root; `ip route`/`ip link` do not. Without this guard an
# unprivileged run prints "no tunnel up" and then "mode: full tunnel" in the
# same breath, and an agent reading the first line re-runs `up`.
[ "$(id -u)" = 0 ] || die "run with sudo"
wg show "$IFACE" 2>/dev/null | grep -E "latest handshake|transfer" || echo "no tunnel up"
# Resolve a name, not an IP literal. A curl to 1.1.1.1 succeeds while DNS is
# completely broken, which is exactly how a dead resolver goes unnoticed.
printf 'DNS: '
if timeout 10 getent hosts api.anthropic.com >/dev/null 2>&1; then
echo "ok (via $(sed -n 's/^nameserver //p' /etc/resolv.conf | tr '\n' ' '))"
else
echo "BROKEN - cannot resolve api.anthropic.com"
fi
# Report the exit per mode. In test mode the probe address is itself the one
# thing inside the tunnel, so a single "public IP" line would print a PIA
# address while every other packet leaves directly -- the exact reading that
# makes a test tunnel look like a full one.
if ip route show 0.0.0.0/1 2>/dev/null | grep -q "$IFACE"; then
echo "mode: full tunnel"
echo " all traffic exits: $(exit_ip "$TRACE_TUNNELLED")"
elif ip link show "$IFACE" >/dev/null 2>&1; then
echo "mode: test route only (1.1.1.1 through the tunnel, nothing else)"
echo " through the tunnel: $(exit_ip "$TRACE_TUNNELLED")"
echo " everything else: $(exit_ip "$TRACE_DIRECT") <- your real address"
else
echo "mode: no tunnel"
echo " all traffic exits: $(exit_ip "$TRACE_DIRECT")"
fi
}
case "${1:-}" in
up) shift; up "${1:-}" ;;
down) down ;;
status) status ;;
*) sed -n '2,26p' "$0" | sed 's/^# \{0,1\}//'; exit 1 ;;
esac
+272
View File
@@ -0,0 +1,272 @@
#!/bin/bash
# triple-c-playwright-heal — make Playwright usable in a Triple-C container.
#
# Idempotent: safe to run on every start and safe to re-run after a partial
# failure. Each step checks for its own result first, so a healthy container is
# a fast no-op that still prints why it is healthy. The last step is the only
# one that means anything: it launches a browser for real.
#
# The things that go wrong, in the order they bite:
#
# 1. @playwright/cli missing — including the case where it was installed and
# then silently removed again. `npm install --no-save <pkg>` in /workspace,
# which has no package.json, prunes packages npm considers extraneous, so
# installing @playwright/cli and then installing playwright wipes the
# first one and leaves an empty node_modules/@playwright/. That directory
# reads as "installed" to a naive check, which is why this script tests
# the package *entry point*.
#
# 2. Bundled chromium missing or the wrong revision. Browsers live in the
# home volume and outlive any single @playwright/cli install, so a stale
# chromium-<old> is routinely present while the installed playwright-core
# wants a newer one. Must be installed AS claude: run as root it lands in
# /root/.cache/ms-playwright where the agent cannot see it.
#
# 3. No cli.config.json — the one that breaks an otherwise clean install.
# With no config, playwright-cli resolves to channel `chrome` (system
# Google Chrome) with the sandbox ON. These containers forbid unprivileged
# user namespaces, so Chrome aborts with "Failed to move to new namespace
# ... Operation not permitted". On newer base images Chrome is not present
# at all and it fails with "Chromium distribution 'chrome' is not found".
# Same root cause both ways: the default channel is wrong here.
#
# 4. The storage-state file the config points at is missing. Playwright
# treats an unreadable storageState as a hard error on every launch, not
# as "no saved state", so the file has to exist from the very first run.
#
# 5. xvfb or socat missing (older base images only). Headless Playwright
# needs neither; the `playwright-cli show` dashboard needs xvfb, and the
# browser-view pane needs socat — without it the pane reports
# "127.0.0.1 sent an invalid response" while the container side is fine.
#
# Usage: triple-c-playwright-heal [--seed-config-only] [--force-config] [--quiet]
# --seed-config-only only ensure the config and its storage-state file
# exist. No npm install, no browser download, no apt, no
# verify launch. Cheap and offline — this is the mode
# entrypoint.sh runs on every container start.
# --force-config overwrite an existing config instead of keeping it
# --quiet print only problems and repairs, not healthy no-ops
set -u
TARGET_USER=claude
TARGET_HOME=/home/claude
PW_DIR=/workspace
CONFIG_DIR="$TARGET_HOME/.playwright"
CONFIG_FILE="$CONFIG_DIR/cli.config.json"
STATE_FILE="$CONFIG_DIR/storage-state.json"
CLI_ENTRY="$PW_DIR/node_modules/@playwright/cli/playwright-cli.js"
FORCE_CONFIG=0
QUIET=0
SEED_ONLY=0
for arg in "$@"; do
case "$arg" in
--force-config) FORCE_CONFIG=1 ;;
--quiet) QUIET=1 ;;
--seed-config-only) SEED_ONLY=1 ;;
*) echo "playwright-heal: unknown option: $arg" >&2; exit 2 ;;
esac
done
changed=0
failed=0
say() { [ "$QUIET" = 1 ] || echo "playwright-heal: $*"; }
warn() { echo "playwright-heal: $*" >&2; }
did() { changed=1; echo "playwright-heal: $*"; }
# Run as claude whether we were invoked as root (docker exec / entrypoint) or
# as claude (terminal session). Nothing user-visible may end up root-owned.
as_claude() {
if [ "$(id -u)" = 0 ]; then
su "$TARGET_USER" -s /bin/sh -c "$1"
else
sh -c "$1"
fi
}
# ── 1. @playwright/cli ───────────────────────────────────────────────────────
if [ "$SEED_ONLY" = 1 ]; then
:
elif [ -f "$CLI_ENTRY" ]; then
say "@playwright/cli present"
else
# An empty leftover @playwright/ can make npm consider the tree settled.
if [ -d "$PW_DIR/node_modules/@playwright" ]; then
say "clearing partial @playwright install"
rm -rf "$PW_DIR/node_modules/@playwright"
fi
say "installing @playwright/cli..."
if as_claude "cd $PW_DIR && npm install --no-save --no-audit --no-fund @playwright/cli" >/tmp/pw-heal-npm.log 2>&1; then
did "installed @playwright/cli"
else
warn "npm install failed; see /tmp/pw-heal-npm.log"
failed=1
fi
fi
# ── 2. bundled chromium ──────────────────────────────────────────────────────
# Ask Playwright where *this* version's chromium belongs rather than globbing
# chromium-*, which would call a stale revision "present" and then fail at
# launch with 'Browser "chrome-for-testing" is not installed'. --dry-run prints
# the install location for the installed version and downloads nothing.
if [ "$SEED_ONLY" = 1 ]; then
:
else
chromium_dir=""
if [ -f "$PW_DIR/node_modules/playwright-core/cli.js" ]; then
chromium_dir=$(as_claude "cd $PW_DIR && node node_modules/playwright-core/cli.js install --dry-run chromium 2>/dev/null" \
| awk '/Install location:/ { print $3; exit }')
fi
if [ -n "$chromium_dir" ] && [ -d "$chromium_dir" ]; then
say "chromium present ($(basename "$chromium_dir"))"
elif [ -f "$PW_DIR/node_modules/playwright-core/cli.js" ]; then
say "downloading chromium (~300 MB)..."
if as_claude "cd $PW_DIR && node node_modules/playwright-core/cli.js install chromium" >/tmp/pw-heal-browser.log 2>&1; then
did "installed chromium"
else
warn "chromium install failed; see /tmp/pw-heal-browser.log"
failed=1
fi
else
warn "playwright-core missing, cannot install chromium"
failed=1
fi
fi
# ── 3. cli.config.json ───────────────────────────────────────────────────────
# The *global* config, not a project-level .playwright/, because the project
# one resolves relative to the current working directory and silently stops
# applying the moment you cd elsewhere.
#
# `chrome-for-testing` is the only recognised chromium alias — "chromium" is
# not one and falls back to system Chrome. chromiumSandbox:false is what
# actually appends --no-sandbox.
write_config() {
mkdir -p "$CONFIG_DIR" || return 1
cat > "$CONFIG_FILE" <<EOF
{
"browser": {
"browserName": "chromium",
"launchOptions": {
"channel": "chrome-for-testing",
"chromiumSandbox": false,
"args": ["--no-sandbox", "--disable-dev-shm-usage"]
},
"contextOptions": {
"storageState": "$STATE_FILE"
}
}
}
EOF
chown -R "$TARGET_USER:$TARGET_USER" "$CONFIG_DIR" 2>/dev/null || true
}
# storageState is a *load* path, and Playwright reads it at context creation.
# A path that does not exist is not treated as "no saved state" — it is a hard
# error, "Error reading storage state from …", on every single launch. So the
# file has to exist before the config that names it can be used at all, and it
# has to be recreated if anything deletes it. An empty state is valid and
# behaves exactly like no state.
#
# The path is read back out of the config rather than assumed, so a
# hand-edited config pointing somewhere else still gets its file created
# instead of being silently broken by ours.
ensure_state_file() {
[ -f "$CONFIG_FILE" ] || return 0
state_path=$(grep -o '"storageState"[[:space:]]*:[[:space:]]*"[^"]*"' "$CONFIG_FILE" 2>/dev/null \
| sed 's/.*"\([^"]*\)"[[:space:]]*$/\1/')
[ -n "$state_path" ] || return 0
[ -f "$state_path" ] && return 0
mkdir -p "$(dirname "$state_path")" 2>/dev/null
printf '{\n "cookies": [],\n "origins": []\n}\n' > "$state_path" || return 1
chown "$TARGET_USER:$TARGET_USER" "$state_path" 2>/dev/null || true
did "created empty $state_path (storageState needs it to exist)"
}
if [ ! -f "$CONFIG_FILE" ]; then
if write_config; then did "wrote $CONFIG_FILE"; else warn "could not write $CONFIG_FILE"; failed=1; fi
elif [ "$FORCE_CONFIG" = 1 ]; then
if write_config; then did "overwrote $CONFIG_FILE (--force-config)"; else warn "could not write $CONFIG_FILE"; failed=1; fi
elif grep -q '"chromiumSandbox"[[:space:]]*:[[:space:]]*false' "$CONFIG_FILE" 2>/dev/null; then
say "config present and disables the sandbox"
else
# Present but hand-edited into a state that will not launch. Do not clobber
# deliberate config silently; say what is wrong and how to replace it.
warn "config at $CONFIG_FILE does not set chromiumSandbox:false — the browser will likely fail to launch. Re-run with --force-config to replace it."
fi
# Unconditional: the config may name a storageState this run did not write —
# one seeded by an older version of this script, or edited by hand — and a
# missing file there breaks every launch.
ensure_state_file || { warn "could not create the storage-state file"; failed=1; }
if [ "$SEED_ONLY" = 1 ]; then
[ "$failed" = 1 ] && exit 1
exit 0
fi
# ── 4. xvfb (headed dashboard only) ──────────────────────────────────────────
# Current base images get this from `playwright install-deps` (its `tools`
# group); older ones predate that layer. Headless never needs it, so a missing
# xvfb is a note, not a failure.
if command -v Xvfb >/dev/null 2>&1; then
say "xvfb present"
elif [ "$(id -u)" = 0 ]; then
say "installing xvfb (needed only for the headed dashboard)..."
if (apt-get update -qq && DEBIAN_FRONTEND=noninteractive apt-get install -y -qq xvfb) >/tmp/pw-heal-xvfb.log 2>&1; then
did "installed xvfb"
else
warn "xvfb install failed (headless still works); see /tmp/pw-heal-xvfb.log"
fi
else
say "xvfb missing and not running as root — skipping (headless still works)"
fi
# ── 4b. socat (the browser-view pane's tunnel) ───────────────────────────────
# Not Playwright's, but the same class of failure and it presents as a
# Playwright problem: the pane's host-side proxy reaches the dashboard by
# running `socat` *inside* the container over a Docker exec. On a container old
# enough to predate socat in the base image, that exec produces something that
# is not an HTTP response, and the webview reports "127.0.0.1 sent an invalid
# response" — with the container side working perfectly. A project keeps the
# base image it was first built from until it is migrated, so this is the
# normal case on an older project, not an exotic one.
if command -v socat >/dev/null 2>&1; then
say "socat present"
elif [ "$(id -u)" = 0 ]; then
say "installing socat (needed by the browser-view pane)..."
if (apt-get update -qq && DEBIAN_FRONTEND=noninteractive apt-get install -y -qq socat) >/tmp/pw-heal-socat.log 2>&1; then
did "installed socat"
else
warn "socat install failed; the browser-view pane will report an invalid response. See /tmp/pw-heal-socat.log"
failed=1
fi
else
warn "socat missing and not running as root — the browser-view pane will report an invalid response"
fi
# ── 5. verify by actually launching ──────────────────────────────────────────
# Every step above can report success while the browser still refuses to
# start — that is precisely how this broke. A dedicated session name keeps
# this clear of whatever the agent already has open.
if [ -f "$CLI_ENTRY" ]; then
verify_out=$(as_claude "cd /tmp && timeout 90 node $CLI_ENTRY -s=heal-verify open 'data:text/html,<h1>ok</h1>' 2>&1")
if printf '%s' "$verify_out" | grep -q 'opened with pid'; then
say "verified: browser launches"
as_claude "cd /tmp && timeout 30 node $CLI_ENTRY -s=heal-verify close" >/dev/null 2>&1
else
warn "browser still fails to launch:"
printf '%s\n' "$verify_out" | grep -m4 -E 'namespace|Check failed|is not installed|is not found|missing dependencies|Error' >&2
failed=1
fi
else
warn "@playwright/cli not installed — nothing to verify"
failed=1
fi
[ "$failed" = 1 ] && exit 1
[ "$changed" = 1 ] && say "done — repairs applied" || say "done — nothing to repair"
exit 0