Compare commits

..
Author SHA1 Message Date
shadow-testandClaude Opus 5 d260f2c7c3 Let a project's container run a VPN client
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m37s
Build App (Preview) / build-linux (pull_request) Successful in 7m45s
Build App (Preview) / build-windows (pull_request) Failing after 13m30s
Build App (Preview) / prune-previews (pull_request) Skipped
A VPN client installed in a container today starts, runs, and then hangs
until its connection times out. Nothing reports an error: a default
container has no /dev/net/tun to open and no CAP_NET_ADMIN to add an
interface or a route with, and clients surface that as a generic timeout
rather than a permissions failure.

Add an opt-in per-project "VPN support" switch granting the three things
a tunnel needs. They are useless individually, which is why
vpn_host_config() defines the set in one place and the tests assert all
of it:

  * CAP_NET_ADMIN — Docker's default bounding set has net_raw but not
    net_admin, so a client can ping but never connect.
  * /dev/net/tun — passed through from the host so the kernel's tun
    module backs it, rather than mknod-ed inside.
  * net.ipv4.conf.all.src_valid_mark — WireGuard's wg-quick sets this and
    cannot from inside a container, /proc/sys being read-only, so its
    handshakes are dropped by reverse-path filtering.

Off by default and deliberately opt-in: NET_ADMIN lets anything in the
container reconfigure that container's network stack. It is namespaced —
no authority over the host's interfaces or any other container.

Capabilities and devices are fixed when a container is created, so this
is container state and takes the label-and-compare treatment.
triple-c.vpn-support is written unconditionally, false included, for the
usual docker commit reason: a true stamped once would ride the snapshot
image into every future container and make the switch impossible to turn
back off. A missing label reads as false and off is byte-identical to
today, so no existing project is churned.

Requesting the device fails at creation when the host kernel has no tun
module, which would otherwise surface as a project that simply refuses to
start. explain_create_failure() rewrites that one error to name the
switch and the Docker-Desktop-VM-versus-your-machine distinction, and
leaves every other failure untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 07:52:39 -07:00
9 changed files with 68 additions and 456 deletions
+14 -98
View File
@@ -39,48 +39,13 @@ jobs:
MAJOR_MINOR=$(cat VERSION | tr -d '[:space:]') MAJOR_MINOR=$(cat VERSION | tr -d '[:space:]')
echo "Major.Minor: ${MAJOR_MINOR}" echo "Major.Minor: ${MAJOR_MINOR}"
# The patch number is **one past the highest patch already used**, and # Find the latest tag matching v{MAJOR_MINOR}.N (exclude -mac, -win suffixes)
# never a distance. # `|| true` so an empty grep result doesn't fail the step under pipefail.
# LATEST_TAG=$(git tag -l "v${MAJOR_MINOR}.*" --sort=-v:refname | grep -E "^v${MAJOR_MINOR}\.[0-9]+$" | head -1 || true)
# It used to be `git rev-list --count <highest tag>..HEAD`, which is
# not a counter at all: it measures how far HEAD has drifted from
# whichever tag sorts highest, and that resets to zero every time a
# tag is cut. The published history is the proof — each of these is
# exactly what the old formula returned at the time:
#
# v0.4.0 -> 3 commits -> v0.4.3 looked fine
# v0.4.3 -> 4 commits -> v0.4.4 fine by luck, 4 > 3
# v0.4.4 -> 2 commits -> v0.4.2 went backwards
# v0.4.4 -> 6 commits -> v0.4.6 jumped, skipping .5
# v0.4.6 -> 3 commits -> v0.4.3 already taken; the upload failed
#
# Reusing a version is worse than failing to publish one: the macOS
# and Windows steps replace assets in place, so a duplicate silently
# rewrote a release that had been public for three days. Monotonic
# numbering is what stops that at the source.
#
# Suffixed tags count too. `create-tag` is skipped when any platform
# job fails, so a run can publish v0.4.7-mac and never create the
# plain v0.4.7 — reading only unsuffixed tags would then hand the
# same number out twice.
HIGHEST=$(git tag -l "v${MAJOR_MINOR}.*" \
| grep -E "^v${MAJOR_MINOR}\.[0-9]+(-mac|-win)?$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+).*/\1/" \
| sort -n | tail -1 || true)
# A re-run of a commit that already released must not mint a new if [ -n "$LATEST_TAG" ]; then
# version just because its own tag now exists. echo "Latest matching tag: ${LATEST_TAG}"
EXISTING=$(git tag --points-at HEAD \ PATCH=$(git rev-list --count "${LATEST_TAG}..HEAD")
| grep -E "^v${MAJOR_MINOR}\.[0-9]+$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+)$/\1/" \
| sort -n | tail -1 || true)
if [ -n "$EXISTING" ]; then
echo "HEAD is already tagged v${MAJOR_MINOR}.${EXISTING} — reusing it"
PATCH="${EXISTING}"
elif [ -n "$HIGHEST" ]; then
echo "Highest patch already used on this line: ${HIGHEST}"
PATCH=$((HIGHEST + 1))
else else
# A minor line nobody has tagged yet is a *new* line, and a new line # A minor line nobody has tagged yet is a *new* line, and a new line
# starts at .0 — that is what "we are moving to 0.4.x" means. The # starts at .0 — that is what "we are moving to 0.4.x" means. The
@@ -200,70 +165,21 @@ jobs:
env: env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }} TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: | run: |
set -euo pipefail
TAG="v${{ needs.compute-version.outputs.version }}" TAG="v${{ needs.compute-version.outputs.version }}"
# Create release
# Idempotent get-or-create, matching build-macos. This step used to curl -s -X POST \
# POST /releases unconditionally: against a tag that already existed
# Gitea answered 409, the grep below found no id, and the run died
# with a bare "exitcode '1'" and not one line of output explaining
# it — `curl -s` with no `-f` swallows the HTTP error, so nothing
# ever said "409" or "duplicate tag". Hence -fsS throughout, and
# pipefail so a failure cannot be stepped over.
HTTP_CODE=$(curl -sS -o release.json -w '%{http_code}' \
-H "Authorization: token ${TOKEN}" \ -H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/tags/${TAG}") -H "Content-Type: application/json" \
case "${HTTP_CODE}" in -d "{\"tag_name\": \"${TAG}\", \"name\": \"Triple-C ${TAG} (Linux)\", \"body\": \"Automated build from commit ${{ gitea.sha }}\"}" \
200) "${GITEA_URL}/api/v1/repos/${REPO}/releases" > release.json
echo "Release ${TAG} already exists, reusing" RELEASE_ID=$(cat release.json | grep -o '"id":[0-9]*' | head -1 | grep -o '[0-9]*')
;;
404)
echo "Creating release ${TAG}"
curl -fsS -X POST \
-H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/json" \
-d "{\"tag_name\": \"${TAG}\", \"name\": \"Triple-C ${TAG} (Linux)\", \"body\": \"Automated build from commit ${{ gitea.sha }}\"}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases" > release.json
;;
*)
echo "Unexpected ${HTTP_CODE} looking up release ${TAG}:" >&2
cat release.json >&2
exit 1
;;
esac
RELEASE_ID=$(python3 -c "import json,sys; print(json.load(open('release.json')).get('id',''))")
if [ -z "${RELEASE_ID}" ]; then
echo "No release id for ${TAG}; refusing to upload into nothing:" >&2
cat release.json >&2
exit 1
fi
echo "Release ID: ${RELEASE_ID}" echo "Release ID: ${RELEASE_ID}"
# Upload each artifact
# Replace-not-conflict, so a retry after a partial upload succeeds.
# Versions are monotonic now (see compute-version), so this can only
# ever be replacing an asset from a failed run of this same commit —
# never one belonging to an already-published version.
for file in artifacts/*; do for file in artifacts/*; do
[ -f "$file" ] || continue [ -f "$file" ] || continue
filename=$(basename "$file") filename=$(basename "$file")
EXISTING_ID=$(curl -sS \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets" \
| python3 -c "import json,sys; t=sys.argv[1]; print(next((a['id'] for a in json.load(sys.stdin) if a.get('name')==t), ''))" "${filename}" || true)
if [ -n "${EXISTING_ID}" ]; then
echo "Deleting existing asset ${filename} (id ${EXISTING_ID})"
curl -fsS -X DELETE \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets/${EXISTING_ID}"
fi
echo "Uploading ${filename}..." echo "Uploading ${filename}..."
curl -fsS --http1.1 \ curl -s -X POST \
--retry 5 --retry-all-errors --retry-delay 5 \
--max-time 600 \
-X POST \
-H "Authorization: token ${TOKEN}" \ -H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/octet-stream" \ -H "Content-Type: application/octet-stream" \
--data-binary "@${file}" \ --data-binary "@${file}" \
+5 -49
View File
@@ -203,8 +203,7 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li
### Container (`container/`) ### Container (`container/`)
- **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared - **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared
libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled` libraries a browser links against (see below)
toggle grants capability for (`iproute2`, `wireguard-tools`, `iptables`)
- **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs - **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs
`npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own `npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own
dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium
@@ -288,58 +287,15 @@ container is created once by a very long function where a dropped capability is
Any two without the third still presents as a connection that hangs to a timeout, which is why Any two without the third still presents as a connection that hangs to a timeout, which is why
the tests assert the whole set. the tests assert the whole set.
- **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun` - **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun`
module has to back it. module has to back it. When the host has no such device the failure lands at *creation* — the
- **A missing device fails at `start`, not `create` — verified against Docker 29.7.** `docker project simply won't start — so `explain_create_failure()` rewrites that one error to name the
create --device /dev/does-not-exist` succeeds and prints an id; runc resolves the device (and switch and the Docker-Desktop-VM-vs-your-machine distinction. Do not let it degrade to a raw
validates sysctls) only when it builds the container. So the guard belongs on the start path: bollard string.
`explain_container_failure()` covers both and is called from `start_container`, where it has a
container id and no project — which is why it keys off the error naming `/dev/net/tun` rather
than off `vpn_support_enabled`. Nothing else in Triple-C requests a device, so that is
unambiguous. A version of this check wired to `create` alone is dead code that looks correct.
- **`NET_ADMIN` here is not user-namespaced.** Docker does not enable userns remapping by default,
so only the *network* namespace confines it: no reach onto host interfaces, but promiscuous
mode, arbitrary addresses/routes/NAT on the shared `docker0` segment (sibling containers, the
LiteLLM gateway among them, are ARP-spoofable), netlink-triggered host module auto-load, and
enough authority to flush in-container netfilter rules that sandbox mode may rely on. Keep the
code comments honest about this — an earlier draft claimed it "confers no authority" outside the
container, which is too strong.
- **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual - **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual
`docker commit` reason: a `true` stamped once would ride the snapshot image into every future `docker commit` reason: a `true` stamped once would ride the snapshot image into every future
container and make the switch impossible to turn off. container and make the switch impossible to turn off.
- Off is byte-identical to a container created before the feature existed, and a missing label - Off is byte-identical to a container created before the feature existed, and a missing label
reads as `false`, so no existing project is churned. reads as `false`, so no existing project is churned.
- **The toggle grants capability and stops there — it routes nothing.** `vpn_host_config()` returns
a cap, a device and a sysctl; no client is installed, no route is touched, no tunnel is started
or restored. Users read the name as "turn the VPN on" and report the default network not routing
through it as a bug. It isn't, and the docs say so explicitly; keep it that way.
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
base-image migration — leaving a project holding the capability with nothing able to exercise it,
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
writable layer, so on a stop/start it is simply the same container's files, and on a recreation
`docker commit` has carried it into the snapshot. Traffic silently reverts to the real address.
Any future autostart or killswitch work starts here.
- **`/run` riding the snapshot means a VPN client's key material can end up in an image.** Verified:
a fresh container off the whp snapshot already contained the `wg.priv` a previous tunnel left in
`/run`. Anything writing key material there inherits the problem — the same `docker commit`
hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks
ephemeral and is not. A VPN client that does this should delete its key on teardown.
- **`iptables` is baked, and picking `nftables` instead would have been wrong.** `Recommends:
nftables | iptables` is stripped by `--no-install-recommends`, and `wg-quick` needs a backend for
any `AllowedIPs = 0.0.0.0/0`. `nftables` is the tempting choice — preferred by `wg-quick`, half
the size — but `wg-quick` picks nft *unconditionally* when present, and its nft ruleset needs
`nft_fib_ipv4`, which LinuxKit (Docker Desktop for Mac) does not build while it *does* build
`xt_CONNMARK`. Shipping nftables would therefore have forfeited Mac. See the Dockerfile comment;
the kernel-config evidence is quoted there.
- **Two `wg-quick` failures remain, and only one is ours to fix.** Full tunnels still need
`xt_CONNMARK`, which WSL2 before 6.6 lacks — nothing installable changes that. And every
provider's stock config carries a `DNS =` line that fails in `set_dns()` before any routing, so it
breaks split tunnels too; `openresolv` has no candidate on noble and `resolvconf` drags in
systemd-resolved, so that one is documented rather than fixed. Driving `wg` and `ip route`
directly avoids both, which is what the skill does.
### Container Lifecycle ### Container Lifecycle
+9 -59
View File
@@ -477,70 +477,21 @@ When enabled, the container is given the three things a VPN client needs to buil
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark` the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
sysctl that WireGuard requires. This is **off by default**. sysctl that WireGuard requires. This is **off by default**.
The `ip`, `wg` and `iptables` commands ship in the container image so there is something able to use Without it, a client such as PIA, WireGuard, OpenVPN or Tailscale installs and its daemon starts
them. If your project's container was created from an older base image it will not have them, and normally, but the connection attempt **hangs until it times out** — a default container has no tun
`wg` will simply not be found — **migrating the project onto the current base image** is what picks device to open and no permission to add an interface or a route, and most clients report that as a
them up. `sudo apt install iproute2 wireguard-tools iptables` works in the meantime, but lives in generic timeout rather than a permissions error.
the writable layer, so it is undone by a **Reset** and by a migration.
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
container's traffic to start leaving through a VPN is the most common misreading of what it does —
configuring a tunnel and routing traffic into it remains yours to do.
With the setting **off**, a client such as PIA or OpenVPN installs and its daemon starts normally,
but the connection attempt **hangs until it times out** — a default container has no tun device to open
and no permission to add an interface or a route, and most clients report that as a generic timeout
rather than a permissions error.
Things worth knowing: Things worth knowing:
- Tailscale is the exception: in its `--tun=userspace-networking` mode it needs neither the - `NET_ADMIN` applies to the container's **own** network namespace. It confers no authority over
capability nor the device, so leave this off if that is all you want. the host's interfaces or over any other container. It does mean anything running in the
container can reconfigure that namespace, which is why it is opt-in.
- `NET_ADMIN` applies to the container's **own** network namespace — it cannot touch the host's
interfaces. It is not nothing, though: within that namespace anything in the container can set
promiscuous mode and add arbitrary addresses, routes and firewall rules on the Docker bridge it
shares with your other containers, and it can flush firewall rules that sandbox mode relies on.
Grant it per project, to projects that need it.
- The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is - The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is
the Linux VM, not your own machine. If it is missing, the container is created but fails to the Linux VM, not your own machine. If it is missing, the container fails to create with an
**start**, with an error naming `/dev/net/tun` and pointing back at this setting. error naming `/dev/net/tun` and pointing back at this setting.
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the - A VPN client's kill switch applies to everything in the container, Claude Code included. If the
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off. tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
- **No tunnel survives a restart.** The network namespace is built fresh every time the container
starts, and there is no service manager inside to reconnect anything. Leftover state under `/run`
makes it *look* like the tunnel is still configured — that directory is in the container's
writable layer, so it is simply still there after a stop/start, and `docker commit` carries it
into the snapshot that a recreation is built from. Either way the interface and its routes are
gone and traffic goes out your real address again, with no error and nothing visibly different.
Re-establish it after every start, and check rather than assume.
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your
resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own
subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures
it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is
`192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`,
which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is
total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`.
Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original
gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own,
or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own
address via the original gateway, or the tunnel's encrypted packets try to route through the
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
while DNS is broken — resolve a name instead.
- **Delete a client's key material when you tear a tunnel down.** Anything written under `/run` is
in the container's writable layer, and recreating or migrating the project runs `docker commit`
over it — so a WireGuard private key left there gets baked into the project's snapshot image and
copied forward from then on. This is not hypothetical; it has already happened here.
- **Strip the `DNS =` line from a provider's `.conf` before `wg-quick up`.** Every commercial
provider ships one, and `wg-quick` hands it to `resolvconf`, which is not installed — so it fails
at `resolvconf: command not found` and deletes the interface again. This happens before any
routing, so it takes **split tunnels down too**. Set the resolver another way instead, or drive
`wg` and `ip route` directly rather than going through `wg-quick`.
- **`wg-quick` full tunnels also need `xt_CONNMARK` from the host kernel.** Native Linux, Docker
Desktop for Mac and WSL2 kernels from 6.6 have it; older WSL2 kernels do not, and a container
cannot load one. There the answer is again to add the routes yourself with `ip route`, which
needs no firewall backend on any platform.
> This setting can only be changed when the container is stopped. Capabilities and devices are > This setting can only be changed when the container is stopped. Capabilities and devices are
> fixed when a container is created, so toggling it recreates the container on the next start. > fixed when a container is created, so toggling it recreates the container on the next start.
@@ -1300,7 +1251,6 @@ The sandbox container (Ubuntu 24.04) comes pre-installed with:
| ruff | Latest | Python linter/formatter | | ruff | Latest | Python linter/formatter |
| Rust | Stable | Rust development (via rustup) | | Rust | Stable | Rust development (via rustup) |
| Docker CLI | Latest | Container management (when spawning is enabled) | | Docker CLI | Latest | Container management (when spawning is enabled) |
| iproute2, WireGuard tools, iptables | Latest | Building a tunnel (when VPN Support is enabled) |
| git | Latest | Version control | | git | Latest | Version control |
| GitHub CLI (gh) | Latest | GitHub integration | | GitHub CLI (gh) | Latest | GitHub integration |
| AWS CLI | v2 | AWS services and Bedrock | | AWS CLI | v2 | AWS services and Bedrock |
+34 -75
View File
@@ -830,16 +830,8 @@ type VpnHostConfigParts = (
/// its handshake packets are dropped by reverse-path filtering. Harmless for /// its handshake packets are dropped by reverse-path filtering. Harmless for
/// OpenVPN-based clients, so it is set unconditionally with the rest. /// OpenVPN-based clients, so it is set unconditionally with the rest.
/// ///
/// What it costs, stated accurately: Docker does not enable user-namespace /// This is namespaced to the container's own network stack: `NET_ADMIN` confers
/// remapping by default, so this is a real `CAP_NET_ADMIN` in the *initial* /// no authority over the host's interfaces or over any other container.
/// user namespace and only the **network** namespace confines it. It cannot
/// touch the host's interfaces, but within its own namespace it can set
/// promiscuous mode and add arbitrary addresses, routes and NAT rules on the
/// shared `docker0` L2 segment — which puts sibling containers (the LiteLLM
/// gateway among them) within reach of ARP spoofing, and lets netlink trigger
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
/// Hence opt-in, per project, rather than on for everyone.
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts { fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
if !enabled { if !enabled {
return (None, None, None); return (None, None, None);
@@ -865,42 +857,31 @@ fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
/// Turn the daemon's device-passthrough failure into an explanation. /// Turn the daemon's device-passthrough failure into an explanation.
/// ///
/// **This fires on `start`, not `create`.** Verified against Docker 29.7: /// Requesting `/dev/net/tun` fails at *creation* when the host kernel has no
/// `docker create --device /dev/does-not-exist` succeeds and prints an id; the /// `tun` module — and the raw bollard error names a path the user will look for
/// device is only resolved when runc builds the container, so the failure lands /// on the wrong machine, since with Docker Desktop the relevant host is the
/// on the *next* call. Sysctls validate at the same point. Anything that /// Linux VM rather than their own. Left unmapped this surfaces as a project
/// inspects only the create path will never see it — which is why both paths /// that simply refuses to start, with nothing pointing back at the switch that
/// route through here and the tests exercise the start-side string. /// caused it.
/// fn explain_create_failure(err: &str, vpn_enabled: bool) -> String {
/// Unmapped, this reads as `Failed to start container: Docker responded with let device_missing = vpn_enabled
/// status code 500: error gathering device information while adding custom && err.contains(TUN_DEVICE)
/// device "/dev/net/tun": no such file or directory` — a path the user will go
/// looking for on the wrong machine, since with Docker Desktop the relevant
/// host is the Linux VM rather than their own, and with nothing pointing back
/// at the switch that caused it.
///
/// Deliberately not gated on `vpn_support_enabled`: nothing else in Triple-C
/// ever asks for a device, so an error naming `/dev/net/tun` can only have come
/// from a container created with the switch on. That keeps the check usable
/// from [`start_container`], which has a container id and no project.
fn explain_container_failure(action: &str, err: &str) -> String {
let device_missing = err.contains(TUN_DEVICE)
&& (err.contains("no such file or directory") && (err.contains("no such file or directory")
|| err.contains("No such file or directory") || err.contains("No such file or directory")
|| err.contains("error gathering device information")); || err.contains("error gathering device information"));
if device_missing { if device_missing {
return format!( return format!(
"Failed to {} container: the Docker host has no {} device, which \ "Failed to create container: the Docker host has no {} device, which \
\"VPN support\" requires. The host kernel needs the `tun` module \ \"VPN support\" requires. The host kernel needs the `tun` module \
loaded (on Docker Desktop that is the Linux VM, not your own \ loaded (on Docker Desktop that is the Linux VM, not your own \
machine). Turn VPN support off in Config → Runtime to start this \ machine). Turn VPN support off in Config → Runtime to start this \
project without it. Original error: {}", project without it. Original error: {}",
action, TUN_DEVICE, err TUN_DEVICE, err
); );
} }
format!("Failed to {} container: {}", action, err) format!("Failed to create container: {}", err)
} }
pub async fn create_container( pub async fn create_container(
@@ -1593,7 +1574,7 @@ pub async fn create_container(
let response = docker let response = docker
.create_container(Some(options), config) .create_container(Some(options), config)
.await .await
.map_err(|e| explain_container_failure("create", &e.to_string()))?; .map_err(|e| explain_create_failure(&e.to_string(), project.vpn_support_enabled))?;
Ok(response.id) Ok(response.id)
} }
@@ -1603,7 +1584,7 @@ pub async fn start_container(container_id: &str) -> Result<(), String> {
docker docker
.start_container(container_id, None::<StartContainerOptions<String>>) .start_container(container_id, None::<StartContainerOptions<String>>)
.await .await
.map_err(|e| explain_container_failure("start", &e.to_string())) .map_err(|e| format!("Failed to start container: {}", e))
} }
pub async fn stop_container(container_id: &str) -> Result<(), String> { pub async fn stop_container(container_id: &str) -> Result<(), String> {
@@ -2789,53 +2770,31 @@ mod tests {
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]); assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
} }
/// What bollard actually hands us when a tun-less host rejects the device.
///
/// Captured verbatim from Docker 29.7: `docker create` with a missing
/// device **succeeds**, and this arrives from the subsequent `start`.
/// `DockerResponseServerError`'s Display is
/// `"Docker responded with status code {code}: {message}"` with the
/// daemon's message unaltered.
const REAL_TUN_ERROR: &str = "Docker responded with status code 500: error \
gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
#[test] #[test]
fn a_missing_tun_device_is_explained_on_the_path_that_actually_fails() { fn a_missing_tun_device_is_explained_rather_than_echoed() {
// The start path is the one that matters: the daemon defers device let raw = "error gathering device information while adding custom device \
// resolution to runc, so create returns an id on a host with no tun \"/dev/net/tun\": no such file or directory";
// module and only start fails. A version of this that checked create let msg = explain_create_failure(raw, true);
// alone would be dead code.
let msg = explain_container_failure("start", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to start container:"), "{}", msg);
assert!(msg.contains("VPN support"), "should name the switch: {}", msg); assert!(msg.contains("VPN support"), "should name the switch: {}", msg);
assert!(msg.contains("tun` module"), "should name the cause: {}", msg); assert!(msg.contains("tun` module"), "should name the cause: {}", msg);
assert!(msg.contains("Config → Runtime"), "should say where to fix it: {}", msg); assert!(msg.contains(raw), "should keep the original error: {}", msg);
assert!(msg.contains(REAL_TUN_ERROR), "should keep the original: {}", msg);
}
#[test]
fn the_same_explanation_covers_create_if_the_daemon_ever_checks_earlier() {
// Belt and braces — older and future daemons may validate at create.
let msg = explain_container_failure("create", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to create container:"), "{}", msg);
assert!(msg.contains("VPN support"), "{}", msg);
} }
#[test] #[test]
fn unrelated_failures_are_left_alone() { fn unrelated_failures_are_left_alone() {
for (action, err) in [ // Including a tun error on a project that never asked for VPN support —
("create", "Conflict. The container name \"/triple-c-x\" is already in use"), // that came from somewhere else and must not be misattributed.
("start", "Docker responded with status code 404: No such container"), let name_clash = "Conflict. The container name \"/triple-c-x\" is already in use";
("start", "error gathering device information while adding custom device \"/dev/dri/card0\""), assert_eq!(
] { explain_create_failure(name_clash, true),
assert_eq!( format!("Failed to create container: {}", name_clash)
explain_container_failure(action, err), );
format!("Failed to {} container: {}", action, err),
"{} should pass through untouched", let tun_err = "no such file or directory: /dev/net/tun";
err assert_eq!(
); explain_create_failure(tun_err, false),
} format!("Failed to create container: {}", tun_err)
);
} }
#[test] #[test]
-1
View File
@@ -135,7 +135,6 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"), ("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"), ("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
("/opt/mission-control", "Mission Control (Flight Control)"), ("/opt/mission-control", "Mission Control (Flight Control)"),
("/usr/bin/wg", "VPN tooling (WireGuard, for the VPN Support toggle)"),
]; ];
/// Headroom demanded on Docker's storage backend on top of the measured /// Headroom demanded on Docker's storage backend on top of the measured
+5 -6
View File
@@ -148,14 +148,13 @@ pub struct Project {
/// Grant the container what a VPN client needs to build a tunnel: /// Grant the container what a VPN client needs to build a tunnel:
/// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard /// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard
/// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard, /// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard,
/// OpenVPN) installs and runs but its connection attempt hangs until it /// OpenVPN, Tailscale) installs and runs but its connection attempt hangs
/// times out, because it cannot create the tunnel interface or touch the /// until it times out, because it cannot create the tunnel interface or
/// routing table. /// touch the routing table.
/// ///
/// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the /// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the
/// container reconfigure its own network stack, which reaches further than /// container reconfigure its own network stack, which is a meaningful step
/// it sounds — see `vpn_host_config` for what it does and does not confer. /// out of the default sandbox. Unlike `auth_bridge_enabled` this *is*
/// Unlike `auth_bridge_enabled` this *is*
/// container state, so it carries a `triple-c.vpn-support` label and is /// container state, so it carries a `triple-c.vpn-support` label and is
/// compared in `container_needs_recreation` — capabilities and devices are /// compared in `container_needs_recreation` — capabilities and devices are
/// fixed at creation and can only change by recreating the container. /// fixed at creation and can only change by recreating the container.
@@ -1,94 +0,0 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
import { render, screen, fireEvent } from "@testing-library/react";
import RuntimeSection from "./RuntimeSection";
import type { Project } from "../../../../lib/types";
const baseProject: Project = {
id: "p1",
name: "api-server",
paths: [{ host_path: "/src/api", mount_name: "api" }],
container_id: null,
status: "stopped",
backend: "anthropic",
bedrock_config: null,
ollama_config: null,
llamacpp_config: null,
openai_compatible_config: null,
allow_docker_access: false,
sandbox_mode_enabled: true,
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: true,
full_permissions: false,
permission_mode: null,
ssh_key_path: null,
ca_cert_path: null,
git_token: null,
git_user_name: null,
git_user_email: null,
custom_env_vars: [],
port_mappings: [],
claude_instructions: null,
claude_code_settings: null,
renamed_session_names: {},
created_at: "2026-01-01T00:00:00Z",
updated_at: "2026-01-01T00:00:00Z",
};
const VPN = "VPN support";
const save = vi.fn().mockResolvedValue(true);
function renderSection(over: Partial<Project> = {}, disabled = false) {
return render(
<RuntimeSection
project={{ ...baseProject, ...over }}
save={save}
disabled={disabled}
disabledReason="Container must be stopped to change this setting."
/>,
);
}
describe("RuntimeSection — VPN support toggle", () => {
beforeEach(() => vi.clearAllMocks());
it("saves only the VPN flag when switched on", () => {
renderSection();
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: true });
});
it("saves the flag off again, rather than dropping the key", () => {
// Off has to be written explicitly: the container carries a
// `triple-c.vpn-support` label either way, and an absent value would leave
// the capability granted.
renderSection({ vpn_support_enabled: true });
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: false });
});
it("reflects the project's current state", () => {
renderSection({ vpn_support_enabled: true });
expect(screen.getByRole("switch", { name: VPN })).toBeChecked();
});
it("cannot be changed while the container is running", () => {
// Capabilities and devices are fixed at creation, so this setting is gated
// on the container being stopped along with the rest of the tab.
renderSection({}, true);
const toggle = screen.getByRole("switch", { name: VPN });
expect(toggle).toBeDisabled();
fireEvent.click(toggle);
expect(save).not.toHaveBeenCalled();
});
it("warns that the change recreates the container", () => {
renderSection();
expect(
screen.getByText(/recreates the container on its next start/i),
).toBeInTheDocument();
});
});
@@ -59,7 +59,7 @@ export default function RuntimeSection({
<SwitchRow <SwitchRow
label="VPN support" label="VPN support"
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched. Changing this recreates the container on its next start — the home and .claude volumes are preserved." hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN, Tailscale) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched."
control={ control={
<Toggle <Toggle
label="VPN support" label="VPN support"
-73
View File
@@ -34,9 +34,6 @@ RUN for i in 1 2 3 4 5; do \
cron \ cron \
bubblewrap \ bubblewrap \
socat \ socat \
iproute2 \
wireguard-tools \
iptables \
&& rm -rf /var/lib/apt/lists/* && rm -rf /var/lib/apt/lists/*
# `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither # `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither
@@ -45,76 +42,6 @@ RUN for i in 1 2 3 4 5; do \
# corporate CA, no matter what the system trust store says. entrypoint.sh # corporate CA, no matter what the system trust store says. entrypoint.sh
# degrades to a warning if it is ever missing. # degrades to a warning if it is ever missing.
# `iproute2`, `wireguard-tools` and `iptables` above are what the VPN support
# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a
# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to
# add a route, and without `wg` no way to build the tunnel those two exist to
# serve — a capability with nothing able to use it.
#
# They are baked rather than left to a runtime `apt-get install` for the same
# reason as the Playwright libraries below: the writable layer is re-paid after
# every Reset and lost on base-image migration. A hand-installed `wg` therefore
# works right up until an upgrade, then disappears and takes the tunnel with it
# — silently, since a VPN that fails to come up looks exactly like one that was
# never started.
#
# Measured against the *current base image*, since a bare ubuntu:24.04 also
# pulls libelf1t64 and netbase, which this base already has, and so over-reports
# by ~258 kB: **+12 packages, 7,203 kB on amd64**. The same set on arm64 is
# ~14.4 MB — the package list is identical on both arches, the binaries are
# simply larger (measured as 14.7 MB on arm64 ubuntu:24.04, less that 258 kB).
#
# ## Why `iptables`, and not `nftables`
#
# `wireguard-tools` declares `Recommends: nftables | iptables`, which the
# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s
# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e.
# every stock full-tunnel config every provider hands out — and it shells out to
# a firewall backend with no `type -p` guard. Measured with neither installed:
#
# [#] iptables-restore -n
# /usr/bin/wg-quick: line 32: iptables-restore: command not found
# wg-quick EXIT=127
#
# `nftables` looks like the better pick — wg-quick prefers it, it is first in
# that Recommends, it is half the size — and it is the wrong one. wg-quick picks
# nft *unconditionally* when present (`if type -p nft`, line 241), so installing
# it makes the iptables path unreachable; and its nft ruleset needs three
# expression families where the iptables path needs one. Isolating them on a
# WSL2 host, the two connmark rules install fine and this is what fails:
#
# nft add rule ... fib saddr type != local drop
# Error: Could not process rule: No such file or directory
# ^^^^^^^^^^^^^^ needs nft_fib_ipv4
#
# That matters because of how the two hosts we ship to are configured. From
# LinuxKit's kernel config — Docker Desktop for Mac, identical on both arches:
#
# CONFIG_NETFILTER_XT_CONNMARK=y <- the iptables path works
# # CONFIG_NFT_FIB_IPV4 is not set <- the nft path does not
#
# So shipping `nftables` would forfeit the platform it was meant to fix. With
# `iptables`, full tunnels work on native Linux, on Docker Desktop for Mac, and
# on WSL2 kernels from 6.6 (which added xt_CONNMARK as a module). Only WSL2
# older than that is left out, and nothing installable here changes it — the way
# out there is to add the routes with `ip route` instead of using `wg-quick`,
# which is what the pia-vpn skill does on every platform.
#
# ## What this still does not fix
#
# `wireguard-tools` only *Suggests* `openresolv | resolvconf`, so neither is
# installed, and every provider's stock config carries a `DNS =` line. That
# fails in `set_dns()`, *before* the firewall step, so it takes split tunnels
# down too:
#
# [#] resolvconf -a wg0 -m 0 -x
# /usr/bin/wg-quick: line 32: resolvconf: command not found
#
# Deliberately not fixed here: `openresolv` has no installation candidate on
# noble, and `resolvconf` resolves only by pulling in systemd-resolved — a
# resolver daemon and systemd units, into a container with no systemd. Strip the
# `DNS =` line and set the resolver another way. Documented in HOW-TO-USE.md.
# Remove default ubuntu user to free UID 1000 for host-user remapping # Remove default ubuntu user to free UID 1000 for host-user remapping
RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \ RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \
&& if getent group ubuntu >/dev/null 2>&1; then groupdel ubuntu 2>/dev/null || true; fi && if getent group ubuntu >/dev/null 2>&1; then groupdel ubuntu 2>/dev/null || true; fi