Compare commits

..
Author SHA1 Message Date
shadow-testandClaude Opus 5 ab2c75d0b2 Ship a firewall backend, and correct three claims review disproved
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 3s
Build App (Preview) / build-macos (pull_request) Successful in 2m46s
Build App (Preview) / build-linux (pull_request) Successful in 7m22s
Build App (Preview) / build-windows (pull_request) Successful in 7m43s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
Build Container / build-container (pull_request) Successful in 11m20s
Review of #28 found the iptables exclusion was justified by a false premise,
and I confirmed it: `wireguard-tools` declares `Recommends: nftables | iptables`,
`--no-install-recommends` strips it, and `wg-quick`'s add_default() shells out
to a firewall backend with no `type -p` guard. Measured on the image as this PR
shipped it:

    [#] iptables-restore -n
    /usr/bin/wg-quick: line 32: iptables-restore: command not found
    wg-quick EXIT=127

That fires for `AllowedIPs = 0.0.0.0/0` — every stock full-tunnel config from
every provider — not for a desktop client's killswitch as the comment claimed.
Split tunnels are unaffected.

Ship `nftables` rather than `iptables`: wg-quick prefers it (`type -p nft`, so
with both installed iptables is dead weight), it is first in the package's own
Recommends, and it is half the size.

The review's proposed fix stopped there; it does not hold. Adding nftables does
not make wg-quick work on this host, and neither does iptables:

    Warning: Extension CONNMARK revision 0 not supported, missing kernel module?

`Table=auto` routes by fwmark and needs xt_CONNMARK from the *host* kernel.
WSL2 has none and containers have no /lib/modules to load one from. So this
fixes native Linux and Docker Desktop for Mac — which other WHP users are on —
and cannot fix Docker Desktop for Windows, where the answer is to add routes
with `ip route` directly. Documented rather than left to be rediscovered.

Also from review:

- "`ip` and `wg` are always present" was false. A project keeps the base image
  it was first built from, so this reaches new projects only. Reworded to match
  the wording already used for the Playwright libraries, and `/usr/bin/wg` added
  to FEATURE_PROBES so an existing project is *told* it is missing VPN tooling
  and prompted to migrate, rather than finding out via `wg: command not found`.
- "no client is installed" contradicted shipping `wg` four lines earlier. The
  true claim is that no tunnel is configured or started.
- The size figure measured against bare ubuntu:24.04, which over-counts by the
  ~209 kB of libelf1t64 the real base already has, and covered one arch. Now
  measured against the current base on amd64 and stated for arm64 too, per the
  standard CLAUDE.md sets for the Playwright layer.
- `/run` persistence conflated two mechanisms: same-container files on a
  stop/start, `docker commit` on a recreation. Both stated, plus the corollary
  that key material written to /run ends up inside a snapshot image — observed,
  a `wg.priv` was already sitting in one.
- The DNS bullet presented a Docker Desktop address as the general case. Now
  leads with the mechanism, notes 127.0.0.11 on a user-defined network is
  unaffected, and adds the two things the advice omitted: a resolver the tunnel
  can reach (or it leaks every query), and pinning the endpoint via the old
  gateway (or the tunnel routes through itself).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 13:55:44 -07:00
shadow-testandClaude Opus 5 00937745f7 Ship the tools the VPN toggle grants capability for
Build Container / build-container (pull_request) Successful in 11m28s
`vpn_support_enabled` hands a project CAP_NET_ADMIN and /dev/net/tun, and the
image then contains no `ip` and no `wg` — a capability with nothing able to
exercise it. Bake `iproute2` and `wireguard-tools` (~4.3 MB with deps).

They belong in the image rather than a runtime install for the reason the
Dockerfile already gives for the Playwright libraries: the writable layer is
lost on base-image migration. A hand-installed `wg` works until an upgrade and
then vanishes, which presents as a tunnel that will not come up rather than as
a missing package. One project only had `ip` at all because MariaDB pulled in
iproute2 as a transitive dependency.

`iptables` stays out. Only a desktop client's killswitch wants it, and those
clients need a GUI the container cannot provide.

Also correct three things the docs left users to discover:

- the toggle grants capability and routes nothing, which is being reported as
  the default network "not routing through the VPN automatically"
- no tunnel survives a restart, and `/run` state riding the snapshot makes it
  look as though one did while traffic goes out the real address
- a full tunnel captures the Docker resolver, which sits outside the
  container's subnet, and takes DNS down with it — Claude Code then reports a
  connection failure because it cannot resolve api.anthropic.com, and a health
  check aimed at an IP literal passes throughout

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-17 12:38:37 -07:00
jknapp 01e72e4785 Merge pull request 'Let a project's container run a VPN client' (#27) from feat/vpn-support into main
Build App / compute-version (push) Successful in 3s
Build App / build-macos (push) Successful in 2m37s
Build App / build-linux (push) Successful in 5m33s
Build App / build-windows (push) Successful in 6m11s
Build App / create-tag (push) Successful in 6s
Build App / sync-to-github (push) Successful in 52s
2026-08-14 15:30:03 +00:00
shadow-testandClaude Opus 5 2b35aa8c16 Explain a missing tun device where the failure actually happens
Build App (Preview) / compute-version (pull_request) Successful in 3s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m37s
Build App (Preview) / build-linux (pull_request) Successful in 5m30s
Build App (Preview) / build-windows (pull_request) Successful in 5m55s
Build App (Preview) / prune-previews (pull_request) Successful in 3s
Review caught that the device guard was wired to the wrong call. The
daemon does not resolve `--device` at create: verified against Docker
29.7, `docker create --device /dev/does-not-exist` succeeds and prints an
id, and runc only resolves the device — and validates sysctls — when it
builds the container. So on a host with no tun module the create returns
fine and `start` fails, which means the explanation never ran and the
user saw the raw daemon string naming a path they would go looking for on
the wrong machine. The unit tests fed the create-side string straight in,
so they confirmed a function no real failure could reach.

Move the guard onto `start_container`, covering create as well in case a
future daemon checks earlier. It no longer takes `vpn_support_enabled` —
`start_container` has a container id and no project, and nothing else in
Triple-C ever requests a device, so an error naming /dev/net/tun is
unambiguous on its own. The test now uses the daemon's verbatim message
via bollard's real Display format.

Also from review:

  * Soften the security claim. Docker does not enable user-namespace
    remapping by default, so this is a real CAP_NET_ADMIN in the initial
    user namespace with only the network namespace confining it. It
    cannot touch host interfaces, but "confers no authority outside the
    container" was too strong: within its namespace it can set
    promiscuous mode and add addresses, routes and NAT on the shared
    docker0 segment, which puts sibling containers — the LiteLLM gateway
    among them — within ARP-spoofing reach, and it can flush netfilter
    rules sandbox mode may rely on. Said plainly in the code, CLAUDE.md
    and HOW-TO-USE.
  * Drop Tailscale from the list of clients needing this. Its
    --tun=userspace-networking mode needs neither the capability nor the
    device, and listing it invites granting NET_ADMIN for nothing.
  * Say in the toggle's own hint that changing it recreates the
    container, matching how every other recreation-triggering setting is
    labelled. The tab's generic "stop the container first" chip does not
    tell the user what is about to happen.
  * Add RuntimeSection tests: saves on, saves off explicitly rather than
    dropping the key, reflects state, is disabled while running, and
    carries the recreation warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:15:05 -07:00
shadow-testandClaude Opus 5 65a3d4eb29 Let a project's container run a VPN client
Build App (Preview) / compute-version (pull_request) Successful in 4s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m39s
Build App (Preview) / build-linux (pull_request) Successful in 7m10s
Build App (Preview) / build-windows (pull_request) Successful in 6m33s
Build App (Preview) / prune-previews (pull_request) Successful in 4s
A VPN client installed in a container today starts, runs, and then hangs
until its connection times out. Nothing reports an error: a default
container has no /dev/net/tun to open and no CAP_NET_ADMIN to add an
interface or a route with, and clients surface that as a generic timeout
rather than a permissions failure.

Add an opt-in per-project "VPN support" switch granting the three things
a tunnel needs. They are useless individually, which is why
vpn_host_config() defines the set in one place and the tests assert all
of it:

  * CAP_NET_ADMIN — Docker's default bounding set has net_raw but not
    net_admin, so a client can ping but never connect.
  * /dev/net/tun — passed through from the host so the kernel's tun
    module backs it, rather than mknod-ed inside.
  * net.ipv4.conf.all.src_valid_mark — WireGuard's wg-quick sets this and
    cannot from inside a container, /proc/sys being read-only, so its
    handshakes are dropped by reverse-path filtering.

Off by default and deliberately opt-in: NET_ADMIN lets anything in the
container reconfigure that container's network stack. It is namespaced —
no authority over the host's interfaces or any other container.

Capabilities and devices are fixed when a container is created, so this
is container state and takes the label-and-compare treatment.
triple-c.vpn-support is written unconditionally, false included, for the
usual docker commit reason: a true stamped once would ride the snapshot
image into every future container and make the switch impossible to turn
back off. A missing label reads as false and off is byte-identical to
today, so no existing project is churned.

Requesting the device fails at creation when the host kernel has no tun
module, which would otherwise surface as a project that simply refuses to
start. explain_create_failure() rewrites that one error to name the
switch and the Docker-Desktop-VM-versus-your-machine distinction, and
leaves every other failure untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 08:06:39 -07:00
10 changed files with 543 additions and 3 deletions
+62 -1
View File
@@ -203,7 +203,8 @@ docker exec stdout → tokio task → emit("terminal-output-{sessionId}") → li
### Container (`container/`)
- **`Dockerfile`** — Ubuntu 24.04 base with Claude Code, Node.js 22, Python 3.12, Rust, Docker CLI, git, gh, AWS CLI v2, ripgrep, pnpm, uv, ruff pre-installed, plus the shared
libraries a browser links against (see below)
libraries a browser links against (see below) and the VPN tooling the `vpn_support_enabled`
toggle grants capability for (`iproute2`, `wireguard-tools`, `nftables`)
- **Browser runtime libraries are baked in; browser *binaries* are not.** A layer runs
`npx --yes playwright@latest install-deps chromium` as root, so Playwright names its own
dependencies and the list cannot rot against Ubuntu 24.04's `t64` renames or a new Chromium
@@ -273,6 +274,66 @@ migration and Reset. Four things here are not obvious:
actively **removes** `triple-c-*.crt` when the setting is cleared — `/usr/local/share` rides the
project's snapshot image, so turning the feature off has to undo, not merely stop.
### VPN support (`vpn_support_enabled`, `docker/container.rs`)
An opt-in per-project switch granting the container what a VPN client needs to build a tunnel.
`vpn_host_config()` is the single definition of what that means, and it is unit-tested because a
container is created once by a very long function where a dropped capability is invisible.
- **All three pieces or none.** `CAP_NET_ADMIN` (Docker's default set has `net_raw` but *not*
`net_admin`, so a client can ping but never connect), the `/dev/net/tun` device (absent
entirely from a default container — nothing to open even with the capability), and
`net.ipv4.conf.all.src_valid_mark=1` (WireGuard's `wg-quick` sets it and cannot from inside a
container, since `/proc/sys` is read-only, so handshake packets die to reverse-path filtering).
Any two without the third still presents as a connection that hangs to a timeout, which is why
the tests assert the whole set.
- **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun`
module has to back it.
- **A missing device fails at `start`, not `create` — verified against Docker 29.7.** `docker
create --device /dev/does-not-exist` succeeds and prints an id; runc resolves the device (and
validates sysctls) only when it builds the container. So the guard belongs on the start path:
`explain_container_failure()` covers both and is called from `start_container`, where it has a
container id and no project — which is why it keys off the error naming `/dev/net/tun` rather
than off `vpn_support_enabled`. Nothing else in Triple-C requests a device, so that is
unambiguous. A version of this check wired to `create` alone is dead code that looks correct.
- **`NET_ADMIN` here is not user-namespaced.** Docker does not enable userns remapping by default,
so only the *network* namespace confines it: no reach onto host interfaces, but promiscuous
mode, arbitrary addresses/routes/NAT on the shared `docker0` segment (sibling containers, the
LiteLLM gateway among them, are ARP-spoofable), netlink-triggered host module auto-load, and
enough authority to flush in-container netfilter rules that sandbox mode may rely on. Keep the
code comments honest about this — an earlier draft claimed it "confers no authority" outside the
container, which is too strong.
- **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual
`docker commit` reason: a `true` stamped once would ride the snapshot image into every future
container and make the switch impossible to turn off.
- Off is byte-identical to a container created before the feature existed, and a missing label
reads as `false`, so no existing project is churned.
- **The toggle grants capability and stops there — it routes nothing.** `vpn_host_config()` returns
a cap, a device and a sysctl; no client is installed, no route is touched, no tunnel is started
or restored. Users read the name as "turn the VPN on" and report the default network not routing
through it as a bug. It isn't, and the docs say so explicitly; keep it that way.
- **The tooling is baked, not installed at runtime.** `iproute2` and `wireguard-tools` are in
`container/Dockerfile` because a runtime install lands in the writable layer and is lost on
base-image migration — leaving a project holding the capability with nothing able to exercise it,
and no error that points at why. `iptables` is deliberately absent; see the Dockerfile comment.
- **Anything built on this fails open.** The network namespace is rebuilt on every start and no
service manager runs inside, so a tunnel never survives stop/start or recreation — while leftover
`/run` state makes it look as though it did. Note the two different mechanisms: `/run` is in the
writable layer, so on a stop/start it is simply the same container's files, and on a recreation
`docker commit` has carried it into the snapshot. Traffic silently reverts to the real address.
Any future autostart or killswitch work starts here.
- **`/run` riding the snapshot means a VPN client's key material can end up in an image.** Verified:
a fresh container off the whp snapshot already contained the `wg.priv` a previous tunnel left in
`/run`. Anything writing key material there inherits the problem — the same `docker commit`
hazard as `triple-c.git-token-hash` and the custom-env fingerprint, in a directory that looks
ephemeral and is not. A VPN client that does this should delete its key on teardown.
- **`wg-quick` full tunnels need `xt_CONNMARK` from the host kernel**, which WSL2 does not have and
a container cannot load; `Recommends: nftables | iptables` is also stripped by
`--no-install-recommends`, so `nftables` is baked explicitly. See the Dockerfile comment — the
short version is that shipping the backend fixes native Linux and Docker Desktop for Mac, nothing
fixes Docker Desktop for Windows, and adding the routes directly with `ip route` sidesteps it on
all three.
### Container Lifecycle
Containers use a **stop/start** model (not create/destroy). Installed packages persist across stops. The `.claude` config dir uses a named Docker volume (`triple-c-claude-config-{projectId}`), nested inside the home volume (`triple-c-home-{projectId}`), so OAuth tokens and Claude Code config survive container stop/start *and* container recreation.
+66
View File
@@ -471,6 +471,72 @@ When enabled, the host Docker socket is mounted into the container so Claude Cod
> Toggling this requires stopping and restarting the container to take effect.
### VPN Support
When enabled, the container is given the three things a VPN client needs to build a tunnel:
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
sysctl that WireGuard requires. This is **off by default**.
The `ip`, `wg` and `nft` commands ship in the container image so there is something able to use
them. If your project's container was created from an older base image it will not have them, and
`wg` will simply not be found — **migrating the project onto the current base image** is what picks
them up. Installing them by hand with `sudo apt install wireguard-tools` works in the meantime, but
lives in the writable layer, so it is undone by a **Reset** and by a migration.
**This setting makes a tunnel possible; it does not make one.** Nothing is connected, no traffic is
redirected, and no tunnel is configured or started on your behalf. Enabling it and expecting the
container's traffic to start leaving through a VPN is the most common misreading of what it does —
configuring a tunnel and routing traffic into it remains yours to do.
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
the connection attempt **hangs until it times out** — a default container has no tun device to open
and no permission to add an interface or a route, and most clients report that as a generic timeout
rather than a permissions error.
Things worth knowing:
- Tailscale is the exception: in its `--tun=userspace-networking` mode it needs neither the
capability nor the device, so leave this off if that is all you want.
- `NET_ADMIN` applies to the container's **own** network namespace — it cannot touch the host's
interfaces. It is not nothing, though: within that namespace anything in the container can set
promiscuous mode and add arbitrary addresses, routes and firewall rules on the Docker bridge it
shares with your other containers, and it can flush firewall rules that sandbox mode relies on.
Grant it per project, to projects that need it.
- The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is
the Linux VM, not your own machine. If it is missing, the container is created but fails to
**start**, with an error naming `/dev/net/tun` and pointing back at this setting.
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
- **No tunnel survives a restart.** The network namespace is built fresh every time the container
starts, and there is no service manager inside to reconnect anything. Leftover state under `/run`
makes it *look* like the tunnel is still configured — that directory is in the container's
writable layer, so it is simply still there after a stop/start, and `docker commit` carries it
into the snapshot that a recreation is built from. Either way the interface and its routes are
gone and traffic goes out your real address again, with no error and nothing visibly different.
Re-establish it after every start, and check rather than assume.
- **A full tunnel breaks DNS unless the client is told to leave private ranges alone.** Your
resolver is whatever `/etc/resolv.conf` says, and if that address is outside the container's own
subnet then a default route of `0.0.0.0/0` — or a `0.0.0.0/1` plus `128.0.0.0/1` pair — captures
it and sends every lookup into a tunnel that cannot carry it. Under Docker Desktop it is
`192.168.65.7`, which is exactly that case; on a user-defined Docker network it is `127.0.0.11`,
which is loopback and unaffected. Check yours rather than assuming. The symptom when it bites is
total: Claude Code reports it cannot connect, because it cannot resolve `api.anthropic.com`.
Route `10.0.0.0/8`, `172.16.0.0/12`, `192.168.0.0/16` and `169.254.0.0/16` via the original
gateway — and give the tunnel a resolver it can actually reach, normally the VPN provider's own,
or you have a tunnel that leaks every DNS query outside itself. Also pin the VPN endpoint's own
address via the original gateway, or the tunnel's encrypted packets try to route through the
tunnel. Note that a health check which fetches an IP literal such as `1.1.1.1` passes cleanly
while DNS is broken — resolve a name instead.
- **`wg-quick` cannot bring up a full tunnel on Docker Desktop for Windows.** Its `Table=auto` mode
routes by firewall mark and needs `xt_CONNMARK` from the host kernel, which WSL2's does not have
and a container cannot load. Split tunnels (a specific `AllowedIPs`) work fine, as does adding
the routes yourself with `ip route`. Native Linux and Docker Desktop for Mac are unaffected.
> This setting can only be changed when the container is stopped. Capabilities and devices are
> fixed when a container is created, so toggling it recreates the container on the next start.
> Recreation preserves the home and `.claude` volumes — it is not a Reset.
### Mission Control
Toggle **Mission Control** to integrate Flight Control — an AI-first development methodology bundled with Triple-C — into the project. When enabled:
+224 -2
View File
@@ -798,6 +798,111 @@ async fn resolve_base_image_id(image_name: &str, base_image_name: &str) -> Strin
.unwrap_or_default()
}
/// The `/dev/net/tun` character device, as it is named on both sides.
const TUN_DEVICE: &str = "/dev/net/tun";
/// The `HostConfig` fields "VPN support" contributes: `CapAdd`, `Devices`,
/// `Sysctls` — in that order.
type VpnHostConfigParts = (
Option<Vec<String>>,
Option<Vec<bollard::models::DeviceMapping>>,
Option<HashMap<String, String>>,
);
/// The three host-config pieces a VPN client needs, or all-`None` when the
/// project has not opted in.
///
/// Returned as a triple rather than set inline so the exact shape is unit
/// testable — a container is created once, by a very long async function, and a
/// silently-dropped capability looks identical to a VPN server that is simply
/// unreachable.
///
/// All three are required together and each fails differently on its own:
/// * **`CAP_NET_ADMIN`** — without it the client cannot create an interface or
/// write a route. Docker's default bounding set grants `net_raw` but not
/// `net_admin`, which is why a client can ping but never connect.
/// * **`/dev/net/tun`** — the device is absent from a default container, so
/// there is nothing to open even with the capability. It is passed through
/// from the host rather than `mknod`-ed inside, so the kernel's `tun` module
/// backs it.
/// * **`net.ipv4.conf.all.src_valid_mark`** — WireGuard's own `wg-quick` sets
/// this, and cannot from inside a container (`/proc/sys` is read-only), so
/// its handshake packets are dropped by reverse-path filtering. Harmless for
/// OpenVPN-based clients, so it is set unconditionally with the rest.
///
/// What it costs, stated accurately: Docker does not enable user-namespace
/// remapping by default, so this is a real `CAP_NET_ADMIN` in the *initial*
/// user namespace and only the **network** namespace confines it. It cannot
/// touch the host's interfaces, but within its own namespace it can set
/// promiscuous mode and add arbitrary addresses, routes and NAT rules on the
/// shared `docker0` L2 segment — which puts sibling containers (the LiteLLM
/// gateway among them) within reach of ARP spoofing, and lets netlink trigger
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
/// Hence opt-in, per project, rather than on for everyone.
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
if !enabled {
return (None, None, None);
}
let devices = vec![bollard::models::DeviceMapping {
path_on_host: Some(TUN_DEVICE.to_string()),
path_in_container: Some(TUN_DEVICE.to_string()),
cgroup_permissions: Some("rwm".to_string()),
}];
let sysctls = HashMap::from([(
"net.ipv4.conf.all.src_valid_mark".to_string(),
"1".to_string(),
)]);
(
Some(vec!["NET_ADMIN".to_string()]),
Some(devices),
Some(sysctls),
)
}
/// Turn the daemon's device-passthrough failure into an explanation.
///
/// **This fires on `start`, not `create`.** Verified against Docker 29.7:
/// `docker create --device /dev/does-not-exist` succeeds and prints an id; the
/// device is only resolved when runc builds the container, so the failure lands
/// on the *next* call. Sysctls validate at the same point. Anything that
/// inspects only the create path will never see it — which is why both paths
/// route through here and the tests exercise the start-side string.
///
/// Unmapped, this reads as `Failed to start container: Docker responded with
/// status code 500: error gathering device information while adding custom
/// device "/dev/net/tun": no such file or directory` — a path the user will go
/// looking for on the wrong machine, since with Docker Desktop the relevant
/// host is the Linux VM rather than their own, and with nothing pointing back
/// at the switch that caused it.
///
/// Deliberately not gated on `vpn_support_enabled`: nothing else in Triple-C
/// ever asks for a device, so an error naming `/dev/net/tun` can only have come
/// from a container created with the switch on. That keeps the check usable
/// from [`start_container`], which has a container id and no project.
fn explain_container_failure(action: &str, err: &str) -> String {
let device_missing = err.contains(TUN_DEVICE)
&& (err.contains("no such file or directory")
|| err.contains("No such file or directory")
|| err.contains("error gathering device information"));
if device_missing {
return format!(
"Failed to {} container: the Docker host has no {} device, which \
\"VPN support\" requires. The host kernel needs the `tun` module \
loaded (on Docker Desktop that is the Linux VM, not your own \
machine). Turn VPN support off in Config → Runtime to start this \
project without it. Original error: {}",
action, TUN_DEVICE, err
);
}
format!("Failed to {} container: {}", action, err)
}
pub async fn create_container(
project: &Project,
docker_socket_path: &str,
@@ -1375,6 +1480,13 @@ pub async fn create_container(
labels.insert("triple-c.image".to_string(), image_name.to_string());
labels.insert("triple-c.timezone".to_string(), timezone.unwrap_or("").to_string());
labels.insert("triple-c.mission-control".to_string(), project.mission_control_enabled.to_string());
// Capabilities, devices and sysctls are fixed at creation, so this is
// container state and gets the label-and-compare treatment. Written
// unconditionally (`false`, not omitted) because `docker commit` copies
// container labels onto the snapshot image: a `true` stamped once would
// otherwise ride that snapshot into every future container and make the
// switch impossible to turn back off.
labels.insert("triple-c.vpn-support".to_string(), project.vpn_support_enabled.to_string());
labels.insert("triple-c.permission-mode".to_string(),
project.effective_permission_mode().as_env_value().to_string());
labels.insert("triple-c.custom-env-fingerprint".to_string(), custom_env_fingerprint.clone());
@@ -1443,10 +1555,15 @@ pub async fn create_container(
labels.insert((*key).to_string(), (*value).to_string());
}
let (cap_add, devices, sysctls) = vpn_host_config(project.vpn_support_enabled);
let host_config = HostConfig {
mounts: Some(mounts),
port_bindings: if port_bindings.is_empty() { None } else { Some(port_bindings) },
init: Some(true),
cap_add,
devices,
sysctls,
..Default::default()
};
@@ -1476,7 +1593,7 @@ pub async fn create_container(
let response = docker
.create_container(Some(options), config)
.await
.map_err(|e| format!("Failed to create container: {}", e))?;
.map_err(|e| explain_container_failure("create", &e.to_string()))?;
Ok(response.id)
}
@@ -1486,7 +1603,7 @@ pub async fn start_container(container_id: &str) -> Result<(), String> {
docker
.start_container(container_id, None::<StartContainerOptions<String>>)
.await
.map_err(|e| format!("Failed to start container: {}", e))
.map_err(|e| explain_container_failure("start", &e.to_string()))
}
pub async fn stop_container(container_id: &str) -> Result<(), String> {
@@ -2367,6 +2484,19 @@ pub async fn container_needs_recreation(
return Ok(true);
}
// ── VPN support (NET_ADMIN + /dev/net/tun + sysctl) ───────────────────
// A container's capabilities, devices and sysctls are set at creation and
// cannot be changed on a running or stopped container, so recreation is the
// only way a toggle here takes effect. A missing label means the container
// predates the feature, which is the same thing as having it off — so
// existing projects are not churned until someone actually turns it on.
let expected_vpn = project.vpn_support_enabled.to_string();
let container_vpn = get_label("triple-c.vpn-support").unwrap_or_else(|| "false".to_string());
if container_vpn != expected_vpn {
log::info!("VPN support mismatch (container={:?}, expected={:?})", container_vpn, expected_vpn);
return Ok(true);
}
// ── Permission mode ────────────────────────────────────────────────────
// The mode is injected as the TRIPLE_C_PERMISSION_MODE env var, and
// container env can only change by recreating the container. A missing
@@ -2616,6 +2746,98 @@ mod tests {
assert_eq!(fp, "");
}
#[test]
fn vpn_support_off_touches_nothing_in_the_host_config() {
// The default must stay byte-identical to a container created before the
// feature existed, or every project recreates on the next start.
let (cap_add, devices, sysctls) = vpn_host_config(false);
assert_eq!(cap_add, None);
assert_eq!(devices, None);
assert_eq!(sysctls, None);
}
#[test]
fn vpn_support_on_grants_all_three_pieces() {
// Each is useless without the others — a client with the capability but
// no device, or the device but no capability, still times out — so this
// asserts the whole set rather than any one of them.
let (cap_add, devices, sysctls) = vpn_host_config(true);
assert_eq!(cap_add, Some(vec!["NET_ADMIN".to_string()]));
let devices = devices.expect("the tun device must be passed through");
assert_eq!(devices.len(), 1);
assert_eq!(devices[0].path_on_host.as_deref(), Some(TUN_DEVICE));
assert_eq!(devices[0].path_in_container.as_deref(), Some(TUN_DEVICE));
assert_eq!(devices[0].cgroup_permissions.as_deref(), Some("rwm"));
assert_eq!(
sysctls
.expect("wireguard needs src_valid_mark")
.get("net.ipv4.conf.all.src_valid_mark")
.map(String::as_str),
Some("1")
);
}
#[test]
fn vpn_support_never_grants_more_than_net_admin() {
// NET_ADMIN is already a step out of the sandbox. Anything else added
// here (SYS_ADMIN, or a blanket privileged flag) would be a much larger
// one, so pin the set.
let (cap_add, _, _) = vpn_host_config(true);
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
}
/// What bollard actually hands us when a tun-less host rejects the device.
///
/// Captured verbatim from Docker 29.7: `docker create` with a missing
/// device **succeeds**, and this arrives from the subsequent `start`.
/// `DockerResponseServerError`'s Display is
/// `"Docker responded with status code {code}: {message}"` with the
/// daemon's message unaltered.
const REAL_TUN_ERROR: &str = "Docker responded with status code 500: error \
gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
#[test]
fn a_missing_tun_device_is_explained_on_the_path_that_actually_fails() {
// The start path is the one that matters: the daemon defers device
// resolution to runc, so create returns an id on a host with no tun
// module and only start fails. A version of this that checked create
// alone would be dead code.
let msg = explain_container_failure("start", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to start container:"), "{}", msg);
assert!(msg.contains("VPN support"), "should name the switch: {}", msg);
assert!(msg.contains("tun` module"), "should name the cause: {}", msg);
assert!(msg.contains("Config → Runtime"), "should say where to fix it: {}", msg);
assert!(msg.contains(REAL_TUN_ERROR), "should keep the original: {}", msg);
}
#[test]
fn the_same_explanation_covers_create_if_the_daemon_ever_checks_earlier() {
// Belt and braces — older and future daemons may validate at create.
let msg = explain_container_failure("create", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to create container:"), "{}", msg);
assert!(msg.contains("VPN support"), "{}", msg);
}
#[test]
fn unrelated_failures_are_left_alone() {
for (action, err) in [
("create", "Conflict. The container name \"/triple-c-x\" is already in use"),
("start", "Docker responded with status code 404: No such container"),
("start", "error gathering device information while adding custom device \"/dev/dri/card0\""),
] {
assert_eq!(
explain_container_failure(action, err),
format!("Failed to {} container: {}", action, err),
"{} should pass through untouched",
err
);
}
}
#[test]
fn the_orphan_sweep_only_ever_looks_at_our_own_untagged_images() {
// Both conditions are load-bearing. Without `dangling` the sweep would
+1
View File
@@ -135,6 +135,7 @@ pub const FEATURE_PROBES: &[(&str, &str)] = &[
("/usr/local/bin/triple-c-task-runner", "Scheduled task runner"),
("/usr/local/bin/triple-c-sso-refresh", "AWS SSO auto-refresh"),
("/opt/mission-control", "Mission Control (Flight Control)"),
("/usr/bin/wg", "VPN support (WireGuard tools)"),
];
/// Headroom demanded on Docker's storage backend on top of the measured
+17
View File
@@ -145,6 +145,22 @@ pub struct Project {
/// container-recreation label.
#[serde(default)]
pub browser_view_enabled: bool,
/// Grant the container what a VPN client needs to build a tunnel:
/// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard
/// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard,
/// OpenVPN) installs and runs but its connection attempt hangs until it
/// times out, because it cannot create the tunnel interface or touch the
/// routing table.
///
/// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the
/// container reconfigure its own network stack, which reaches further than
/// it sounds — see `vpn_host_config` for what it does and does not confer.
/// Unlike `auth_bridge_enabled` this *is*
/// container state, so it carries a `triple-c.vpn-support` label and is
/// compared in `container_needs_recreation` — capabilities and devices are
/// fixed at creation and can only change by recreating the container.
#[serde(default)]
pub vpn_support_enabled: bool,
/// Use the shared, long-lived Claude Code OAuth token (from
/// `claude setup-token`, held in the OS keychain) for this project instead
/// of requiring its own `claude login`. Only consulted when `backend` is
@@ -366,6 +382,7 @@ impl Project {
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: default_use_shared_auth_token(),
full_permissions: false,
permission_mode: None,
@@ -120,6 +120,15 @@ export default function OverviewTab({
{project.mission_control_enabled ? "ON" : "OFF"}
</span>
</span>
{/* Only when granted. It is off for nearly every project and an
always-present "VPN OFF" would be noise, but where it *is* on the
container holds NET_ADMIN, which is worth seeing at a glance. */}
{project.vpn_support_enabled && (
<span className="text-[var(--text-secondary)]">
VPN support{" "}
<span className="text-[var(--text-primary)] font-medium">ON</span>
</span>
)}
<button
type="button"
onClick={() => onOpenTab("config")}
@@ -0,0 +1,94 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
import { render, screen, fireEvent } from "@testing-library/react";
import RuntimeSection from "./RuntimeSection";
import type { Project } from "../../../../lib/types";
const baseProject: Project = {
id: "p1",
name: "api-server",
paths: [{ host_path: "/src/api", mount_name: "api" }],
container_id: null,
status: "stopped",
backend: "anthropic",
bedrock_config: null,
ollama_config: null,
llamacpp_config: null,
openai_compatible_config: null,
allow_docker_access: false,
sandbox_mode_enabled: true,
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: true,
full_permissions: false,
permission_mode: null,
ssh_key_path: null,
ca_cert_path: null,
git_token: null,
git_user_name: null,
git_user_email: null,
custom_env_vars: [],
port_mappings: [],
claude_instructions: null,
claude_code_settings: null,
renamed_session_names: {},
created_at: "2026-01-01T00:00:00Z",
updated_at: "2026-01-01T00:00:00Z",
};
const VPN = "VPN support";
const save = vi.fn().mockResolvedValue(true);
function renderSection(over: Partial<Project> = {}, disabled = false) {
return render(
<RuntimeSection
project={{ ...baseProject, ...over }}
save={save}
disabled={disabled}
disabledReason="Container must be stopped to change this setting."
/>,
);
}
describe("RuntimeSection — VPN support toggle", () => {
beforeEach(() => vi.clearAllMocks());
it("saves only the VPN flag when switched on", () => {
renderSection();
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: true });
});
it("saves the flag off again, rather than dropping the key", () => {
// Off has to be written explicitly: the container carries a
// `triple-c.vpn-support` label either way, and an absent value would leave
// the capability granted.
renderSection({ vpn_support_enabled: true });
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: false });
});
it("reflects the project's current state", () => {
renderSection({ vpn_support_enabled: true });
expect(screen.getByRole("switch", { name: VPN })).toBeChecked();
});
it("cannot be changed while the container is running", () => {
// Capabilities and devices are fixed at creation, so this setting is gated
// on the container being stopped along with the rest of the tab.
renderSection({}, true);
const toggle = screen.getByRole("switch", { name: VPN });
expect(toggle).toBeDisabled();
fireEvent.click(toggle);
expect(save).not.toHaveBeenCalled();
});
it("warns that the change recreates the container", () => {
renderSection();
expect(
screen.getByText(/recreates the container on its next start/i),
).toBeInTheDocument();
});
});
@@ -57,6 +57,19 @@ export default function RuntimeSection({
}
/>
<SwitchRow
label="VPN support"
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched. Changing this recreates the container on its next start — the home and .claude volumes are preserved."
control={
<Toggle
label="VPN support"
checked={project.vpn_support_enabled}
disabled={disabled}
onChange={(v) => save({ vpn_support_enabled: v })}
/>
}
/>
<SwitchRow
label="Mission Control"
hint="A web dashboard for monitoring and managing Claude sessions remotely."
+4
View File
@@ -33,6 +33,10 @@ export interface Project {
auth_bridge_enabled: boolean;
/** Opt in to the browser-view pane. Host-side only, like `auth_bridge_enabled`. */
browser_view_enabled: boolean;
/** Grant NET_ADMIN, /dev/net/tun and the WireGuard `src_valid_mark` sysctl so
* a VPN client inside the container can build a tunnel. Unlike the two flags
* above this is container state — changing it recreates the container. */
vpn_support_enabled: boolean;
/** Use the shared long-lived Claude Code token (from `claude setup-token`,
* held in the OS keychain) instead of this project's own `claude login`.
* Defaults to true; only applies when `backend` is "anthropic" and a token
+53
View File
@@ -34,6 +34,9 @@ RUN for i in 1 2 3 4 5; do \
cron \
bubblewrap \
socat \
iproute2 \
wireguard-tools \
nftables \
&& rm -rf /var/lib/apt/lists/*
# `libnss3-tools` above provides `certutil`. Chrome/Chromium read neither
@@ -42,6 +45,56 @@ RUN for i in 1 2 3 4 5; do \
# corporate CA, no matter what the system trust store says. entrypoint.sh
# degrades to a warning if it is ever missing.
# `iproute2`, `wireguard-tools` and `nftables` above are what the VPN support
# toggle (`vpn_support_enabled`) grants capability *for*. That toggle hands a
# project CAP_NET_ADMIN and /dev/net/tun; without `ip` there is then no way to
# add a route, and without `wg` no way to build the tunnel those two exist to
# serve — a capability with nothing able to use it.
#
# They are baked rather than left to a runtime `apt-get install` for the same
# reason as the Playwright libraries below: the writable layer is re-paid after
# every Reset and lost on base-image migration. A hand-installed `wg` therefore
# works right up until an upgrade, then disappears and takes the tunnel with it
# — silently, since a VPN that fails to come up looks exactly like one that was
# never started.
#
# Measured against the *current base image*, not a bare ubuntu:24.04 — the base
# already ships libelf1t64, so measuring on bare ubuntu over-counts by ~209 kB:
# +9 packages, 5,614 kB on amd64 (4,153 kB of that is iproute2+wireguard-tools,
# 1,461 kB is nftables). The same set on arm64 is 7,422 kB, measured against
# ubuntu:24.04 since the arm64 base is not cached here.
#
# ## Why `nftables` specifically
#
# `wireguard-tools` declares `Recommends: nftables | iptables`, which the
# `--no-install-recommends` above strips. That is not cosmetic: `wg-quick`'s
# `add_default()` runs whenever a config has `AllowedIPs = 0.0.0.0/0` — i.e.
# every stock full-tunnel config every provider hands out — and it shells out to
# a firewall backend with no `type -p` guard. Measured without one:
#
# [#] iptables-restore -n
# /usr/bin/wg-quick: line 32: iptables-restore: command not found
# wg-quick EXIT=127 (interface rolled back, split tunnels unaffected)
#
# `nftables` rather than `iptables` because `wg-quick` prefers it (`if type -p
# nft`, so with both installed iptables is dead weight), it is the first
# alternative in the package's own Recommends, and it is roughly half the size.
#
# This does NOT make `wg-quick`'s full-tunnel mode work everywhere. `Table=auto`
# routes by fwmark and needs connection-mark tracking from the *host* kernel:
#
# Warning: Extension CONNMARK revision 0 not supported, missing kernel module?
#
# WSL2's kernel has no `xt_CONNMARK` and containers have no /lib/modules to load
# one from, so on Docker Desktop for Windows `wg-quick up` on a full tunnel fails
# regardless of what is installed here. Native Linux and Docker Desktop for Mac
# have it. Shipping the backend is what makes the difference on those two;
# nothing shipped here can make the difference on WSL2, where the way out is to
# add the routes with `ip route` instead of going through `wg-quick` at all.
#
# `iptables` is deliberately still NOT here: with `nftables` present `wg-quick`
# never reaches for it, so it would add size and firewall surface for nothing.
# Remove default ubuntu user to free UID 1000 for host-user remapping
RUN if id ubuntu >/dev/null 2>&1; then userdel -r ubuntu 2>/dev/null || userdel ubuntu; fi \
&& if getent group ubuntu >/dev/null 2>&1; then groupdel ubuntu 2>/dev/null || true; fi