Compare commits

..
1 Commits
Author SHA1 Message Date
shadow-testandClaude Opus 5 d260f2c7c3 Let a project's container run a VPN client
Build App (Preview) / compute-version (pull_request) Successful in 5s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m37s
Build App (Preview) / build-linux (pull_request) Successful in 7m45s
Build App (Preview) / build-windows (pull_request) Failing after 13m30s
Build App (Preview) / prune-previews (pull_request) Skipped
A VPN client installed in a container today starts, runs, and then hangs
until its connection times out. Nothing reports an error: a default
container has no /dev/net/tun to open and no CAP_NET_ADMIN to add an
interface or a route with, and clients surface that as a generic timeout
rather than a permissions failure.

Add an opt-in per-project "VPN support" switch granting the three things
a tunnel needs. They are useless individually, which is why
vpn_host_config() defines the set in one place and the tests assert all
of it:

  * CAP_NET_ADMIN — Docker's default bounding set has net_raw but not
    net_admin, so a client can ping but never connect.
  * /dev/net/tun — passed through from the host so the kernel's tun
    module backs it, rather than mknod-ed inside.
  * net.ipv4.conf.all.src_valid_mark — WireGuard's wg-quick sets this and
    cannot from inside a container, /proc/sys being read-only, so its
    handshakes are dropped by reverse-path filtering.

Off by default and deliberately opt-in: NET_ADMIN lets anything in the
container reconfigure that container's network stack. It is namespaced —
no authority over the host's interfaces or any other container.

Capabilities and devices are fixed when a container is created, so this
is container state and takes the label-and-compare treatment.
triple-c.vpn-support is written unconditionally, false included, for the
usual docker commit reason: a true stamped once would ride the snapshot
image into every future container and make the switch impossible to turn
back off. A missing label reads as false and off is byte-identical to
today, so no existing project is churned.

Requesting the device fails at creation when the host kernel has no tun
module, which would otherwise surface as a project that simply refuses to
start. explain_create_failure() rewrites that one error to name the
switch and the Docker-Desktop-VM-versus-your-machine distinction, and
leaves every other failure untouched.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 07:52:39 -07:00
7 changed files with 67 additions and 303 deletions
+14 -98
View File
@@ -39,48 +39,13 @@ jobs:
MAJOR_MINOR=$(cat VERSION | tr -d '[:space:]')
echo "Major.Minor: ${MAJOR_MINOR}"
# The patch number is **one past the highest patch already used**, and
# never a distance.
#
# It used to be `git rev-list --count <highest tag>..HEAD`, which is
# not a counter at all: it measures how far HEAD has drifted from
# whichever tag sorts highest, and that resets to zero every time a
# tag is cut. The published history is the proof — each of these is
# exactly what the old formula returned at the time:
#
# v0.4.0 -> 3 commits -> v0.4.3 looked fine
# v0.4.3 -> 4 commits -> v0.4.4 fine by luck, 4 > 3
# v0.4.4 -> 2 commits -> v0.4.2 went backwards
# v0.4.4 -> 6 commits -> v0.4.6 jumped, skipping .5
# v0.4.6 -> 3 commits -> v0.4.3 already taken; the upload failed
#
# Reusing a version is worse than failing to publish one: the macOS
# and Windows steps replace assets in place, so a duplicate silently
# rewrote a release that had been public for three days. Monotonic
# numbering is what stops that at the source.
#
# Suffixed tags count too. `create-tag` is skipped when any platform
# job fails, so a run can publish v0.4.7-mac and never create the
# plain v0.4.7 — reading only unsuffixed tags would then hand the
# same number out twice.
HIGHEST=$(git tag -l "v${MAJOR_MINOR}.*" \
| grep -E "^v${MAJOR_MINOR}\.[0-9]+(-mac|-win)?$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+).*/\1/" \
| sort -n | tail -1 || true)
# Find the latest tag matching v{MAJOR_MINOR}.N (exclude -mac, -win suffixes)
# `|| true` so an empty grep result doesn't fail the step under pipefail.
LATEST_TAG=$(git tag -l "v${MAJOR_MINOR}.*" --sort=-v:refname | grep -E "^v${MAJOR_MINOR}\.[0-9]+$" | head -1 || true)
# A re-run of a commit that already released must not mint a new
# version just because its own tag now exists.
EXISTING=$(git tag --points-at HEAD \
| grep -E "^v${MAJOR_MINOR}\.[0-9]+$" \
| sed -E "s/^v${MAJOR_MINOR}\.([0-9]+)$/\1/" \
| sort -n | tail -1 || true)
if [ -n "$EXISTING" ]; then
echo "HEAD is already tagged v${MAJOR_MINOR}.${EXISTING} — reusing it"
PATCH="${EXISTING}"
elif [ -n "$HIGHEST" ]; then
echo "Highest patch already used on this line: ${HIGHEST}"
PATCH=$((HIGHEST + 1))
if [ -n "$LATEST_TAG" ]; then
echo "Latest matching tag: ${LATEST_TAG}"
PATCH=$(git rev-list --count "${LATEST_TAG}..HEAD")
else
# A minor line nobody has tagged yet is a *new* line, and a new line
# starts at .0 — that is what "we are moving to 0.4.x" means. The
@@ -200,70 +165,21 @@ jobs:
env:
TOKEN: ${{ secrets.REGISTRY_TOKEN }}
run: |
set -euo pipefail
TAG="v${{ needs.compute-version.outputs.version }}"
# Idempotent get-or-create, matching build-macos. This step used to
# POST /releases unconditionally: against a tag that already existed
# Gitea answered 409, the grep below found no id, and the run died
# with a bare "exitcode '1'" and not one line of output explaining
# it — `curl -s` with no `-f` swallows the HTTP error, so nothing
# ever said "409" or "duplicate tag". Hence -fsS throughout, and
# pipefail so a failure cannot be stepped over.
HTTP_CODE=$(curl -sS -o release.json -w '%{http_code}' \
# Create release
curl -s -X POST \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/tags/${TAG}")
case "${HTTP_CODE}" in
200)
echo "Release ${TAG} already exists, reusing"
;;
404)
echo "Creating release ${TAG}"
curl -fsS -X POST \
-H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/json" \
-d "{\"tag_name\": \"${TAG}\", \"name\": \"Triple-C ${TAG} (Linux)\", \"body\": \"Automated build from commit ${{ gitea.sha }}\"}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases" > release.json
;;
*)
echo "Unexpected ${HTTP_CODE} looking up release ${TAG}:" >&2
cat release.json >&2
exit 1
;;
esac
RELEASE_ID=$(python3 -c "import json,sys; print(json.load(open('release.json')).get('id',''))")
if [ -z "${RELEASE_ID}" ]; then
echo "No release id for ${TAG}; refusing to upload into nothing:" >&2
cat release.json >&2
exit 1
fi
-H "Content-Type: application/json" \
-d "{\"tag_name\": \"${TAG}\", \"name\": \"Triple-C ${TAG} (Linux)\", \"body\": \"Automated build from commit ${{ gitea.sha }}\"}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases" > release.json
RELEASE_ID=$(cat release.json | grep -o '"id":[0-9]*' | head -1 | grep -o '[0-9]*')
echo "Release ID: ${RELEASE_ID}"
# Replace-not-conflict, so a retry after a partial upload succeeds.
# Versions are monotonic now (see compute-version), so this can only
# ever be replacing an asset from a failed run of this same commit —
# never one belonging to an already-published version.
# Upload each artifact
for file in artifacts/*; do
[ -f "$file" ] || continue
filename=$(basename "$file")
EXISTING_ID=$(curl -sS \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets" \
| python3 -c "import json,sys; t=sys.argv[1]; print(next((a['id'] for a in json.load(sys.stdin) if a.get('name')==t), ''))" "${filename}" || true)
if [ -n "${EXISTING_ID}" ]; then
echo "Deleting existing asset ${filename} (id ${EXISTING_ID})"
curl -fsS -X DELETE \
-H "Authorization: token ${TOKEN}" \
"${GITEA_URL}/api/v1/repos/${REPO}/releases/${RELEASE_ID}/assets/${EXISTING_ID}"
fi
echo "Uploading ${filename}..."
curl -fsS --http1.1 \
--retry 5 --retry-all-errors --retry-delay 5 \
--max-time 600 \
-X POST \
curl -s -X POST \
-H "Authorization: token ${TOKEN}" \
-H "Content-Type: application/octet-stream" \
--data-binary "@${file}" \
+4 -15
View File
@@ -287,21 +287,10 @@ container is created once by a very long function where a dropped capability is
Any two without the third still presents as a connection that hangs to a timeout, which is why
the tests assert the whole set.
- **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun`
module has to back it.
- **A missing device fails at `start`, not `create` — verified against Docker 29.7.** `docker
create --device /dev/does-not-exist` succeeds and prints an id; runc resolves the device (and
validates sysctls) only when it builds the container. So the guard belongs on the start path:
`explain_container_failure()` covers both and is called from `start_container`, where it has a
container id and no project — which is why it keys off the error naming `/dev/net/tun` rather
than off `vpn_support_enabled`. Nothing else in Triple-C requests a device, so that is
unambiguous. A version of this check wired to `create` alone is dead code that looks correct.
- **`NET_ADMIN` here is not user-namespaced.** Docker does not enable userns remapping by default,
so only the *network* namespace confines it: no reach onto host interfaces, but promiscuous
mode, arbitrary addresses/routes/NAT on the shared `docker0` segment (sibling containers, the
LiteLLM gateway among them, are ARP-spoofable), netlink-triggered host module auto-load, and
enough authority to flush in-container netfilter rules that sandbox mode may rely on. Keep the
code comments honest about this — an earlier draft claimed it "confers no authority" outside the
container, which is too strong.
module has to back it. When the host has no such device the failure lands at *creation* — the
project simply won't start — so `explain_create_failure()` rewrites that one error to name the
switch and the Docker-Desktop-VM-vs-your-machine distinction. Do not let it degrade to a raw
bollard string.
- **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual
`docker commit` reason: a `true` stamped once would ride the snapshot image into every future
container and make the switch impossible to turn off.
+9 -14
View File
@@ -477,24 +477,19 @@ When enabled, the container is given the three things a VPN client needs to buil
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
sysctl that WireGuard requires. This is **off by default**.
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
the connection attempt **hangs until it times out** — a default container has no tun device to open
and no permission to add an interface or a route, and most clients report that as a generic timeout
rather than a permissions error.
Without it, a client such as PIA, WireGuard, OpenVPN or Tailscale installs and its daemon starts
normally, but the connection attempt **hangs until it times out** — a default container has no tun
device to open and no permission to add an interface or a route, and most clients report that as a
generic timeout rather than a permissions error.
Things worth knowing:
- Tailscale is the exception: in its `--tun=userspace-networking` mode it needs neither the
capability nor the device, so leave this off if that is all you want.
- `NET_ADMIN` applies to the container's **own** network namespace — it cannot touch the host's
interfaces. It is not nothing, though: within that namespace anything in the container can set
promiscuous mode and add arbitrary addresses, routes and firewall rules on the Docker bridge it
shares with your other containers, and it can flush firewall rules that sandbox mode relies on.
Grant it per project, to projects that need it.
- `NET_ADMIN` applies to the container's **own** network namespace. It confers no authority over
the host's interfaces or over any other container. It does mean anything running in the
container can reconfigure that namespace, which is why it is opt-in.
- The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is
the Linux VM, not your own machine. If it is missing, the container is created but fails to
**start**, with an error naming `/dev/net/tun` and pointing back at this setting.
the Linux VM, not your own machine. If it is missing, the container fails to create with an
error naming `/dev/net/tun` and pointing back at this setting.
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
+34 -75
View File
@@ -830,16 +830,8 @@ type VpnHostConfigParts = (
/// its handshake packets are dropped by reverse-path filtering. Harmless for
/// OpenVPN-based clients, so it is set unconditionally with the rest.
///
/// What it costs, stated accurately: Docker does not enable user-namespace
/// remapping by default, so this is a real `CAP_NET_ADMIN` in the *initial*
/// user namespace and only the **network** namespace confines it. It cannot
/// touch the host's interfaces, but within its own namespace it can set
/// promiscuous mode and add arbitrary addresses, routes and NAT rules on the
/// shared `docker0` L2 segment — which puts sibling containers (the LiteLLM
/// gateway among them) within reach of ARP spoofing, and lets netlink trigger
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
/// Hence opt-in, per project, rather than on for everyone.
/// This is namespaced to the container's own network stack: `NET_ADMIN` confers
/// no authority over the host's interfaces or over any other container.
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
if !enabled {
return (None, None, None);
@@ -865,42 +857,31 @@ fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
/// Turn the daemon's device-passthrough failure into an explanation.
///
/// **This fires on `start`, not `create`.** Verified against Docker 29.7:
/// `docker create --device /dev/does-not-exist` succeeds and prints an id; the
/// device is only resolved when runc builds the container, so the failure lands
/// on the *next* call. Sysctls validate at the same point. Anything that
/// inspects only the create path will never see it — which is why both paths
/// route through here and the tests exercise the start-side string.
///
/// Unmapped, this reads as `Failed to start container: Docker responded with
/// status code 500: error gathering device information while adding custom
/// device "/dev/net/tun": no such file or directory` — a path the user will go
/// looking for on the wrong machine, since with Docker Desktop the relevant
/// host is the Linux VM rather than their own, and with nothing pointing back
/// at the switch that caused it.
///
/// Deliberately not gated on `vpn_support_enabled`: nothing else in Triple-C
/// ever asks for a device, so an error naming `/dev/net/tun` can only have come
/// from a container created with the switch on. That keeps the check usable
/// from [`start_container`], which has a container id and no project.
fn explain_container_failure(action: &str, err: &str) -> String {
let device_missing = err.contains(TUN_DEVICE)
/// Requesting `/dev/net/tun` fails at *creation* when the host kernel has no
/// `tun` module — and the raw bollard error names a path the user will look for
/// on the wrong machine, since with Docker Desktop the relevant host is the
/// Linux VM rather than their own. Left unmapped this surfaces as a project
/// that simply refuses to start, with nothing pointing back at the switch that
/// caused it.
fn explain_create_failure(err: &str, vpn_enabled: bool) -> String {
let device_missing = vpn_enabled
&& err.contains(TUN_DEVICE)
&& (err.contains("no such file or directory")
|| err.contains("No such file or directory")
|| err.contains("error gathering device information"));
if device_missing {
return format!(
"Failed to {} container: the Docker host has no {} device, which \
"Failed to create container: the Docker host has no {} device, which \
\"VPN support\" requires. The host kernel needs the `tun` module \
loaded (on Docker Desktop that is the Linux VM, not your own \
machine). Turn VPN support off in Config Runtime to start this \
project without it. Original error: {}",
action, TUN_DEVICE, err
TUN_DEVICE, err
);
}
format!("Failed to {} container: {}", action, err)
format!("Failed to create container: {}", err)
}
pub async fn create_container(
@@ -1593,7 +1574,7 @@ pub async fn create_container(
let response = docker
.create_container(Some(options), config)
.await
.map_err(|e| explain_container_failure("create", &e.to_string()))?;
.map_err(|e| explain_create_failure(&e.to_string(), project.vpn_support_enabled))?;
Ok(response.id)
}
@@ -1603,7 +1584,7 @@ pub async fn start_container(container_id: &str) -> Result<(), String> {
docker
.start_container(container_id, None::<StartContainerOptions<String>>)
.await
.map_err(|e| explain_container_failure("start", &e.to_string()))
.map_err(|e| format!("Failed to start container: {}", e))
}
pub async fn stop_container(container_id: &str) -> Result<(), String> {
@@ -2789,53 +2770,31 @@ mod tests {
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
}
/// What bollard actually hands us when a tun-less host rejects the device.
///
/// Captured verbatim from Docker 29.7: `docker create` with a missing
/// device **succeeds**, and this arrives from the subsequent `start`.
/// `DockerResponseServerError`'s Display is
/// `"Docker responded with status code {code}: {message}"` with the
/// daemon's message unaltered.
const REAL_TUN_ERROR: &str = "Docker responded with status code 500: error \
gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
#[test]
fn a_missing_tun_device_is_explained_on_the_path_that_actually_fails() {
// The start path is the one that matters: the daemon defers device
// resolution to runc, so create returns an id on a host with no tun
// module and only start fails. A version of this that checked create
// alone would be dead code.
let msg = explain_container_failure("start", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to start container:"), "{}", msg);
fn a_missing_tun_device_is_explained_rather_than_echoed() {
let raw = "error gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
let msg = explain_create_failure(raw, true);
assert!(msg.contains("VPN support"), "should name the switch: {}", msg);
assert!(msg.contains("tun` module"), "should name the cause: {}", msg);
assert!(msg.contains("Config → Runtime"), "should say where to fix it: {}", msg);
assert!(msg.contains(REAL_TUN_ERROR), "should keep the original: {}", msg);
}
#[test]
fn the_same_explanation_covers_create_if_the_daemon_ever_checks_earlier() {
// Belt and braces — older and future daemons may validate at create.
let msg = explain_container_failure("create", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to create container:"), "{}", msg);
assert!(msg.contains("VPN support"), "{}", msg);
assert!(msg.contains(raw), "should keep the original error: {}", msg);
}
#[test]
fn unrelated_failures_are_left_alone() {
for (action, err) in [
("create", "Conflict. The container name \"/triple-c-x\" is already in use"),
("start", "Docker responded with status code 404: No such container"),
("start", "error gathering device information while adding custom device \"/dev/dri/card0\""),
] {
assert_eq!(
explain_container_failure(action, err),
format!("Failed to {} container: {}", action, err),
"{} should pass through untouched",
err
);
}
// Including a tun error on a project that never asked for VPN support —
// that came from somewhere else and must not be misattributed.
let name_clash = "Conflict. The container name \"/triple-c-x\" is already in use";
assert_eq!(
explain_create_failure(name_clash, true),
format!("Failed to create container: {}", name_clash)
);
let tun_err = "no such file or directory: /dev/net/tun";
assert_eq!(
explain_create_failure(tun_err, false),
format!("Failed to create container: {}", tun_err)
);
}
#[test]
+5 -6
View File
@@ -148,14 +148,13 @@ pub struct Project {
/// Grant the container what a VPN client needs to build a tunnel:
/// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard
/// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard,
/// OpenVPN) installs and runs but its connection attempt hangs until it
/// times out, because it cannot create the tunnel interface or touch the
/// routing table.
/// OpenVPN, Tailscale) installs and runs but its connection attempt hangs
/// until it times out, because it cannot create the tunnel interface or
/// touch the routing table.
///
/// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the
/// container reconfigure its own network stack, which reaches further than
/// it sounds — see `vpn_host_config` for what it does and does not confer.
/// Unlike `auth_bridge_enabled` this *is*
/// container reconfigure its own network stack, which is a meaningful step
/// out of the default sandbox. Unlike `auth_bridge_enabled` this *is*
/// container state, so it carries a `triple-c.vpn-support` label and is
/// compared in `container_needs_recreation` — capabilities and devices are
/// fixed at creation and can only change by recreating the container.
@@ -1,94 +0,0 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
import { render, screen, fireEvent } from "@testing-library/react";
import RuntimeSection from "./RuntimeSection";
import type { Project } from "../../../../lib/types";
const baseProject: Project = {
id: "p1",
name: "api-server",
paths: [{ host_path: "/src/api", mount_name: "api" }],
container_id: null,
status: "stopped",
backend: "anthropic",
bedrock_config: null,
ollama_config: null,
llamacpp_config: null,
openai_compatible_config: null,
allow_docker_access: false,
sandbox_mode_enabled: true,
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: true,
full_permissions: false,
permission_mode: null,
ssh_key_path: null,
ca_cert_path: null,
git_token: null,
git_user_name: null,
git_user_email: null,
custom_env_vars: [],
port_mappings: [],
claude_instructions: null,
claude_code_settings: null,
renamed_session_names: {},
created_at: "2026-01-01T00:00:00Z",
updated_at: "2026-01-01T00:00:00Z",
};
const VPN = "VPN support";
const save = vi.fn().mockResolvedValue(true);
function renderSection(over: Partial<Project> = {}, disabled = false) {
return render(
<RuntimeSection
project={{ ...baseProject, ...over }}
save={save}
disabled={disabled}
disabledReason="Container must be stopped to change this setting."
/>,
);
}
describe("RuntimeSection — VPN support toggle", () => {
beforeEach(() => vi.clearAllMocks());
it("saves only the VPN flag when switched on", () => {
renderSection();
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: true });
});
it("saves the flag off again, rather than dropping the key", () => {
// Off has to be written explicitly: the container carries a
// `triple-c.vpn-support` label either way, and an absent value would leave
// the capability granted.
renderSection({ vpn_support_enabled: true });
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: false });
});
it("reflects the project's current state", () => {
renderSection({ vpn_support_enabled: true });
expect(screen.getByRole("switch", { name: VPN })).toBeChecked();
});
it("cannot be changed while the container is running", () => {
// Capabilities and devices are fixed at creation, so this setting is gated
// on the container being stopped along with the rest of the tab.
renderSection({}, true);
const toggle = screen.getByRole("switch", { name: VPN });
expect(toggle).toBeDisabled();
fireEvent.click(toggle);
expect(save).not.toHaveBeenCalled();
});
it("warns that the change recreates the container", () => {
renderSection();
expect(
screen.getByText(/recreates the container on its next start/i),
).toBeInTheDocument();
});
});
@@ -59,7 +59,7 @@ export default function RuntimeSection({
<SwitchRow
label="VPN support"
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched. Changing this recreates the container on its next start — the home and .claude volumes are preserved."
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN, Tailscale) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched."
control={
<Toggle
label="VPN support"