Explain a missing tun device where the failure actually happens
Build App (Preview) / compute-version (pull_request) Successful in 3s
Build App (Preview) / create-release (pull_request) Successful in 1s
Build App (Preview) / build-macos (pull_request) Successful in 2m37s
Build App (Preview) / build-linux (pull_request) Successful in 5m30s
Build App (Preview) / build-windows (pull_request) Successful in 5m55s
Build App (Preview) / prune-previews (pull_request) Successful in 3s

Review caught that the device guard was wired to the wrong call. The
daemon does not resolve `--device` at create: verified against Docker
29.7, `docker create --device /dev/does-not-exist` succeeds and prints an
id, and runc only resolves the device — and validates sysctls — when it
builds the container. So on a host with no tun module the create returns
fine and `start` fails, which means the explanation never ran and the
user saw the raw daemon string naming a path they would go looking for on
the wrong machine. The unit tests fed the create-side string straight in,
so they confirmed a function no real failure could reach.

Move the guard onto `start_container`, covering create as well in case a
future daemon checks earlier. It no longer takes `vpn_support_enabled` —
`start_container` has a container id and no project, and nothing else in
Triple-C ever requests a device, so an error naming /dev/net/tun is
unambiguous on its own. The test now uses the daemon's verbatim message
via bollard's real Display format.

Also from review:

  * Soften the security claim. Docker does not enable user-namespace
    remapping by default, so this is a real CAP_NET_ADMIN in the initial
    user namespace with only the network namespace confining it. It
    cannot touch host interfaces, but "confers no authority outside the
    container" was too strong: within its namespace it can set
    promiscuous mode and add addresses, routes and NAT on the shared
    docker0 segment, which puts sibling containers — the LiteLLM gateway
    among them — within ARP-spoofing reach, and it can flush netfilter
    rules sandbox mode may rely on. Said plainly in the code, CLAUDE.md
    and HOW-TO-USE.
  * Drop Tailscale from the list of clients needing this. Its
    --tun=userspace-networking mode needs neither the capability nor the
    device, and listing it invites granting NET_ADMIN for nothing.
  * Say in the toggle's own hint that changing it recreates the
    container, matching how every other recreation-triggering setting is
    labelled. The tab's generic "stop the container first" chip does not
    tell the user what is about to happen.
  * Add RuntimeSection tests: saves on, saves off explicitly rather than
    dropping the key, reflects state, is disabled while running, and
    carries the recreation warning.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-14 08:15:05 -07:00
co-authored by Claude Opus 5
parent 65a3d4eb29
commit 2b35aa8c16
6 changed files with 205 additions and 53 deletions
+15 -4
View File
@@ -287,10 +287,21 @@ container is created once by a very long function where a dropped capability is
Any two without the third still presents as a connection that hangs to a timeout, which is why
the tests assert the whole set.
- **The device is passed through from the host, never `mknod`-ed inside.** The kernel's `tun`
module has to back it. When the host has no such device the failure lands at *creation* — the
project simply won't start — so `explain_create_failure()` rewrites that one error to name the
switch and the Docker-Desktop-VM-vs-your-machine distinction. Do not let it degrade to a raw
bollard string.
module has to back it.
- **A missing device fails at `start`, not `create` — verified against Docker 29.7.** `docker
create --device /dev/does-not-exist` succeeds and prints an id; runc resolves the device (and
validates sysctls) only when it builds the container. So the guard belongs on the start path:
`explain_container_failure()` covers both and is called from `start_container`, where it has a
container id and no project — which is why it keys off the error naming `/dev/net/tun` rather
than off `vpn_support_enabled`. Nothing else in Triple-C requests a device, so that is
unambiguous. A version of this check wired to `create` alone is dead code that looks correct.
- **`NET_ADMIN` here is not user-namespaced.** Docker does not enable userns remapping by default,
so only the *network* namespace confines it: no reach onto host interfaces, but promiscuous
mode, arbitrary addresses/routes/NAT on the shared `docker0` segment (sibling containers, the
LiteLLM gateway among them, are ARP-spoofable), netlink-triggered host module auto-load, and
enough authority to flush in-container netfilter rules that sandbox mode may rely on. Keep the
code comments honest about this — an earlier draft claimed it "confers no authority" outside the
container, which is too strong.
- **`triple-c.vpn-support` is written unconditionally, including `false`.** The usual
`docker commit` reason: a `true` stamped once would ride the snapshot image into every future
container and make the switch impossible to turn off.
+14 -9
View File
@@ -477,19 +477,24 @@ When enabled, the container is given the three things a VPN client needs to buil
the `NET_ADMIN` capability, the `/dev/net/tun` device, and the `net.ipv4.conf.all.src_valid_mark`
sysctl that WireGuard requires. This is **off by default**.
Without it, a client such as PIA, WireGuard, OpenVPN or Tailscale installs and its daemon starts
normally, but the connection attempt **hangs until it times out** — a default container has no tun
device to open and no permission to add an interface or a route, and most clients report that as a
generic timeout rather than a permissions error.
Without it, a client such as PIA, WireGuard or OpenVPN installs and its daemon starts normally, but
the connection attempt **hangs until it times out** — a default container has no tun device to open
and no permission to add an interface or a route, and most clients report that as a generic timeout
rather than a permissions error.
Things worth knowing:
- `NET_ADMIN` applies to the container's **own** network namespace. It confers no authority over
the host's interfaces or over any other container. It does mean anything running in the
container can reconfigure that namespace, which is why it is opt-in.
- Tailscale is the exception: in its `--tun=userspace-networking` mode it needs neither the
capability nor the device, so leave this off if that is all you want.
- `NET_ADMIN` applies to the container's **own** network namespace — it cannot touch the host's
interfaces. It is not nothing, though: within that namespace anything in the container can set
promiscuous mode and add arbitrary addresses, routes and firewall rules on the Docker bridge it
shares with your other containers, and it can flush firewall rules that sandbox mode relies on.
Grant it per project, to projects that need it.
- The **Docker host's** kernel must have the `tun` module available. With Docker Desktop that is
the Linux VM, not your own machine. If it is missing, the container fails to create with an
error naming `/dev/net/tun` and pointing back at this setting.
the Linux VM, not your own machine. If it is missing, the container is created but fails to
**start**, with an error naming `/dev/net/tun` and pointing back at this setting.
- A VPN client's kill switch applies to everything in the container, Claude Code included. If the
tunnel drops, expect API calls to fail until it reconnects or the kill switch is turned off.
+75 -34
View File
@@ -830,8 +830,16 @@ type VpnHostConfigParts = (
/// its handshake packets are dropped by reverse-path filtering. Harmless for
/// OpenVPN-based clients, so it is set unconditionally with the rest.
///
/// This is namespaced to the container's own network stack: `NET_ADMIN` confers
/// no authority over the host's interfaces or over any other container.
/// What it costs, stated accurately: Docker does not enable user-namespace
/// remapping by default, so this is a real `CAP_NET_ADMIN` in the *initial*
/// user namespace and only the **network** namespace confines it. It cannot
/// touch the host's interfaces, but within its own namespace it can set
/// promiscuous mode and add arbitrary addresses, routes and NAT rules on the
/// shared `docker0` L2 segment — which puts sibling containers (the LiteLLM
/// gateway among them) within reach of ARP spoofing, and lets netlink trigger
/// host-kernel module auto-loading. It is also enough to flush netfilter rules
/// inside the container, so pair it with `sandbox_mode_enabled` advisedly.
/// Hence opt-in, per project, rather than on for everyone.
fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
if !enabled {
return (None, None, None);
@@ -857,31 +865,42 @@ fn vpn_host_config(enabled: bool) -> VpnHostConfigParts {
/// Turn the daemon's device-passthrough failure into an explanation.
///
/// Requesting `/dev/net/tun` fails at *creation* when the host kernel has no
/// `tun` module — and the raw bollard error names a path the user will look for
/// on the wrong machine, since with Docker Desktop the relevant host is the
/// Linux VM rather than their own. Left unmapped this surfaces as a project
/// that simply refuses to start, with nothing pointing back at the switch that
/// caused it.
fn explain_create_failure(err: &str, vpn_enabled: bool) -> String {
let device_missing = vpn_enabled
&& err.contains(TUN_DEVICE)
/// **This fires on `start`, not `create`.** Verified against Docker 29.7:
/// `docker create --device /dev/does-not-exist` succeeds and prints an id; the
/// device is only resolved when runc builds the container, so the failure lands
/// on the *next* call. Sysctls validate at the same point. Anything that
/// inspects only the create path will never see it — which is why both paths
/// route through here and the tests exercise the start-side string.
///
/// Unmapped, this reads as `Failed to start container: Docker responded with
/// status code 500: error gathering device information while adding custom
/// device "/dev/net/tun": no such file or directory` — a path the user will go
/// looking for on the wrong machine, since with Docker Desktop the relevant
/// host is the Linux VM rather than their own, and with nothing pointing back
/// at the switch that caused it.
///
/// Deliberately not gated on `vpn_support_enabled`: nothing else in Triple-C
/// ever asks for a device, so an error naming `/dev/net/tun` can only have come
/// from a container created with the switch on. That keeps the check usable
/// from [`start_container`], which has a container id and no project.
fn explain_container_failure(action: &str, err: &str) -> String {
let device_missing = err.contains(TUN_DEVICE)
&& (err.contains("no such file or directory")
|| err.contains("No such file or directory")
|| err.contains("error gathering device information"));
if device_missing {
return format!(
"Failed to create container: the Docker host has no {} device, which \
"Failed to {} container: the Docker host has no {} device, which \
\"VPN support\" requires. The host kernel needs the `tun` module \
loaded (on Docker Desktop that is the Linux VM, not your own \
machine). Turn VPN support off in Config → Runtime to start this \
project without it. Original error: {}",
TUN_DEVICE, err
action, TUN_DEVICE, err
);
}
format!("Failed to create container: {}", err)
format!("Failed to {} container: {}", action, err)
}
pub async fn create_container(
@@ -1574,7 +1593,7 @@ pub async fn create_container(
let response = docker
.create_container(Some(options), config)
.await
.map_err(|e| explain_create_failure(&e.to_string(), project.vpn_support_enabled))?;
.map_err(|e| explain_container_failure("create", &e.to_string()))?;
Ok(response.id)
}
@@ -1584,7 +1603,7 @@ pub async fn start_container(container_id: &str) -> Result<(), String> {
docker
.start_container(container_id, None::<StartContainerOptions<String>>)
.await
.map_err(|e| format!("Failed to start container: {}", e))
.map_err(|e| explain_container_failure("start", &e.to_string()))
}
pub async fn stop_container(container_id: &str) -> Result<(), String> {
@@ -2770,31 +2789,53 @@ mod tests {
assert_eq!(cap_add.unwrap(), vec!["NET_ADMIN"]);
}
/// What bollard actually hands us when a tun-less host rejects the device.
///
/// Captured verbatim from Docker 29.7: `docker create` with a missing
/// device **succeeds**, and this arrives from the subsequent `start`.
/// `DockerResponseServerError`'s Display is
/// `"Docker responded with status code {code}: {message}"` with the
/// daemon's message unaltered.
const REAL_TUN_ERROR: &str = "Docker responded with status code 500: error \
gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
#[test]
fn a_missing_tun_device_is_explained_rather_than_echoed() {
let raw = "error gathering device information while adding custom device \
\"/dev/net/tun\": no such file or directory";
let msg = explain_create_failure(raw, true);
fn a_missing_tun_device_is_explained_on_the_path_that_actually_fails() {
// The start path is the one that matters: the daemon defers device
// resolution to runc, so create returns an id on a host with no tun
// module and only start fails. A version of this that checked create
// alone would be dead code.
let msg = explain_container_failure("start", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to start container:"), "{}", msg);
assert!(msg.contains("VPN support"), "should name the switch: {}", msg);
assert!(msg.contains("tun` module"), "should name the cause: {}", msg);
assert!(msg.contains(raw), "should keep the original error: {}", msg);
assert!(msg.contains("Config → Runtime"), "should say where to fix it: {}", msg);
assert!(msg.contains(REAL_TUN_ERROR), "should keep the original: {}", msg);
}
#[test]
fn the_same_explanation_covers_create_if_the_daemon_ever_checks_earlier() {
// Belt and braces — older and future daemons may validate at create.
let msg = explain_container_failure("create", REAL_TUN_ERROR);
assert!(msg.starts_with("Failed to create container:"), "{}", msg);
assert!(msg.contains("VPN support"), "{}", msg);
}
#[test]
fn unrelated_failures_are_left_alone() {
// Including a tun error on a project that never asked for VPN support —
// that came from somewhere else and must not be misattributed.
let name_clash = "Conflict. The container name \"/triple-c-x\" is already in use";
assert_eq!(
explain_create_failure(name_clash, true),
format!("Failed to create container: {}", name_clash)
);
let tun_err = "no such file or directory: /dev/net/tun";
assert_eq!(
explain_create_failure(tun_err, false),
format!("Failed to create container: {}", tun_err)
);
for (action, err) in [
("create", "Conflict. The container name \"/triple-c-x\" is already in use"),
("start", "Docker responded with status code 404: No such container"),
("start", "error gathering device information while adding custom device \"/dev/dri/card0\""),
] {
assert_eq!(
explain_container_failure(action, err),
format!("Failed to {} container: {}", action, err),
"{} should pass through untouched",
err
);
}
}
#[test]
+6 -5
View File
@@ -148,13 +148,14 @@ pub struct Project {
/// Grant the container what a VPN client needs to build a tunnel:
/// `CAP_NET_ADMIN`, the `/dev/net/tun` device, and the WireGuard
/// `src_valid_mark` sysctl. Without all three a client (PIA, WireGuard,
/// OpenVPN, Tailscale) installs and runs but its connection attempt hangs
/// until it times out, because it cannot create the tunnel interface or
/// touch the routing table.
/// OpenVPN) installs and runs but its connection attempt hangs until it
/// times out, because it cannot create the tunnel interface or touch the
/// routing table.
///
/// Off by default and deliberately opt-in: `NET_ADMIN` lets anything in the
/// container reconfigure its own network stack, which is a meaningful step
/// out of the default sandbox. Unlike `auth_bridge_enabled` this *is*
/// container reconfigure its own network stack, which reaches further than
/// it sounds — see `vpn_host_config` for what it does and does not confer.
/// Unlike `auth_bridge_enabled` this *is*
/// container state, so it carries a `triple-c.vpn-support` label and is
/// compared in `container_needs_recreation` — capabilities and devices are
/// fixed at creation and can only change by recreating the container.
@@ -0,0 +1,94 @@
import { describe, it, expect, vi, beforeEach } from "vitest";
import { render, screen, fireEvent } from "@testing-library/react";
import RuntimeSection from "./RuntimeSection";
import type { Project } from "../../../../lib/types";
const baseProject: Project = {
id: "p1",
name: "api-server",
paths: [{ host_path: "/src/api", mount_name: "api" }],
container_id: null,
status: "stopped",
backend: "anthropic",
bedrock_config: null,
ollama_config: null,
llamacpp_config: null,
openai_compatible_config: null,
allow_docker_access: false,
sandbox_mode_enabled: true,
mission_control_enabled: false,
auth_bridge_enabled: false,
browser_view_enabled: false,
vpn_support_enabled: false,
use_shared_auth_token: true,
full_permissions: false,
permission_mode: null,
ssh_key_path: null,
ca_cert_path: null,
git_token: null,
git_user_name: null,
git_user_email: null,
custom_env_vars: [],
port_mappings: [],
claude_instructions: null,
claude_code_settings: null,
renamed_session_names: {},
created_at: "2026-01-01T00:00:00Z",
updated_at: "2026-01-01T00:00:00Z",
};
const VPN = "VPN support";
const save = vi.fn().mockResolvedValue(true);
function renderSection(over: Partial<Project> = {}, disabled = false) {
return render(
<RuntimeSection
project={{ ...baseProject, ...over }}
save={save}
disabled={disabled}
disabledReason="Container must be stopped to change this setting."
/>,
);
}
describe("RuntimeSection — VPN support toggle", () => {
beforeEach(() => vi.clearAllMocks());
it("saves only the VPN flag when switched on", () => {
renderSection();
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: true });
});
it("saves the flag off again, rather than dropping the key", () => {
// Off has to be written explicitly: the container carries a
// `triple-c.vpn-support` label either way, and an absent value would leave
// the capability granted.
renderSection({ vpn_support_enabled: true });
fireEvent.click(screen.getByRole("switch", { name: VPN }));
expect(save).toHaveBeenCalledWith({ vpn_support_enabled: false });
});
it("reflects the project's current state", () => {
renderSection({ vpn_support_enabled: true });
expect(screen.getByRole("switch", { name: VPN })).toBeChecked();
});
it("cannot be changed while the container is running", () => {
// Capabilities and devices are fixed at creation, so this setting is gated
// on the container being stopped along with the rest of the tab.
renderSection({}, true);
const toggle = screen.getByRole("switch", { name: VPN });
expect(toggle).toBeDisabled();
fireEvent.click(toggle);
expect(save).not.toHaveBeenCalled();
});
it("warns that the change recreates the container", () => {
renderSection();
expect(
screen.getByText(/recreates the container on its next start/i),
).toBeInTheDocument();
});
});
@@ -59,7 +59,7 @@ export default function RuntimeSection({
<SwitchRow
label="VPN support"
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN, Tailscale) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched."
hint="Grants NET_ADMIN and the /dev/net/tun device so a VPN client (PIA, WireGuard, OpenVPN) can build a tunnel inside the container. Without it a client installs and runs but its connection hangs until it times out. Anything in the container can then reconfigure the container's own network stack; the host's is untouched. Changing this recreates the container on its next start — the home and .claude volumes are preserved."
control={
<Toggle
label="VPN support"