Fix review findings: secrets in snapshots, URL spoofing, migration data loss

Adversarial review of the branch produced findings across four areas.
This addresses them, plus the Windows CI environment.

Secrets. commit_container_snapshot baked the container's full env into
the per-project snapshot image, so the shared OAuth token — and the AWS
keys, git token and gateway master key — outlived revocation and were
readable via docker inspect. Verified against Engine 29.6 that a commit
body's config merges over the container's: keys cannot be dropped but
can be overwritten, so all of them now commit as KEY=. clear_claude_token
additionally rewrites images from earlier builds and reports honestly
when a tag could not be rewritten.

The recommendation to move the token out of env entirely was not taken,
with reasoning: apiKeyHelper is a different auth method that outranks
CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no
file-based delivery exists. The durable exposure — the image — is what
is closed here. Separately noted, not fixed: entrypoint.sh captures the
token into the scheduler's .env inside the persisted volume.

URL spoofing. Three call sites reached openUrl with container-controlled
strings, one of which the review missed (the WebLinksAddon handler).
The sign-in URL was scraped from container output with a longest-match
tie-break and no userinfo check, so claude.ai@evil.tld rendered as
"claude.ai…" in a truncating element. There is now one sanitizer in
front of every sink — scheme allowlist, no userinfo, C0/C1 and quote
rejection, host allowlist for the sign-in case, first-match — and the
origin renders un-truncated. The toast is keyed so a changed URL
remounts, closing a bait-and-switch where the user read one URL and
clicked another.

Migration. The rollback pin was best-effort: a tag failure was logged
and the migration continued past remove_container, after which the
final commit overwrote the only copy of the old system layer. It now
aborts before anything destructive and reads the tag back. /var was
destroyed while the ordinary recreate path preserves it — making the
"safe" alternative to Reset more destructive than Reset's alternative;
data-bearing subtrees are now detected and disclosed in the pre-flight
rather than copied, since tarring a live database onto a different
base's packages is a corruption risk. resume_migration now verifies the
migration-state label instead of reporting success for a container that
never swapped. dismiss actually resolves the record rather than leaving
the feature permanently refusing to migrate. Start and Reset are guarded
while a migration is live.

Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and
advertised URL are derived together so they cannot drift. Disabling it
now stops it. App exit runs teardown concurrently under a budget with a
visible shutting-down state instead of blocking for minutes. Auto-starts
retry when Docker is not up yet, and the polling-recovery path now
reconciles, so interrupted migrations are still recovered. Auth-bridge
forwards are capped, closing a container-driven fd exhaustion.

Windows CI. build-windows failed on this branch with "linker link.exe
not found". The runner had no MSVC build tools and the workflow assumed
a hand-provisioned machine, so a bare runner registers, accepts jobs and
fails at link time after downloading the whole crate graph. The job now
installs the VC++ workload when vswhere cannot find it, matching how it
already conditionally installs Rust and Node.

192 Rust tests, 274 frontend tests, both builds clean, zero warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-09 19:35:39 -07:00
co-authored by Claude Opus 5
parent eb1324cb16
commit 2de00b3c55
43 changed files with 4348 additions and 346 deletions
+341 -36
View File
@@ -7,11 +7,14 @@
//!
//! Two things differ from STT, both deliberate:
//!
//! * **The port is published on `0.0.0.0`, not `127.0.0.1`.** STT is consumed
//! by the Tauri host process, so loopback is enough. The gateway is consumed
//! by *project containers*, which sit on Docker's default bridge and reach
//! the host through the bridge gateway — a loopback-only bind is invisible to
//! them. See [`gateway_base_url`].
//! * **The published host address is *detected*, not fixed.** STT is consumed
//! by the Tauri host process, so loopback is always enough. The gateway is
//! consumed by *project containers*, and how a container reaches the host
//! depends on the engine — so the bind address does too. See
//! [`GatewayBinding`]. It is never `0.0.0.0`: the config behind this port
//! holds a billed provider key, and Docker's published-port rules land in the
//! `DOCKER` iptables chain *ahead* of a host firewall, so a wildcard bind is
//! genuinely LAN-reachable even with `ufw` enabled.
//! * **The rendered config is uploaded into the container over the Docker
//! API** rather than passed as env. It holds the provider API key, and both
//! env vars and labels are readable by anything on the host via
@@ -23,10 +26,14 @@ use bollard::container::{
};
use bollard::image::BuildImageOptions;
use bollard::models::{HostConfig, Mount, MountTypeEnum, PortBinding};
use bollard::network::InspectNetworkOptions;
use bollard::Docker;
use futures_util::StreamExt;
use sha2::{Digest, Sha256};
use std::collections::HashMap;
use std::io::Write;
use std::sync::OnceLock;
use tokio::sync::{Mutex, OnceCell};
use super::client::get_docker;
use crate::models::gateway_settings::{GatewaySettings, GatewayStatus};
@@ -58,21 +65,139 @@ const GATEWAY_INTERNAL_PORT: u16 = 4000;
const CONFIG_FINGERPRINT_LABEL: &str = "triple-c.gateway.config-fingerprint";
/// The value a project should use as its base URL (`ANTHROPIC_BASE_URL`).
/// The default bridge gateway address on a stock native-Linux engine. Only a
/// fallback: the real value is read from the `bridge` network's IPAM config.
const DEFAULT_BRIDGE_GATEWAY: &str = "172.17.0.1";
/// Where the gateway's published port is bound on the host, and the address a
/// *project container* uses to reach it.
///
/// Project containers run on Docker's default bridge with no user-defined
/// network and no `--add-host`, so the only address they share with the
/// gateway is the host itself. Publishing the gateway on `0.0.0.0:<port>`
/// makes it reachable from every container network on the machine:
/// network and no `--add-host`, so the only address they share with the gateway
/// is the host itself — but *which* host address works is engine-specific, and
/// the whole point of this type is that the two answers are derived together so
/// they cannot drift apart:
///
/// * Docker Desktop (macOS / Windows / WSL2) resolves `host.docker.internal`
/// from inside containers automatically — that is the portable value and the
/// one already suggested by the existing OpenAI-compatible placeholder text.
/// * On native Linux Docker `host.docker.internal` is not injected, and the
/// equivalent address is the default bridge gateway, normally
/// `http://172.17.0.1:<port>`.
pub fn gateway_base_url(port: u16) -> String {
format!("http://host.docker.internal:{}", port)
/// * **Docker Desktop** (macOS / Windows / WSL2) resolves `host.docker.internal`
/// from inside containers automatically, and its port forwarder reaches the
/// host's *loopback*. So: bind `127.0.0.1`, hand out `host.docker.internal`.
/// * **Native Linux Docker** injects no `host.docker.internal`, and the address
/// containers share with the host is the default bridge gateway (normally
/// `172.17.0.1`). So: bind that address, and hand out the same literal.
///
/// Neither case binds `0.0.0.0`. The bridge-gateway bind is reachable from
/// every container on the default bridge — which is the requirement — without
/// publishing a key-bearing proxy to the LAN.
#[derive(Debug, Clone, PartialEq, Eq)]
pub struct GatewayBinding {
/// Host address the published port is bound to (`HostIp`).
pub host_ip: String,
/// Host address a project container should dial.
pub container_host: String,
}
impl GatewayBinding {
fn desktop() -> Self {
Self {
host_ip: "127.0.0.1".to_string(),
container_host: "host.docker.internal".to_string(),
}
}
fn bridge(gateway_ip: &str) -> Self {
Self {
host_ip: gateway_ip.to_string(),
container_host: gateway_ip.to_string(),
}
}
/// The value a project should use as its base URL (`ANTHROPIC_BASE_URL`).
pub fn base_url(&self, port: u16) -> String {
format!("http://{}:{}", self.container_host, port)
}
/// The address the *host* process (health checks) should dial.
fn host_url(&self, port: u16) -> String {
format!("http://{}:{}", self.host_ip, port)
}
}
/// Decide the binding from what the daemon reports. Pure, so the engine-shape
/// matrix is testable without a daemon.
fn binding_for(operating_system: &str, bridge_gateway: Option<&str>) -> GatewayBinding {
// Docker Desktop reports exactly "Docker Desktop" here on every platform it
// ships for; matched loosely so a future suffix doesn't silently flip us
// onto the bridge path.
if operating_system.to_ascii_lowercase().contains("docker desktop") {
return GatewayBinding::desktop();
}
GatewayBinding::bridge(
bridge_gateway
.map(str::trim)
.filter(|g| !g.is_empty())
.unwrap_or(DEFAULT_BRIDGE_GATEWAY),
)
}
/// Detection is one `info` + one `inspect_network` per process; the answer
/// cannot change without the engine being replaced under us.
static GATEWAY_BINDING: OnceCell<GatewayBinding> = OnceCell::const_new();
/// The gateway's host binding, detected once and cached.
///
/// When Docker is unreachable the *loopback* answer is returned without being
/// cached: it is the conservative one (nothing is published anywhere yet, and
/// the only caller in that state is status reporting), and the next call
/// re-detects once the daemon is up.
pub async fn gateway_binding() -> GatewayBinding {
if let Some(binding) = GATEWAY_BINDING.get() {
return binding.clone();
}
match detect_binding().await {
Ok(binding) => {
let _ = GATEWAY_BINDING.set(binding.clone());
binding
}
Err(e) => {
log::debug!("Gateway bind detection deferred ({}), assuming loopback", e);
GatewayBinding::desktop()
}
}
}
async fn detect_binding() -> Result<GatewayBinding, String> {
let docker = get_docker()?;
let info = docker
.info()
.await
.map_err(|e| format!("Failed to query the Docker daemon: {}", e))?;
let operating_system = info.operating_system.unwrap_or_default();
let gateway_ip = bridge_gateway_ip(&docker).await;
let binding = binding_for(&operating_system, gateway_ip.as_deref());
log::info!(
"Model gateway will publish on {} (engine OS: {})",
binding.host_ip,
if operating_system.is_empty() {
"unknown"
} else {
&operating_system
}
);
Ok(binding)
}
/// The default bridge's gateway address, straight from its IPAM config, so a
/// host whose bridge subnet was customised still gets a reachable bind.
async fn bridge_gateway_ip(docker: &Docker) -> Option<String> {
let network = docker
.inspect_network("bridge", None::<InspectNetworkOptions<String>>)
.await
.ok()?;
network
.ipam?
.config?
.into_iter()
.find_map(|c| c.gateway.filter(|g| !g.trim().is_empty()))
}
fn sha256_hex(input: &str) -> String {
@@ -101,10 +226,31 @@ pub async fn get_gateway_status(settings: &GatewaySettings) -> Result<GatewaySta
image_exists,
model_count: settings.valid_models().len(),
has_api_key: secure::has_gateway_api_key(),
base_url: gateway_base_url(settings.port),
base_url: gateway_binding().await.base_url(settings.port),
})
}
/// Whether a gateway container exists, and whether it is running. Used by the
/// settings reconcile, which must not start anything the user never started.
pub async fn gateway_container_presence() -> Result<(bool, bool), String> {
Ok(match find_gateway_container().await? {
Some((_, state, _)) => (true, state == "running"),
None => (false, false),
})
}
/// Whether a container summary's names contain *exactly* our container.
///
/// Docker's `name` filter is an unanchored regex, so listing with it also
/// returns `triple-c-gateway-backup`, `my-triple-c-gateway`, and anything else
/// containing the string. Taking `.first()` of that would let this module
/// adopt — and then force-remove — a container it does not own.
/// `container::find_existing_container` matches exactly for the same reason.
fn is_gateway_container(names: Option<&Vec<String>>) -> bool {
let expected = format!("/{}", GATEWAY_CONTAINER_NAME);
names.is_some_and(|names| names.iter().any(|n| n == &expected))
}
/// `(id, state, config fingerprint label)` for the gateway container, if any.
async fn find_gateway_container() -> Result<Option<(String, String, String)>, String> {
let docker = get_docker()?;
@@ -123,7 +269,11 @@ async fn find_gateway_container() -> Result<Option<(String, String, String)>, St
.await
.map_err(|e| format!("Failed to list containers: {}", e))?;
if let Some(container) = containers.first() {
// The filter is a prefilter only — the exact-name check is what decides.
for container in &containers {
if !is_gateway_container(container.names.as_ref()) {
continue;
}
let id = container.id.clone().unwrap_or_default();
let state = container.state.clone().unwrap_or_default();
let fingerprint = container
@@ -169,17 +319,21 @@ fn yaml_str(value: &str) -> String {
/// The parts of the config that are safe to hash into a Docker label — i.e.
/// everything except the two secrets, whose changes are tracked by the
/// keychain rotation id instead.
fn config_shape(settings: &GatewaySettings) -> String {
fn config_shape(settings: &GatewaySettings, binding: &GatewayBinding) -> String {
let models: Vec<String> = settings
.valid_models()
.iter()
.map(|m| format!("{}={}", m.name.trim(), m.model_id.trim()))
.collect();
// `bind` is part of the shape so that moving between engines (or a bridge
// subnet change) recreates the container instead of leaving it published on
// an address the new environment doesn't use.
format!(
"provider={};api_base={};port={};models={}",
"provider={};api_base={};port={};bind={};models={}",
settings.provider.trim(),
settings.api_base.as_deref().unwrap_or("").trim(),
settings.port,
binding.host_ip,
models.join(",")
)
}
@@ -273,6 +427,7 @@ async fn upload_config(container_id: &str, config: &str) -> Result<(), String> {
async fn create_gateway_container(
settings: &GatewaySettings,
binding: &GatewayBinding,
fingerprint: &str,
) -> Result<String, String> {
let docker = get_docker()?;
@@ -299,9 +454,9 @@ async fn create_gateway_container(
port_bindings.insert(
format!("{}/tcp", GATEWAY_INTERNAL_PORT),
Some(vec![PortBinding {
// Not loopback — project containers reach this through the host.
// See `gateway_base_url`.
host_ip: Some("0.0.0.0".to_string()),
// Never `0.0.0.0`: the narrowest host address project containers
// can still reach. See `GatewayBinding`.
host_ip: Some(binding.host_ip.clone()),
host_port: Some(settings.port.to_string()),
}]),
);
@@ -328,6 +483,7 @@ async fn create_gateway_container(
"triple-c.gateway.port".to_string(),
settings.port.to_string(),
);
labels.insert("triple-c.gateway.bind".to_string(), binding.host_ip.clone());
labels.insert(
"triple-c.gateway.provider".to_string(),
settings.provider.trim().to_string(),
@@ -365,7 +521,29 @@ async fn create_gateway_container(
Ok(response.id)
}
/// Serialises every mutation of the single fixed-name gateway container.
///
/// `ensure_gateway_running` is check-then-act over one container name, so two
/// concurrent callers — the setup auto-start and the user's Start button is the
/// realistic pair — would both see `None` and both try to create it, and the
/// loser would surface a raw Docker 409. Migration guards the same shape with
/// `ActiveGuard`; here the right behaviour is to *serialise* rather than
/// refuse, because the second caller then observes the first's container, finds
/// a matching fingerprint, and returns its status — which is exactly what it
/// asked for.
fn gateway_lock() -> &'static Mutex<()> {
static LOCK: OnceLock<Mutex<()>> = OnceLock::new();
LOCK.get_or_init(|| Mutex::new(()))
}
pub async fn ensure_gateway_running(settings: &GatewaySettings) -> Result<GatewayStatus, String> {
let _guard = gateway_lock().lock().await;
ensure_gateway_running_locked(settings).await
}
async fn ensure_gateway_running_locked(
settings: &GatewaySettings,
) -> Result<GatewayStatus, String> {
let docker = get_docker()?;
if settings.valid_models().is_empty() {
@@ -382,11 +560,13 @@ pub async fn ensure_gateway_running(settings: &GatewaySettings) -> Result<Gatewa
})?;
let master_key = secure::get_or_create_gateway_master_key()?;
let binding = gateway_binding().await;
// Rotation id, not a hash of either secret — see `storage::secure`.
let secret_version = secure::get_gateway_secret_version()?.unwrap_or_default();
let fingerprint = sha256_hex(&format!(
"{}|{}",
config_shape(settings),
config_shape(settings, &binding),
secret_version
));
@@ -421,7 +601,7 @@ pub async fn ensure_gateway_running(settings: &GatewaySettings) -> Result<Gatewa
.map_err(|e| format!("Failed to remove gateway container: {}", e))?;
}
let id = create_gateway_container(settings, &fingerprint).await?;
let id = create_gateway_container(settings, &binding, &fingerprint).await?;
// Upload before the first start: LiteLLM reads the config once at boot.
let rendered = render_config(settings, &api_key, &master_key);
@@ -446,7 +626,8 @@ pub async fn ensure_gateway_running(settings: &GatewaySettings) -> Result<Gatewa
.map_err(|e| format!("Failed to start gateway container: {}", e))?;
log::info!(
"Model gateway started on port {} ({} model(s))",
"Model gateway started on {}:{} ({} model(s))",
binding.host_ip,
settings.port,
settings.valid_models().len()
);
@@ -454,13 +635,25 @@ pub async fn ensure_gateway_running(settings: &GatewaySettings) -> Result<Gatewa
get_gateway_status(settings).await
}
/// Grace period given to LiteLLM on stop. The Docker default is 10s, which app
/// exit cannot afford to spend on a proxy that holds no state worth flushing.
const GATEWAY_STOP_GRACE_SECS: i64 = 3;
pub async fn stop_gateway_container() -> Result<(), String> {
// Same lock as `ensure_gateway_running`, so a stop can't interleave with a
// create/start and leave a container running behind a "stopped" return.
let _guard = gateway_lock().lock().await;
let docker = get_docker()?;
if let Some((id, state, _)) = find_gateway_container().await? {
if state == "running" {
docker
.stop_container(&id, None::<StopContainerOptions>)
.stop_container(
&id,
Some(StopContainerOptions {
t: GATEWAY_STOP_GRACE_SECS,
}),
)
.await
.map_err(|e| format!("Failed to stop gateway container: {}", e))?;
}
@@ -477,8 +670,12 @@ pub async fn check_gateway_health(port: u16) -> Result<bool, String> {
.build()
.map_err(|e| format!("Failed to build HTTP client: {}", e))?;
// Dial whatever the container is actually published on — with a
// bridge-gateway bind, the host's loopback answers nothing.
let base = gateway_binding().await.host_url(port);
match client
.get(format!("http://127.0.0.1:{}/health/liveliness", port))
.get(format!("{}/health/liveliness", base))
.send()
.await
{
@@ -627,18 +824,126 @@ mod tests {
#[test]
fn config_shape_excludes_secrets_and_tracks_changes() {
let a = config_shape(&settings());
let binding = GatewayBinding::desktop();
let a = config_shape(&settings(), &binding);
let mut s = settings();
s.models[0].model_id = "gpt-4.1".to_string();
assert_ne!(a, config_shape(&s));
assert_ne!(a, config_shape(&s, &binding));
assert!(!a.contains("sk-"));
}
#[test]
fn base_url_points_at_the_host_not_loopback() {
// A project container cannot reach the host's loopback interface.
let url = gateway_base_url(4000);
assert_eq!(url, "http://host.docker.internal:4000");
assert!(!url.contains("127.0.0.1"));
fn config_shape_tracks_the_bind_address() {
// Moving between engines must recreate the container rather than leave
// it published on an address the new environment doesn't use.
let s = settings();
assert_ne!(
config_shape(&s, &GatewayBinding::desktop()),
config_shape(&s, &GatewayBinding::bridge("172.17.0.1"))
);
}
#[test]
fn docker_desktop_binds_loopback_and_hands_out_host_docker_internal() {
let binding = binding_for("Docker Desktop", None);
assert_eq!(binding.host_ip, "127.0.0.1");
assert_eq!(binding.base_url(4000), "http://host.docker.internal:4000");
// Detection must not depend on the bridge answer on this engine.
assert_eq!(binding, binding_for("Docker Desktop", Some("172.17.0.1")));
}
#[test]
fn native_linux_binds_the_bridge_gateway_it_reports() {
// A project container can't reach the host's loopback here, but it can
// reach the bridge gateway — and so can nothing on the LAN.
let binding = binding_for("Ubuntu 24.04.1 LTS", Some("172.19.0.1"));
assert_eq!(binding.host_ip, "172.19.0.1");
assert_eq!(binding.base_url(4000), "http://172.19.0.1:4000");
assert_eq!(binding.host_url(4000), "http://172.19.0.1:4000");
}
#[test]
fn a_missing_bridge_answer_falls_back_to_the_documented_default() {
for reported in [None, Some(""), Some(" ")] {
assert_eq!(
binding_for("Ubuntu 24.04.1 LTS", reported).host_ip,
"172.17.0.1"
);
}
}
#[test]
fn no_engine_shape_ever_binds_a_wildcard_address() {
// The regression this guards: the published port fronts a container
// config holding a billed provider key, and Docker's rules sit ahead of
// the host firewall.
for os in ["Docker Desktop", "Ubuntu 24.04.1 LTS", "", "Rancher Desktop"] {
for gw in [None, Some("172.17.0.1"), Some("10.0.0.1")] {
let host_ip = binding_for(os, gw).host_ip;
assert_ne!(host_ip, "0.0.0.0", "os={:?} gw={:?}", os, gw);
assert_ne!(host_ip, "::", "os={:?} gw={:?}", os, gw);
assert!(!host_ip.is_empty());
}
}
}
#[test]
fn only_the_exact_container_name_is_adopted() {
// Docker's `name` filter is an unanchored regex: all of these come back
// from a filtered list. Adopting one would force-remove a user's
// container.
assert!(is_gateway_container(Some(&vec![
"/triple-c-gateway".to_string()
])));
assert!(is_gateway_container(Some(&vec![
"/something-else".to_string(),
"/triple-c-gateway".to_string(),
])));
for impostor in [
"/triple-c-gateway-backup",
"/my-triple-c-gateway",
"/triple-c-gateway2",
"triple-c-gateway",
] {
assert!(
!is_gateway_container(Some(&vec![impostor.to_string()])),
"{} must not be adopted",
impostor
);
}
assert!(!is_gateway_container(None));
assert!(!is_gateway_container(Some(&vec![])));
}
#[tokio::test]
async fn the_gateway_lock_serialises_concurrent_callers() {
// The auto-start racing the Start button: both would otherwise see no
// container and both create one, and the loser gets a Docker 409.
use std::sync::atomic::{AtomicUsize, Ordering};
use std::sync::Arc;
let inside = Arc::new(AtomicUsize::new(0));
let overlaps = Arc::new(AtomicUsize::new(0));
let mut tasks = Vec::new();
for _ in 0..8 {
let inside = inside.clone();
let overlaps = overlaps.clone();
tasks.push(tokio::spawn(async move {
let _guard = gateway_lock().lock().await;
if inside.fetch_add(1, Ordering::SeqCst) != 0 {
overlaps.fetch_add(1, Ordering::SeqCst);
}
tokio::task::yield_now().await;
tokio::time::sleep(std::time::Duration::from_millis(1)).await;
inside.fetch_sub(1, Ordering::SeqCst);
}));
}
for t in tasks {
t.await.unwrap();
}
assert_eq!(overlaps.load(Ordering::SeqCst), 0);
assert_eq!(inside.load(Ordering::SeqCst), 0);
}
}