Fix review findings: secrets in snapshots, URL spoofing, migration data loss
Adversarial review of the branch produced findings across four areas. This addresses them, plus the Windows CI environment. Secrets. commit_container_snapshot baked the container's full env into the per-project snapshot image, so the shared OAuth token — and the AWS keys, git token and gateway master key — outlived revocation and were readable via docker inspect. Verified against Engine 29.6 that a commit body's config merges over the container's: keys cannot be dropped but can be overwritten, so all of them now commit as KEY=. clear_claude_token additionally rewrites images from earlier builds and reports honestly when a tag could not be rewritten. The recommendation to move the token out of env entirely was not taken, with reasoning: apiKeyHelper is a different auth method that outranks CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no file-based delivery exists. The durable exposure — the image — is what is closed here. Separately noted, not fixed: entrypoint.sh captures the token into the scheduler's .env inside the persisted volume. URL spoofing. Three call sites reached openUrl with container-controlled strings, one of which the review missed (the WebLinksAddon handler). The sign-in URL was scraped from container output with a longest-match tie-break and no userinfo check, so claude.ai@evil.tld rendered as "claude.ai…" in a truncating element. There is now one sanitizer in front of every sink — scheme allowlist, no userinfo, C0/C1 and quote rejection, host allowlist for the sign-in case, first-match — and the origin renders un-truncated. The toast is keyed so a changed URL remounts, closing a bait-and-switch where the user read one URL and clicked another. Migration. The rollback pin was best-effort: a tag failure was logged and the migration continued past remove_container, after which the final commit overwrote the only copy of the old system layer. It now aborts before anything destructive and reads the tag back. /var was destroyed while the ordinary recreate path preserves it — making the "safe" alternative to Reset more destructive than Reset's alternative; data-bearing subtrees are now detected and disclosed in the pre-flight rather than copied, since tarring a live database onto a different base's packages is a corruption risk. resume_migration now verifies the migration-state label instead of reporting success for a container that never swapped. dismiss actually resolves the record rather than leaving the feature permanently refusing to migrate. Start and Reset are guarded while a migration is live. Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and advertised URL are derived together so they cannot drift. Disabling it now stops it. App exit runs teardown concurrently under a budget with a visible shutting-down state instead of blocking for minutes. Auto-starts retry when Docker is not up yet, and the polling-recovery path now reconciles, so interrupted migrations are still recovered. Auth-bridge forwards are capped, closing a container-driven fd exhaustion. Windows CI. build-windows failed on this branch with "linker link.exe not found". The runner had no MSVC build tools and the workflow assumed a hand-provisioned machine, so a bare runner registers, accepts jobs and fails at link time after downloading the whole crate graph. The job now installs the VC++ workload when vswhere cannot find it, matching how it already conditionally installs Rust and Node. 192 Rust tests, 274 frontend tests, both builds clean, zero warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -432,11 +432,51 @@ pub async fn upload_bytes_to_container(
|
||||
Ok(format!("{}/{}", dest_dir.trim_end_matches('/'), file_name))
|
||||
}
|
||||
|
||||
/// Ceiling on how much container output a one-shot exec will buffer into the
|
||||
/// host process.
|
||||
///
|
||||
/// Every `exec_oneshot*` call reads the whole stream into a `String` before any
|
||||
/// caller sees a byte, and what it is reading is *container-controlled* — the
|
||||
/// scheduler notifications reader `cat`s up to 50 files with no size cap, and
|
||||
/// the auth bridge reads `/proc/net/tcp` every two seconds. Neither has an
|
||||
/// upstream bound, so this is where the bound goes. Generous enough that no
|
||||
/// legitimate reader (the largest is a package manifest of a full image) comes
|
||||
/// close.
|
||||
pub const MAX_ONESHOT_OUTPUT: usize = 8 * 1024 * 1024;
|
||||
|
||||
/// The auth bridge's per-tick budget. It reads two procfs files whose rows are
|
||||
/// ~150 bytes; a real container has tens of listeners, and the parser only ever
|
||||
/// yields at most one entry per port number. 1 MiB is thousands of rows — far
|
||||
/// past anything genuine, far short of a problem.
|
||||
pub const PROC_NET_OUTPUT_LIMIT: usize = 1024 * 1024;
|
||||
|
||||
/// Append to `buf` while it stays inside `limit`. Returns `false` once the
|
||||
/// limit is exceeded, at which point the caller must stop reading.
|
||||
fn push_capped(buf: &mut String, chunk: &str, limit: usize) -> bool {
|
||||
if buf.len() + chunk.len() > limit {
|
||||
return false;
|
||||
}
|
||||
buf.push_str(chunk);
|
||||
true
|
||||
}
|
||||
|
||||
/// Run a one-shot (non-interactive) exec command in a container and collect stdout.
|
||||
pub async fn exec_oneshot(container_id: &str, cmd: Vec<String>) -> Result<String, String> {
|
||||
exec_oneshot_env(container_id, cmd, Vec::new()).await
|
||||
}
|
||||
|
||||
/// [`exec_oneshot`] with a caller-chosen output ceiling, for readers whose
|
||||
/// input is fully container-controlled and whose legitimate output is small.
|
||||
pub async fn exec_oneshot_limited(
|
||||
container_id: &str,
|
||||
cmd: Vec<String>,
|
||||
limit: usize,
|
||||
) -> Result<String, String> {
|
||||
exec_oneshot_inner(container_id, "claude", cmd, Vec::new(), limit)
|
||||
.await
|
||||
.map(|(output, _)| output)
|
||||
}
|
||||
|
||||
/// Like `exec_oneshot`, but passes additional environment variables to the exec
|
||||
/// process. Secrets passed this way live only in `/proc/<pid>/environ` (readable
|
||||
/// by the same user / root) rather than in the process argv, so they are not
|
||||
@@ -477,6 +517,16 @@ pub async fn exec_oneshot_as(
|
||||
user: &str,
|
||||
cmd: Vec<String>,
|
||||
env: Vec<String>,
|
||||
) -> Result<(String, i64), String> {
|
||||
exec_oneshot_inner(container_id, user, cmd, env, MAX_ONESHOT_OUTPUT).await
|
||||
}
|
||||
|
||||
async fn exec_oneshot_inner(
|
||||
container_id: &str,
|
||||
user: &str,
|
||||
cmd: Vec<String>,
|
||||
env: Vec<String>,
|
||||
limit: usize,
|
||||
) -> Result<(String, i64), String> {
|
||||
let docker = get_docker()?;
|
||||
|
||||
@@ -505,7 +555,19 @@ pub async fn exec_oneshot_as(
|
||||
StartExecResults::Attached { mut output, .. } => {
|
||||
while let Some(msg) = output.next().await {
|
||||
match msg {
|
||||
Ok(data) => combined.push_str(&String::from_utf8_lossy(&data.into_bytes())),
|
||||
Ok(data) => {
|
||||
let chunk = String::from_utf8_lossy(&data.into_bytes()).into_owned();
|
||||
if !push_capped(&mut combined, &chunk, limit) {
|
||||
// Stop reading rather than truncate silently: every
|
||||
// caller parses this output, and a half-read
|
||||
// manifest or JSON array is worse than an error.
|
||||
// Dropping `output` kills the exec's stream.
|
||||
return Err(format!(
|
||||
"Command output exceeded {} bytes and was abandoned",
|
||||
limit
|
||||
));
|
||||
}
|
||||
}
|
||||
Err(e) => return Err(format!("Exec output error: {}", e)),
|
||||
}
|
||||
}
|
||||
@@ -540,3 +602,42 @@ pub async fn wait_for_exec_exit(exec_id: &str) -> Option<i64> {
|
||||
}
|
||||
None
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
#[test]
|
||||
fn output_under_the_limit_is_buffered_whole() {
|
||||
let mut buf = String::new();
|
||||
assert!(push_capped(&mut buf, "hello ", 16));
|
||||
assert!(push_capped(&mut buf, "world", 16));
|
||||
assert_eq!(buf, "hello world");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn output_over_the_limit_is_refused_rather_than_truncated() {
|
||||
// The abandoned chunk must not land in the buffer either: a caller that
|
||||
// ignored the error would otherwise parse a half-read document.
|
||||
let mut buf = String::new();
|
||||
assert!(push_capped(&mut buf, "0123456789", 12));
|
||||
assert!(!push_capped(&mut buf, "0123456789", 12));
|
||||
assert_eq!(buf, "0123456789");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_single_oversized_chunk_is_refused() {
|
||||
let mut buf = String::new();
|
||||
assert!(!push_capped(&mut buf, "0123456789", 4));
|
||||
assert!(buf.is_empty());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn the_bridge_budget_is_far_smaller_than_the_general_one() {
|
||||
// The auth bridge re-reads container-controlled procfs every 2s, so it
|
||||
// gets a tighter ceiling than one-shot readers that run on demand.
|
||||
assert!(PROC_NET_OUTPUT_LIMIT < MAX_ONESHOT_OUTPUT);
|
||||
// …but still comfortably above a genuine /proc/net/tcp{,6} pair.
|
||||
assert!(PROC_NET_OUTPUT_LIMIT > 100 * 150);
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user