Fix review findings: secrets in snapshots, URL spoofing, migration data loss

Adversarial review of the branch produced findings across four areas.
This addresses them, plus the Windows CI environment.

Secrets. commit_container_snapshot baked the container's full env into
the per-project snapshot image, so the shared OAuth token — and the AWS
keys, git token and gateway master key — outlived revocation and were
readable via docker inspect. Verified against Engine 29.6 that a commit
body's config merges over the container's: keys cannot be dropped but
can be overwritten, so all of them now commit as KEY=. clear_claude_token
additionally rewrites images from earlier builds and reports honestly
when a tag could not be rewritten.

The recommendation to move the token out of env entirely was not taken,
with reasoning: apiKeyHelper is a different auth method that outranks
CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no
file-based delivery exists. The durable exposure — the image — is what
is closed here. Separately noted, not fixed: entrypoint.sh captures the
token into the scheduler's .env inside the persisted volume.

URL spoofing. Three call sites reached openUrl with container-controlled
strings, one of which the review missed (the WebLinksAddon handler).
The sign-in URL was scraped from container output with a longest-match
tie-break and no userinfo check, so claude.ai@evil.tld rendered as
"claude.ai…" in a truncating element. There is now one sanitizer in
front of every sink — scheme allowlist, no userinfo, C0/C1 and quote
rejection, host allowlist for the sign-in case, first-match — and the
origin renders un-truncated. The toast is keyed so a changed URL
remounts, closing a bait-and-switch where the user read one URL and
clicked another.

Migration. The rollback pin was best-effort: a tag failure was logged
and the migration continued past remove_container, after which the
final commit overwrote the only copy of the old system layer. It now
aborts before anything destructive and reads the tag back. /var was
destroyed while the ordinary recreate path preserves it — making the
"safe" alternative to Reset more destructive than Reset's alternative;
data-bearing subtrees are now detected and disclosed in the pre-flight
rather than copied, since tarring a live database onto a different
base's packages is a corruption risk. resume_migration now verifies the
migration-state label instead of reporting success for a container that
never swapped. dismiss actually resolves the record rather than leaving
the feature permanently refusing to migrate. Start and Reset are guarded
while a migration is live.

Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and
advertised URL are derived together so they cannot drift. Disabling it
now stops it. App exit runs teardown concurrently under a budget with a
visible shutting-down state instead of blocking for minutes. Auto-starts
retry when Docker is not up yet, and the polling-recovery path now
reconciles, so interrupted migrations are still recovered. Auth-bridge
forwards are capped, closing a container-driven fd exhaustion.

Windows CI. build-windows failed on this branch with "linker link.exe
not found". The runner had no MSVC build tools and the workflow assumed
a hand-provisioned machine, so a bare runner registers, accepts jobs and
fails at link time after downloading the whole crate graph. The job now
installs the VC++ workload when vswhere cannot find it, matching how it
already conditionally installs Rust and Node.

192 Rust tests, 274 frontend tests, both builds clean, zero warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-09 19:35:39 -07:00
co-authored by Claude Opus 5
parent eb1324cb16
commit 2de00b3c55
43 changed files with 4348 additions and 346 deletions
+57 -1
View File
@@ -145,6 +145,11 @@ pub async fn remove_project(
// before the container (and the project record) go away.
state.auth_bridge.stop(&project_id).await;
// A migration record outliving its project leaks a state file, a staged
// payload tar that can run to several GB, and a `:pre-migration-<ts>` tag
// holding an entire snapshot image that nothing will ever reference again.
crate::commands::migration_commands::purge_migration_artifacts(&project_id).await;
// Stop and remove container if it exists
if let Some(ref project) = state.projects_store.get(&project_id) {
if let Some(ref container_id) = project.container_id {
@@ -215,6 +220,20 @@ pub async fn start_project_container(
app_handle: tauri::AppHandle,
state: State<'_, AppState>,
) -> Result<Project, String> {
// A migration removes the container and creates its replacement moments
// later. Starting in that window finds no container, creates a second one
// under the same name, and the migration's own create then fails on the
// name conflict — which sends it into an auto-rollback that also cannot
// create. The UI already refuses (`canMigrate` gates on the container being
// stopped and no run being in flight); this is the same gate on the side
// that actually owns the invariant.
if crate::commands::migration_commands::is_migrating(&project_id) {
return Err(
"A container base update is running for this project. Wait for it to finish, then start the project."
.to_string(),
);
}
let mut project = state
.projects_store
.get(&project_id)
@@ -536,11 +555,28 @@ pub async fn rebuild_project_container(
app_handle: tauri::AppHandle,
state: State<'_, AppState>,
) -> Result<Project, String> {
// Reset deletes both volumes and the snapshot image. Doing that while a
// migration is mid-flight pulls the ground out from under it and leaves an
// orphan migration record pointing at images that no longer exist.
if crate::commands::migration_commands::is_migrating(&project_id) {
return Err(
"A container base update is running for this project. Wait for it to finish before resetting."
.to_string(),
);
}
let project = state
.projects_store
.get(&project_id)
.ok_or_else(|| format!("Project {} not found", project_id))?;
// Reset supersedes any migration decision that was still pending: the
// snapshot image and both volumes are about to go, so a surviving record
// could only describe things that no longer exist — while its
// `:pre-migration-<ts>` tag held a whole snapshot image (multiple GB) alive
// with nothing left that could ever use it.
crate::commands::migration_commands::purge_migration_artifacts(&project_id).await;
// The bridge is bound to the container that is about to be destroyed;
// `start_project_container` below re-arms it against the new one.
state.auth_bridge.stop(&project_id).await;
@@ -588,7 +624,27 @@ pub async fn reconcile_project_statuses(
}
for project in &projects {
if project.status != ProjectStatus::Running && project.status != ProjectStatus::Error {
// `Starting` and `Stopping` are in here as a backstop, not because
// anything is expected to leave a project in one. They are transitional
// states owned by an in-flight command, so a project still wearing one
// is a project whose command died — a crash mid-start, or a migration
// that bailed out between the stop and the swap. Skipping them, as this
// loop used to, meant nothing in the app ever put such a project right:
// it sat at "Stopping" with the Start button disabled, permanently.
// Docker is the authority either way, so the check below is correct for
// all four.
if !matches!(
project.status,
ProjectStatus::Running
| ProjectStatus::Error
| ProjectStatus::Starting
| ProjectStatus::Stopping
) {
continue;
}
// ...but never for a project this process is actively migrating: the
// container is legitimately absent for part of that run.
if crate::commands::migration_commands::is_migrating(&project.id) {
continue;
}