Fix review findings: secrets in snapshots, URL spoofing, migration data loss

Adversarial review of the branch produced findings across four areas.
This addresses them, plus the Windows CI environment.

Secrets. commit_container_snapshot baked the container's full env into
the per-project snapshot image, so the shared OAuth token — and the AWS
keys, git token and gateway master key — outlived revocation and were
readable via docker inspect. Verified against Engine 29.6 that a commit
body's config merges over the container's: keys cannot be dropped but
can be overwritten, so all of them now commit as KEY=. clear_claude_token
additionally rewrites images from earlier builds and reports honestly
when a tag could not be rewritten.

The recommendation to move the token out of env entirely was not taken,
with reasoning: apiKeyHelper is a different auth method that outranks
CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no
file-based delivery exists. The durable exposure — the image — is what
is closed here. Separately noted, not fixed: entrypoint.sh captures the
token into the scheduler's .env inside the persisted volume.

URL spoofing. Three call sites reached openUrl with container-controlled
strings, one of which the review missed (the WebLinksAddon handler).
The sign-in URL was scraped from container output with a longest-match
tie-break and no userinfo check, so claude.ai@evil.tld rendered as
"claude.ai…" in a truncating element. There is now one sanitizer in
front of every sink — scheme allowlist, no userinfo, C0/C1 and quote
rejection, host allowlist for the sign-in case, first-match — and the
origin renders un-truncated. The toast is keyed so a changed URL
remounts, closing a bait-and-switch where the user read one URL and
clicked another.

Migration. The rollback pin was best-effort: a tag failure was logged
and the migration continued past remove_container, after which the
final commit overwrote the only copy of the old system layer. It now
aborts before anything destructive and reads the tag back. /var was
destroyed while the ordinary recreate path preserves it — making the
"safe" alternative to Reset more destructive than Reset's alternative;
data-bearing subtrees are now detected and disclosed in the pre-flight
rather than copied, since tarring a live database onto a different
base's packages is a corruption risk. resume_migration now verifies the
migration-state label instead of reporting success for a container that
never swapped. dismiss actually resolves the record rather than leaving
the feature permanently refusing to migrate. Start and Reset are guarded
while a migration is live.

Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and
advertised URL are derived together so they cannot drift. Disabling it
now stops it. App exit runs teardown concurrently under a budget with a
visible shutting-down state instead of blocking for minutes. Auto-starts
retry when Docker is not up yet, and the polling-recovery path now
reconciles, so interrupted migrations are still recovered. Auth-bridge
forwards are capped, closing a container-driven fd exhaustion.

Windows CI. build-windows failed on this branch with "linker link.exe
not found". The runner had no MSVC build tools and the workflow assumed
a hand-provisioned machine, so a bare runner registers, accepts jobs and
fails at link time after downloading the whole crate graph. The job now
installs the VC++ workload when vswhere cannot find it, matching how it
already conditionally installs Rust and Node.

192 Rust tests, 274 frontend tests, both builds clean, zero warnings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-08-09 19:35:39 -07:00
co-authored by Claude Opus 5
parent eb1324cb16
commit 2de00b3c55
43 changed files with 4348 additions and 346 deletions
+66 -4
View File
@@ -29,6 +29,17 @@ export interface ContainerMigration {
/** Null until the first probe returns, or when the container has never been created. */
staleness: ContainerStaleness | null;
probing: boolean;
/**
* The probe has landed with a complete answer.
*
* Until it does, `apt_delta`, `verbatim_paths` and `unpreserved_data` are all
* "not known", which is indistinguishable from "empty" at every call site
* that reads them. Starting a migration in that state means the modal telling
* the user there was nothing to copy while the backend quietly skips copying
* — so the action is gated on this, not on the probe merely having been
* kicked off.
*/
probeSettled: boolean;
/** True while a migration is running — whether we started it or found it. */
running: boolean;
/** True when the run in progress was recovered from disk, not started here. */
@@ -51,8 +62,15 @@ export interface ContainerMigration {
start: (options: MigrationOptions) => Promise<void>;
keep: () => Promise<void>;
rollback: () => Promise<void>;
/** Clear a report we cannot act on (failed / rolled back). Local only. */
dismiss: () => void;
/**
* Acknowledge a report there is nothing to keep or roll back.
*
* It has to reach the backend, not just clear local state: an
* `awaiting-confirmation` record that is never resolved comes back on the
* next mount *and* makes every future migration refuse with "already has a
* finished migration waiting for a decision".
*/
dismiss: () => Promise<void>;
refresh: () => Promise<void>;
}
@@ -186,6 +204,25 @@ export function useContainerMigration(project: Project): ContainerMigration {
);
}, [progress, running]);
/**
* Re-read the persisted record after a run settles.
*
* A migration that got past the container swap and then failed leaves the
* record at `interrupted` — the container is mid-swap and the only correct
* next actions are Resume and Roll back. Without this the hook would show the
* failure report's Keep button over a half-migrated container, and a *failed
* resume* would clear `interrupted` and never look again, hiding the mid-swap
* container for the rest of the session.
*/
const adoptRecordAfterRun = useCallback(async () => {
try {
const state = await commands.getMigrationState(projectId);
setInterrupted(state?.phase === INTERRUPTED ? state : null);
} catch {
/* Leave whatever we had; a transient IPC failure is not an outcome. */
}
}, [projectId]);
const start = useCallback(
async (options: MigrationOptions) => {
setLog([]);
@@ -213,10 +250,11 @@ export function useContainerMigration(project: Project): ContainerMigration {
} finally {
setRunning(false);
useAppState.getState().setContainerProgress(projectId, null);
await adoptRecordAfterRun();
void refresh();
}
},
[projectId, refresh],
[projectId, refresh, adoptRecordAfterRun],
);
/**
@@ -271,11 +309,35 @@ export function useContainerMigration(project: Project): ContainerMigration {
}
}, [projectId, project.name, refresh, pushToast]);
const dismiss = useCallback(() => setReport(null), []);
/**
* Dismiss resolves the record; it is not a local hide.
*
* `confirm_migration` is the backend's "this decision is made": it drops the
* rollback tag (there is none in this case), deletes the staged payload and
* removes the state file. Skipping it left an `awaiting-confirmation` record
* on disk that reappeared on every mount and made `migrate_project_to_base`
* refuse forever — recoverable only by deleting JSON by hand.
*/
const dismiss = useCallback(async () => {
setBusy(true);
try {
await commands.confirmMigration(projectId);
setReport(null);
} catch (e) {
pushToast({
kind: "error",
message: `Could not clear the update record for “${project.name}`,
detail: String(e),
});
} finally {
setBusy(false);
}
}, [projectId, project.name, pushToast]);
return {
staleness,
probing,
probeSettled: !probing && staleness !== null && !staleness.probe_error,
running,
recovered,
interrupted,