Adversarial review of the branch produced findings across four areas. This addresses them, plus the Windows CI environment. Secrets. commit_container_snapshot baked the container's full env into the per-project snapshot image, so the shared OAuth token — and the AWS keys, git token and gateway master key — outlived revocation and were readable via docker inspect. Verified against Engine 29.6 that a commit body's config merges over the container's: keys cannot be dropped but can be overwritten, so all of them now commit as KEY=. clear_claude_token additionally rewrites images from earlier builds and reports honestly when a tag could not be rewritten. The recommendation to move the token out of env entirely was not taken, with reasoning: apiKeyHelper is a different auth method that outranks CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no file-based delivery exists. The durable exposure — the image — is what is closed here. Separately noted, not fixed: entrypoint.sh captures the token into the scheduler's .env inside the persisted volume. URL spoofing. Three call sites reached openUrl with container-controlled strings, one of which the review missed (the WebLinksAddon handler). The sign-in URL was scraped from container output with a longest-match tie-break and no userinfo check, so claude.ai@evil.tld rendered as "claude.ai…" in a truncating element. There is now one sanitizer in front of every sink — scheme allowlist, no userinfo, C0/C1 and quote rejection, host allowlist for the sign-in case, first-match — and the origin renders un-truncated. The toast is keyed so a changed URL remounts, closing a bait-and-switch where the user read one URL and clicked another. Migration. The rollback pin was best-effort: a tag failure was logged and the migration continued past remove_container, after which the final commit overwrote the only copy of the old system layer. It now aborts before anything destructive and reads the tag back. /var was destroyed while the ordinary recreate path preserves it — making the "safe" alternative to Reset more destructive than Reset's alternative; data-bearing subtrees are now detected and disclosed in the pre-flight rather than copied, since tarring a live database onto a different base's packages is a corruption risk. resume_migration now verifies the migration-state label instead of reporting success for a container that never swapped. dismiss actually resolves the record rather than leaving the feature permanently refusing to migrate. Start and Reset are guarded while a migration is live. Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and advertised URL are derived together so they cannot drift. Disabling it now stops it. App exit runs teardown concurrently under a budget with a visible shutting-down state instead of blocking for minutes. Auto-starts retry when Docker is not up yet, and the polling-recovery path now reconciles, so interrupted migrations are still recovered. Auth-bridge forwards are capped, closing a container-driven fd exhaustion. Windows CI. build-windows failed on this branch with "linker link.exe not found". The runner had no MSVC build tools and the workflow assumed a hand-provisioned machine, so a bare runner registers, accepts jobs and fails at link time after downloading the whole crate graph. The job now installs the VC++ workload when vswhere cannot find it, matching how it already conditionally installs Rust and Node. 192 Rust tests, 274 frontend tests, both builds clean, zero warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
356 lines
13 KiB
TypeScript
356 lines
13 KiB
TypeScript
import { useCallback, useEffect, useRef, useState } from "react";
|
|
import type {
|
|
ContainerStaleness,
|
|
MigrationOptions,
|
|
MigrationReport,
|
|
MigrationState,
|
|
Project,
|
|
} from "../lib/types";
|
|
import {
|
|
MIGRATION_PHASE_AWAITING_CONFIRMATION,
|
|
MIGRATION_PHASE_IN_PROGRESS,
|
|
MIGRATION_PHASE_INTERRUPTED,
|
|
} from "../lib/types";
|
|
import * as commands from "../lib/tauri-commands";
|
|
import { useAppState } from "../store/appState";
|
|
|
|
/**
|
|
* Unsettled phases from `MigrationState.phase` (hyphenated, unlike the
|
|
* outcome phases on `MigrationReport`). Compared as strings on purpose: the
|
|
* backend types this loosely so an unrecognised value from a future build
|
|
* cannot crash the UI, and neither can it here — an unknown phase simply
|
|
* surfaces nothing rather than throwing.
|
|
*/
|
|
const IN_PROGRESS = MIGRATION_PHASE_IN_PROGRESS;
|
|
const INTERRUPTED = MIGRATION_PHASE_INTERRUPTED;
|
|
const AWAITING = MIGRATION_PHASE_AWAITING_CONFIRMATION;
|
|
|
|
export interface ContainerMigration {
|
|
/** Null until the first probe returns, or when the container has never been created. */
|
|
staleness: ContainerStaleness | null;
|
|
probing: boolean;
|
|
/**
|
|
* The probe has landed with a complete answer.
|
|
*
|
|
* Until it does, `apt_delta`, `verbatim_paths` and `unpreserved_data` are all
|
|
* "not known", which is indistinguishable from "empty" at every call site
|
|
* that reads them. Starting a migration in that state means the modal telling
|
|
* the user there was nothing to copy while the backend quietly skips copying
|
|
* — so the action is gated on this, not on the probe merely having been
|
|
* kicked off.
|
|
*/
|
|
probeSettled: boolean;
|
|
/** True while a migration is running — whether we started it or found it. */
|
|
running: boolean;
|
|
/** True when the run in progress was recovered from disk, not started here. */
|
|
recovered: boolean;
|
|
/**
|
|
* A migration the app died in the middle of. It is not running and it has no
|
|
* report: the container is mid-swap until someone resumes or rolls it back.
|
|
*/
|
|
interrupted: MigrationState | null;
|
|
/** Re-enter an interrupted migration. The backend continues the same run. */
|
|
resume: () => Promise<void>;
|
|
/** The settled report, kept until the user keeps, rolls back or dismisses it. */
|
|
report: MigrationReport | null;
|
|
/** Progress lines from `container-progress`, oldest first. */
|
|
log: string[];
|
|
/** The most recent progress line, or null before the first one arrives. */
|
|
phaseMessage: string | null;
|
|
/** True while confirm/rollback is in flight. */
|
|
busy: boolean;
|
|
start: (options: MigrationOptions) => Promise<void>;
|
|
keep: () => Promise<void>;
|
|
rollback: () => Promise<void>;
|
|
/**
|
|
* Acknowledge a report there is nothing to keep or roll back.
|
|
*
|
|
* It has to reach the backend, not just clear local state: an
|
|
* `awaiting-confirmation` record that is never resolved comes back on the
|
|
* next mount *and* makes every future migration refuse with "already has a
|
|
* finished migration waiting for a decision".
|
|
*/
|
|
dismiss: () => Promise<void>;
|
|
refresh: () => Promise<void>;
|
|
}
|
|
|
|
/**
|
|
* Container base-image migration for one project.
|
|
*
|
|
* Three things have to survive a closed modal: the run itself, the progress
|
|
* log, and the report. A migration takes minutes, so the modal is a *view* onto
|
|
* this hook rather than the thing that owns the work — closing it must not
|
|
* cancel anything. The hook lives in `ProjectHome`, above both the modal and
|
|
* the Overview banner, so either surface can be showing at any point.
|
|
*
|
|
* A migration the app died in the middle of is picked up from
|
|
* `getMigrationState` on mount — as `interrupted`, which is offered for resume,
|
|
* or as `awaiting-confirmation`, whose report is put back on screen. Without
|
|
* that, a half-migrated container would look identical to a healthy one, which
|
|
* is the exact failure mode this whole feature exists to fix.
|
|
*/
|
|
export function useContainerMigration(project: Project): ContainerMigration {
|
|
const projectId = project.id;
|
|
const [staleness, setStaleness] = useState<ContainerStaleness | null>(null);
|
|
const [probing, setProbing] = useState(false);
|
|
const [running, setRunning] = useState(false);
|
|
const [recovered, setRecovered] = useState(false);
|
|
const [interrupted, setInterrupted] = useState<MigrationState | null>(null);
|
|
const [report, setReport] = useState<MigrationReport | null>(null);
|
|
const [log, setLog] = useState<string[]>([]);
|
|
const [busy, setBusy] = useState(false);
|
|
const pushToast = useAppState((s) => s.pushToast);
|
|
const progress = useAppState((s) => s.containerProgress[projectId]);
|
|
|
|
// Guards a late response from an earlier project overwriting a newer one.
|
|
const generation = useRef(0);
|
|
|
|
const refresh = useCallback(async () => {
|
|
const gen = ++generation.current;
|
|
if (!project.container_id) {
|
|
setStaleness(null);
|
|
return;
|
|
}
|
|
setProbing(true);
|
|
try {
|
|
const next = await commands.getContainerStaleness(projectId);
|
|
if (gen === generation.current) setStaleness(next);
|
|
} catch {
|
|
// A probe that cannot reach the container is "we do not know", which is
|
|
// an absent banner rather than an error one — the same call is retried
|
|
// whenever the container's status changes.
|
|
if (gen === generation.current) setStaleness(null);
|
|
} finally {
|
|
if (gen === generation.current) setProbing(false);
|
|
}
|
|
}, [projectId, project.container_id]);
|
|
|
|
// Probe staleness when the container settles into a new state. The probe runs
|
|
// two filesystem walks and is explicitly not for polling, so it is skipped
|
|
// mid-transition and mid-run — a reading taken while the container is being
|
|
// swapped describes neither the old system layer nor the new one.
|
|
const settled = project.status !== "starting" && project.status !== "stopping";
|
|
useEffect(() => {
|
|
if (running || !settled) return;
|
|
void refresh();
|
|
}, [refresh, settled, running]);
|
|
|
|
// Crash recovery: adopt whatever the backend still has on record.
|
|
useEffect(() => {
|
|
let cancelled = false;
|
|
commands
|
|
.getMigrationState(projectId)
|
|
.then((state) => {
|
|
if (cancelled || !state) return;
|
|
if (state.phase === IN_PROGRESS) {
|
|
// Something is still driving it; watch rather than restart.
|
|
setRunning(true);
|
|
setRecovered(true);
|
|
} else if (state.phase === INTERRUPTED) {
|
|
// Nothing is driving it. The container is mid-swap and will stay that
|
|
// way until someone resumes — so this must be visible, not silent.
|
|
setInterrupted(state);
|
|
} else if (state.phase === AWAITING && state.report) {
|
|
setReport(state.report);
|
|
}
|
|
})
|
|
.catch(() => {
|
|
/* No recorded state is the normal case. */
|
|
});
|
|
return () => {
|
|
cancelled = true;
|
|
};
|
|
}, [projectId]);
|
|
|
|
// A recovered run has no promise to await, so poll it to completion.
|
|
useEffect(() => {
|
|
if (!running || !recovered) return;
|
|
let cancelled = false;
|
|
const timer = setInterval(() => {
|
|
commands
|
|
.getMigrationState(projectId)
|
|
.then((state: MigrationState | null) => {
|
|
if (cancelled || state?.phase === IN_PROGRESS) return;
|
|
setRunning(false);
|
|
setRecovered(false);
|
|
// A cleared record means it was confirmed or rolled back elsewhere.
|
|
if (!state) {
|
|
void refresh();
|
|
return;
|
|
}
|
|
if (state.phase === INTERRUPTED) {
|
|
setInterrupted(state);
|
|
return;
|
|
}
|
|
if (state.report) setReport(state.report);
|
|
void refresh();
|
|
})
|
|
.catch(() => {
|
|
/* Keep polling; a transient IPC failure is not an outcome. */
|
|
});
|
|
}, 2500);
|
|
return () => {
|
|
cancelled = true;
|
|
clearInterval(timer);
|
|
};
|
|
}, [running, recovered, projectId, refresh]);
|
|
|
|
// Accumulate the shared progress line into a scrollback the modal can show.
|
|
// The store collapses repeats, so identical consecutive apt lines appear once.
|
|
useEffect(() => {
|
|
if (!running || !progress) return;
|
|
setLog((prev) =>
|
|
prev[prev.length - 1] === progress ? prev : [...prev, progress],
|
|
);
|
|
}, [progress, running]);
|
|
|
|
/**
|
|
* Re-read the persisted record after a run settles.
|
|
*
|
|
* A migration that got past the container swap and then failed leaves the
|
|
* record at `interrupted` — the container is mid-swap and the only correct
|
|
* next actions are Resume and Roll back. Without this the hook would show the
|
|
* failure report's Keep button over a half-migrated container, and a *failed
|
|
* resume* would clear `interrupted` and never look again, hiding the mid-swap
|
|
* container for the rest of the session.
|
|
*/
|
|
const adoptRecordAfterRun = useCallback(async () => {
|
|
try {
|
|
const state = await commands.getMigrationState(projectId);
|
|
setInterrupted(state?.phase === INTERRUPTED ? state : null);
|
|
} catch {
|
|
/* Leave whatever we had; a transient IPC failure is not an outcome. */
|
|
}
|
|
}, [projectId]);
|
|
|
|
const start = useCallback(
|
|
async (options: MigrationOptions) => {
|
|
setLog([]);
|
|
setReport(null);
|
|
setRecovered(false);
|
|
setInterrupted(null);
|
|
setRunning(true);
|
|
try {
|
|
const result = await commands.migrateProjectToBase(projectId, options);
|
|
setReport(result);
|
|
} catch (e) {
|
|
// A rejected call means the backend never produced a report. Synthesise
|
|
// the failed shape so the report surface — not a toast that scrolls
|
|
// away — is still what tells the user.
|
|
setReport({
|
|
phase: "failed",
|
|
packages_requested: [],
|
|
packages_installed: [],
|
|
packages_failed: [],
|
|
paths_copied: [],
|
|
features_restored: [],
|
|
rollback_available: false,
|
|
message: String(e),
|
|
});
|
|
} finally {
|
|
setRunning(false);
|
|
useAppState.getState().setContainerProgress(projectId, null);
|
|
await adoptRecordAfterRun();
|
|
void refresh();
|
|
}
|
|
},
|
|
[projectId, refresh, adoptRecordAfterRun],
|
|
);
|
|
|
|
/**
|
|
* Re-enter an interrupted migration. The backend continues that run rather
|
|
* than starting a new one, and the recorded options are replayed as-is — the
|
|
* deltas cannot be recomputed once the container has already been swapped.
|
|
*/
|
|
const resume = useCallback(async () => {
|
|
const pending = interrupted;
|
|
if (!pending) return;
|
|
await start(pending.options);
|
|
}, [interrupted, start]);
|
|
|
|
const keep = useCallback(async () => {
|
|
setBusy(true);
|
|
try {
|
|
await commands.confirmMigration(projectId);
|
|
setReport(null);
|
|
await refresh();
|
|
} catch (e) {
|
|
pushToast({
|
|
kind: "error",
|
|
message: `Could not discard the rollback image for “${project.name}”`,
|
|
detail: String(e),
|
|
});
|
|
} finally {
|
|
setBusy(false);
|
|
}
|
|
}, [projectId, project.name, refresh, pushToast]);
|
|
|
|
const rollback = useCallback(async () => {
|
|
setBusy(true);
|
|
try {
|
|
await commands.rollbackMigration(projectId);
|
|
setReport(null);
|
|
setInterrupted(null);
|
|
pushToast({
|
|
kind: "success",
|
|
message: `“${project.name}” is back on its previous system layer.`,
|
|
detail:
|
|
"Volumes were not touched, so anything written to your home directory or workspace during the update is still there.",
|
|
});
|
|
await refresh();
|
|
} catch (e) {
|
|
pushToast({
|
|
kind: "error",
|
|
message: `Rollback failed for “${project.name}”`,
|
|
detail: String(e),
|
|
});
|
|
} finally {
|
|
setBusy(false);
|
|
}
|
|
}, [projectId, project.name, refresh, pushToast]);
|
|
|
|
/**
|
|
* Dismiss resolves the record; it is not a local hide.
|
|
*
|
|
* `confirm_migration` is the backend's "this decision is made": it drops the
|
|
* rollback tag (there is none in this case), deletes the staged payload and
|
|
* removes the state file. Skipping it left an `awaiting-confirmation` record
|
|
* on disk that reappeared on every mount and made `migrate_project_to_base`
|
|
* refuse forever — recoverable only by deleting JSON by hand.
|
|
*/
|
|
const dismiss = useCallback(async () => {
|
|
setBusy(true);
|
|
try {
|
|
await commands.confirmMigration(projectId);
|
|
setReport(null);
|
|
} catch (e) {
|
|
pushToast({
|
|
kind: "error",
|
|
message: `Could not clear the update record for “${project.name}”`,
|
|
detail: String(e),
|
|
});
|
|
} finally {
|
|
setBusy(false);
|
|
}
|
|
}, [projectId, project.name, pushToast]);
|
|
|
|
return {
|
|
staleness,
|
|
probing,
|
|
probeSettled: !probing && staleness !== null && !staleness.probe_error,
|
|
running,
|
|
recovered,
|
|
interrupted,
|
|
report,
|
|
log,
|
|
phaseMessage: log.length > 0 ? log[log.length - 1] : null,
|
|
busy,
|
|
start,
|
|
resume,
|
|
keep,
|
|
rollback,
|
|
dismiss,
|
|
refresh,
|
|
};
|
|
}
|