Migrate a project onto a new base image without losing its volumes
Projects were pinned to the image they were first created from. Both create paths preferred triple-c-snapshot-<id>:latest whenever it existed, and container_needs_recreation compared the container's live image against the triple-c.image label — which create_container wrote from the same image it created from. A tautology that could never fire. The only escape was Reset, which calls remove_project_volumes and destroys the login, skills and transcripts. Measured consequences on this host: real projects are missing socat (so the auth bridge cannot tunnel) and bubblewrap (so sandbox mode does not work), plus Mission Control and triple-c-sso-refresh, and sit 61 packages behind the base including ca-certificates, openssl and curl. Detection. create_container now writes triple-c.base-image-id (the image ID, not RepoDigests, which local-built and custom images do not have) and triple-c.create-image. container_needs_recreation takes the expected create-image and compares against the latter, so the check means something. base-image-id is deliberately NOT compared: a base bump would otherwise silently recreate from the snapshot, consuming the "you should migrate" signal without migrating. Staleness is surfaced, never acted on automatically. Migration keeps the volumes. /home/claude and ~/.claude are volumes and the image's copy is seed-only — permanently masked after first mount — so the login, ~/.claude.json, skills, transcripts, scheduler tasks, SSH keys, cargo, uv, ruff and Claude Code itself re-attach untouched. Only root-level state is rebuilt: apt packages are replayed against the new base rather than copied, so no stale libc is dragged forward, and /usr/local, /opt and the non-bind-mounted parts of /workspace are copied verbatim with tar --skip-old-files so they can never clobber a newer base binary. docker diff is not used: on a snapshot-derived container it reports only changes since the last commit. Raw image-vs-image diffing is filtered through dpkg ownership because it otherwise lies — 8,677 raw path differences on a real project reduced to 2 genuinely user-authored files, both loose /workspace-root files. Crash safety. snapshot:latest keeps pointing at the old image until the final commit, so any crash before it self-heals on next start. Later crashes are caught by reconcile_project_statuses. The rollback pin is a docker tag: 0.057s and 0 bytes. Rollback restores the system layer only — volumes are never touched — and the UI says so rather than implying a time machine. Fixes an infinite recreation loop shipped with the MCP removal. docker commit propagates labels to the image, so a container created from a snapshot inherited its non-empty triple-c.mcp-fingerprint and the one-shot shim recreated it again on every start, forever. Lineage labels are now always written explicitly. Documents the second, separate bug this uncovered: Dockerfile changes under /home/claude never reach an existing project, migration or not, because the volume masks them. Anything that must stay upgradable belongs in /usr/local/bin or /opt, or must be seeded by entrypoint.sh. 145 Rust tests, 227 frontend tests, both builds clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -37,6 +37,22 @@ pub async fn create_attached_exec(
|
||||
container_id: &str,
|
||||
cmd: Vec<String>,
|
||||
tty: bool,
|
||||
) -> Result<AttachedExec, String> {
|
||||
create_attached_exec_as(container_id, cmd, tty, "claude", "/workspace").await
|
||||
}
|
||||
|
||||
/// [`create_attached_exec`] with the user and working directory spelled out.
|
||||
///
|
||||
/// Only base-image migration needs this: replaying `apt` and unpacking a
|
||||
/// payload tar at `/` have to run as **root**, and every other caller wants the
|
||||
/// `claude` / `/workspace` defaults that [`create_attached_exec`] supplies. It
|
||||
/// stays the single place an attached exec is opened.
|
||||
pub async fn create_attached_exec_as(
|
||||
container_id: &str,
|
||||
cmd: Vec<String>,
|
||||
tty: bool,
|
||||
user: &str,
|
||||
working_dir: &str,
|
||||
) -> Result<AttachedExec, String> {
|
||||
let docker = get_docker()?;
|
||||
|
||||
@@ -49,8 +65,8 @@ pub async fn create_attached_exec(
|
||||
attach_stderr: Some(true),
|
||||
tty: Some(tty),
|
||||
cmd: Some(cmd),
|
||||
user: Some("claude".to_string()),
|
||||
working_dir: Some("/workspace".to_string()),
|
||||
user: Some(user.to_string()),
|
||||
working_dir: Some(working_dir.to_string()),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
@@ -371,6 +387,51 @@ pub async fn upload_host_file_to_container(
|
||||
Ok(format!("/tmp/{}", dest_name))
|
||||
}
|
||||
|
||||
/// Write `data` into the container at `<dest_dir>/<file_name>` with `mode`.
|
||||
///
|
||||
/// For small, generated files — migration uses it for the `tar -T` include
|
||||
/// list, which can be too long to pass as argv. Anything large should be
|
||||
/// streamed through an attached exec's stdin instead, since this buffers the
|
||||
/// whole payload in memory twice (once raw, once tarred).
|
||||
pub async fn upload_bytes_to_container(
|
||||
container_id: &str,
|
||||
dest_dir: &str,
|
||||
file_name: &str,
|
||||
data: &[u8],
|
||||
mode: u32,
|
||||
) -> Result<String, String> {
|
||||
let docker = get_docker()?;
|
||||
|
||||
let mut tar_buf = Vec::with_capacity(data.len() + 1024);
|
||||
{
|
||||
let mut builder = tar::Builder::new(&mut tar_buf);
|
||||
let mut header = tar::Header::new_gnu();
|
||||
header.set_size(data.len() as u64);
|
||||
header.set_mode(mode);
|
||||
header.set_cksum();
|
||||
builder
|
||||
.append_data(&mut header, file_name, data)
|
||||
.map_err(|e| format!("Failed to create tar entry: {}", e))?;
|
||||
builder
|
||||
.finish()
|
||||
.map_err(|e| format!("Failed to finalize tar: {}", e))?;
|
||||
}
|
||||
|
||||
docker
|
||||
.upload_to_container(
|
||||
container_id,
|
||||
Some(UploadToContainerOptions {
|
||||
path: dest_dir.to_string(),
|
||||
..Default::default()
|
||||
}),
|
||||
tar_buf.into(),
|
||||
)
|
||||
.await
|
||||
.map_err(|e| format!("Failed to upload file to container: {}", e))?;
|
||||
|
||||
Ok(format!("{}/{}", dest_dir.trim_end_matches('/'), file_name))
|
||||
}
|
||||
|
||||
/// Run a one-shot (non-interactive) exec command in a container and collect stdout.
|
||||
pub async fn exec_oneshot(container_id: &str, cmd: Vec<String>) -> Result<String, String> {
|
||||
exec_oneshot_env(container_id, cmd, Vec::new()).await
|
||||
@@ -400,6 +461,22 @@ pub async fn exec_oneshot_env_status(
|
||||
container_id: &str,
|
||||
cmd: Vec<String>,
|
||||
env: Vec<String>,
|
||||
) -> Result<(String, i64), String> {
|
||||
exec_oneshot_as(container_id, "claude", cmd, env).await
|
||||
}
|
||||
|
||||
/// [`exec_oneshot_env_status`] with the user spelled out.
|
||||
///
|
||||
/// Base-image migration is the only caller that needs anything but `claude`:
|
||||
/// `apt-get`, `npm -g` and the payload unpack all run as **root**. Note that
|
||||
/// the container does grant `claude` passwordless sudo, but going through
|
||||
/// `sudo` would put the whole command in `ps` output and add a second failure
|
||||
/// mode to interpret, so the exec is simply created as root.
|
||||
pub async fn exec_oneshot_as(
|
||||
container_id: &str,
|
||||
user: &str,
|
||||
cmd: Vec<String>,
|
||||
env: Vec<String>,
|
||||
) -> Result<(String, i64), String> {
|
||||
let docker = get_docker()?;
|
||||
|
||||
@@ -411,7 +488,7 @@ pub async fn exec_oneshot_env_status(
|
||||
attach_stderr: Some(true),
|
||||
cmd: Some(cmd),
|
||||
env: if env.is_empty() { None } else { Some(env) },
|
||||
user: Some("claude".to_string()),
|
||||
user: Some(user.to_string()),
|
||||
..Default::default()
|
||||
},
|
||||
)
|
||||
|
||||
Reference in New Issue
Block a user