Commit Graph
3 Commits
Author SHA1 Message Date
shadow-testandClaude Opus 5 6abc7f27a4 Fix disk/migration defects and add a real per-project lock
The compaction panel's headline action had never worked, three reclaim
paths could delete data with no confirmation and no grace period, and the
app's only mutual-exclusion primitive was one-way.

**A per-project lock (`project_lock.rs`).** `ACTIVE_MIGRATIONS` was the
app's only exclusion and everything but migration merely *polled* it once
at entry. Compaction, start/stop/recreate, Reset and destroy now
**acquire** a `ProjectGuard` and hold it for the whole operation;
`is_migrating` is a view onto the same registry. Closes the three verified
interleavings where a compaction commits `flat(A)` over a `:latest` that a
migration, a recreate or a Reset had already moved. In-process only — the
two-instance case is documented in the module, not solved, and the
daemon-wide reapers gained age gates to bound it.

**H1: compaction never ran.** `fold_shell_script` joined the scrub script's
lines with a space, so every build died on `syntax error: unexpected "do"`.
Replaced with the JSON exec form, which carries any script verbatim;
`sh -n` and a real end-to-end build now cover it (159.5 MB / 9 layers ->
33.7 MB / 1 layer, setuid and multi-line env preserved).

**H2/H4:** `reclaim_migration_pins` and `survey_rollback_pins` apply
`parse_rollback_tag` and `pin_is_reapable` like every other path, and stop
double-counting an image with two pin tags. The 14-day grace period is
re-anchored from the tag's timestamp (when the migration *started*) to a
tombstone recording when the record went missing, with clock skew handled
in both directions.

**H3:** `migration_store::load` no longer renames a corrupt record aside —
that destroyed the `has_record` signal both pin reapers depend on. `save`
fsyncs the file and the directory, and corruption backups are timestamped.

**H2b/M2:** a crashed compaction's `:compacting` tag and `triple-c-compact-*`
container are reaped at startup; the stale-container sweep moved from the
end of a compaction to the start, where its doc always claimed it was.

**M5/M6:** orphan-volume deletion moved from a `Safety::Safe` tick to the
destructive path with a typed volume name; `project_store_trust` reads the
real `projects.json` so a second instance's project is not offered as an
orphan.

Numbers: `images_total_bytes` uses `df()`'s deduplicated `layers_size`; the
Total column is derived from the same figure the Snapshot column shows;
partial container/staging reclaims report their failure count; `human()`
no longer prints "1000.0 KB"; `docker_cli` has a timeout; blocking `fs`
calls moved to `spawn_blocking`.

Also fixes `ProbeContainerGuard::remove_now`, which disarmed before the
await and so did nothing on the cancellation path it exists for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GBq2rGum6GX7xXgsas1fDc
2026-08-23 11:45:28 -07:00
shadow-testandClaude Opus 5 611f67cca7 Fix what review found in the Disk section
Safety:

- `destroy`'s rollback-pin arm took a tag over IPC and interpolated it
  straight into an image reference it then removed. `tag: "latest"` named
  the project's live snapshot, deleted under a dialog saying "rollback
  pin". It is the one destructive variant carrying a free-form string, so
  it now goes through `parse_rollback_tag`.
- The compaction's scratch container was named `triple-c-scrub-*`, which
  is what the scrub reclaim bucket hunts and force-removes. A reclaim from
  a second window would have destroyed the container a running compaction
  was about to commit. It gets `triple-c-compact-*`, swept at the start of
  the next compaction rather than from a bucket anything else can fire.
- Deleting a home or config volume only refused a *running* container, but
  a stopped one still pins its volumes — the resting state of every
  project ever started — so the user typed the project name and met a raw
  409. The container is now removed first and `loses` says so.

Correctness:

- The compaction Dockerfile emitted no `LABEL`, so the flattened
  intermediate could never match the sweep's `dangling` + `triple-c.managed`
  filter that three cleanup paths rely on. Verified on Docker 29.7.2 that
  the label lands on the final stage, the build still yields one layer, and
  untagging the staging tag after the commit leaves the committed snapshot
  intact and startable.
- `snapshot_commit_layers` silently meant something else when
  `triple-c.base-image-id` was absent — the normal case for a pre-label
  project — counting the base's own layers and letting a never-recreated
  project qualify for compaction. `base_lineage_known` now carries that,
  the column says "unknown", and the plan does not offer the rewrite.
- `destroy` returned a `ReclaimResult` wearing a `ReclaimTarget` that named
  work it had not done (a home-volume deletion came back as
  `OrphanVolume`). Split into `target` / `destroyed`, exactly one set.
- `formatBytes` ran `toFixed` after the divide loop, so 999,999 rendered as
  "1000.0 KB" — in the app's only byte formatter, in a panel full of
  near-boundary sizes.
- `is_base_image_reference` split on the first colon, so a registry port
  ate the repo name.

UI:

- `snapshot_above_base_bytes: null` — deliberately unmeasurable — rendered
  as "0 B", the one guessed number in the table.
- Layer count was flagged by colour alone; it now says "stacked".
- The tick list survived a reclaim, so the same call could be re-fired at
  objects that no longer existed. The plan is dropped after any action and
  the panel says the totals predate it.
- `setReport` landed before the plan call was awaited, so a plan failure
  rendered fresh totals above the previous scan's rows.
- Both confirmation modals unmounted before awaiting, making the entire
  busy path dead code during multi-second work.
- `buildx du` failures silently showed `docker system df`'s under-reported
  build-cache figure with no explanation.
- Tooltip text reached no assistive tech, so two headers announced as
  "Help"; hardcoded input id; error-toned glyph in warning-toned panels;
  `sweepOrphanedSnapshots` and `clearOutcome` had no callers.
- Four docstrings claimed things the code did not do, and two tests were
  named for behaviour they did not assert.

Tests: 513 frontend (was 502), 370 Rust (was 365).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GBq2rGum6GX7xXgsas1fDc
2026-08-23 09:51:09 -07:00
shadow-testandClaude Opus 5 77ef2291d7 Add a Disk section: see where the bytes went, and get them back
Every recreation runs `docker commit`, which stacks a layer and never
rewrites one, and 24 conditions in `container_needs_recreation` trigger a
recreation. Prevention landed earlier on this branch; this is the half a
user can act on.

The per-project table leads with the two numbers that explain the
mechanism rather than just the total: how many commit layers a snapshot
has stacked above its base, and what the container's writable layer will
add at the next commit.

Backend (`docker/disk.rs`, commands in `docker_commands.rs`):
- `get_docker_disk_usage` — one `df()` joined against the project store,
  behind an explicit Scan button because it walks the whole daemon.
- `list_reclaimable` / `reclaim` — classified buckets with measured bytes,
  planned off the existing report so re-planning costs no second scan.
- `destroy_project_disk_object` — one object, typed confirmation.
- `sweep_orphaned_snapshots` — exposed, so its report is finally visible.

Safety is structural: `reclaim` takes `ReclaimTarget`, which has no
variant that can name a live project's data. Destructive work is a
separate type reached only through `destroy`. No unfiltered prune is
called anywhere, and nothing outside a `triple-c*` name or `triple-c.*`
label is touched.

Orphan detection subtracts ids from the project store and consults
nothing else. From the daemon's side an idle live project and a deleted
one are indistinguishable — volumes present, no container, no image — so
inferring from container or image absence would offer a live project's
credentials and transcripts for deletion. A store that loaded empty from
an existing `projects.json` is treated as a failed load, not as "no
projects", because `ProjectsStore::new()` recovers from a corrupt file by
starting empty.

Three things verified against a live Docker 29.7.2 rather than assumed:

- Compaction is a two-stage build (`FROM scratch` + `COPY --from`), which
  keeps every byte inside the daemon; bollard's import buffers a whole
  image into memory. uid/gid and setuid survive; a 192.6 MB/4-layer
  synthetic came out 45.7 MB/1 layer. Image config does not survive, so it
  is replayed via create+commit, which round-trips a multi-line env var
  that a Dockerfile `ENV` could not.
- Flattening breaks base-layer sharing, so the result carries its own copy
  of the base. Eight of ten real projects had a 0.10–1.32 GB delta over a
  4.72 GB shared base — compacting those costs ~4 GB. The bound now
  subtracts that penalty, such projects are not offered at all, and the
  run compares unique bytes and abandons a rewrite that would grow.
- `docker builder prune` reports `Total:`, not `Total reclaimed space:`,
  so the first parser scored every prune as freeing nothing.

The Windows/WSL2 note is mandatory and its copy lives in Rust beside the
tests that pin it: pruning frees space inside `ext4.vhdx`, which never
shrinks on its own, so C: does not change until the disk is compacted.

Also adds `lib/formatBytes.ts` — the app had four disagreeing copies, and
`projects/home/format.ts` and `migrationCopy.ts` now delegate to it with
byte-identical output. Base 1000 by default, matching what Docker prints.

Tests: 502 frontend (was 453), 365 Rust (was 322).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GBq2rGum6GX7xXgsas1fDc
2026-08-23 09:39:54 -07:00