Stop Docker disk growth, and fix the Claude Code settings that never worked

Two independent sets of fixes.

## Disk: stop the growth, no UI this round

The dangling-snapshot sweep was already correct and was never the leak. The
leak is that every `docker commit` **stacks** a layer and nothing compacts one:
a file deleted after it has been committed becomes a whiteout, not free bytes.
24 conditions trigger recreation+commit, so changing one settings field costs a
multi-gigabyte layer for the life of the project. One project was measured with
14 stacked commit layers, ~5.1 GB above its base.

* **Scrub the writable layer before every commit** (`docker/container.rs`,
  `SNAPSHOT_SCRUB_PATHS` / `scrub_writable_layer`). The one moment those bytes
  are still free to drop is before the commit that captures them. Measured on
  one container's 4.48 GB pending layer: 3.0 GB of agent scratchpad under
  `/tmp/claude-*`, the terminal drag-drop staging area (256 MiB per file, with
  no `rm` for it anywhere in the repo), a PNG per pasted image, and the apt
  lists/cache/logs that `browser_view/install.rs` and `triple-c-playwright-heal`
  leave behind with no `apt-get clean`. A hardcoded list, never a heuristic:
  `/workspace/{mount_name}` is a host bind mount and nothing here may reach one,
  and the three `/tmp` globs cannot select the read-only `.host-ca`/`.host-aws`
  mounts. Failure is a log line — a scrub must never block a snapshot.

* **Cap container logs** (`capped_log_config`). There was no `LogConfig`
  anywhere, so containers ran on the daemon's unbounded `json-file` default.
  Deliberately *not* wired into `container_needs_recreation`: participating
  would recreate every project once, and a recreation costs a commit, which is
  the thing being fixed. Picked up on the next natural recreation.

* **Make superseded base images sweepable** (`container/Dockerfile`). It carried
  no `LABEL` at all, so `orphan_sweep_filters`' `dangling` + `triple-c.managed`
  pair provably could not match one — ~11.9 GB observed stranded. Stamping
  `triple-c.managed=true` is the whole fix; the sweep needed no change.
  `create_container` writes the new `triple-c.base` key explicitly empty, or
  Docker's label inheritance plus `docker commit` would make every snapshot
  claim to be a base image. `force: false` stays, and now says why.

* **Sweep at startup** (`lib.rs`), not only after recreation: probes first
  (a probe pins an image the unforced sweep then refuses), pins second, sweep
  last. `sweep_orphaned_snapshots_logged` exists because all three callers threw
  the report away — `reclaimed_bytes`, `failed` and `unavailable` included.

* **Reap migration leftovers.** `rollback_migration` retagged and orphaned the
  migrated snapshot with no sweep. Stale `pre-migration-*` pins are now
  age-reaped by scanning the tag pattern rather than trusting the state file —
  `migration_store::load` reports an unparseable record as absent, which
  stranded a 4-12 GB pin nothing could name again; `load` now moves a corrupt
  record aside so `has_record` is trustworthy. A pin whose migration is still
  awaiting confirmation is never reaped at any age. The probe container's
  removal was a plain statement after an await, so a dropped future (an app quit
  mid-migration) leaked a container pinning a multi-gigabyte image; it is a
  `Drop` guard now, with `reap_probe_containers` for the case where the process
  itself dies.

* **Prune scheduler logs.** `remove` deleted a task's JSON but never its log
  directory, and the task runner appended uncapped `claude -p` output.

* **Fix the delete copy.** It said "the container, config volume, and stored
  credentials"; it removes *both* volumes and the snapshot image.

No prune UI, and no unfiltered `prune_images`/`prune_volumes` anywhere — the
daemon is shared with the user's unrelated work.

## Claude Code settings: two invented keys, one inverted default, one sticky bug

Verified against code.claude.com/docs/en/settings-reference.md and env-vars.md.

* `effort` -> **`effortLevel`**, the key Claude Code actually reads; the old one
  was written and silently ignored. `xhigh` added to the dropdown.
* `focusMode` -> **`viewMode: "focus"`**. `focusMode` was invented. The real key
  does exactly what the existing UI hint already described.
* **Session recap was inverted.** Claude Code's recap is on by default, so
  `CLAUDE_CODE_ENABLE_AWAY_SUMMARY=1`-when-enabled was a no-op and the control
  could never turn the recap *off*. The field is renamed to
  `session_recap_disabled` rather than reused: reusing the name with the
  opposite meaning would have read every stored `enable_session_recap: false` —
  which is every project that never touched the control — as "the user turned
  this off".
* **The stickiness, which is the important one.** Keys were emitted only when
  non-default, and the entrypoint *merges* into a settings.json on a persisted
  volume, so switching a setting off omitted its key, the merge preserved the
  stale on-value, and the setting stayed on until a destructive Reset. The fix
  already existed in the same file — the sandbox block is emitted
  unconditionally for exactly this reason — and is now applied to all five keys.
  A key whose neutral state is *unset* (`tui`, `effortLevel`, `viewMode`,
  `awaySummaryEnabled`) is emitted as JSON `null` and the entrypoint deletes it,
  because a stand-in value is not neutral: `tui: "default"` pins the classic
  renderer where unset lets Claude Code choose, and `viewMode: "default"`
  overrides the user's own sticky `/focus` choice.
* The same stickiness existed, unnoticed, in the **env vars**: `docker commit`
  bakes container env into the snapshot image, so a `=1` written once rode it
  forever. All four are now emitted on every create, extracted into
  `claude_code_env_vars` and unit tested. Two use an empty value for "off"
  rather than `0`, because they outrank a setting the user can change from
  inside their own container and Triple-C's default must not overrule a
  `/config` choice it never asked about.
* TUI mode is now a genuine three-way choice (automatic / classic / fullscreen),
  which the always-emitted key makes both necessary and possible.

`merge_claude_code_settings` is untouched by choice: a project-level OFF still
cannot override a globally-ON setting.

Tests: 364 frontend (+5), 308 Rust (+23), covering the scrub path list and
script, log rotation, pin reaping, and that toggling a setting off actually
clears a previously-set ON value.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GBq2rGum6GX7xXgsas1fDc
This commit is contained in:
2026-08-23 08:35:10 -07:00
co-authored by Claude Opus 5
parent 75cace7dde
commit dd2894cc60
19 changed files with 1469 additions and 112 deletions
+19
View File
@@ -1,5 +1,24 @@
FROM ubuntu:24.04
# ── Provenance labels ────────────────────────────────────────────────────────
# Without these the base image carries no labels at all, and
# `sweep_orphaned_snapshots` (app/src-tauri/src/docker/container.rs) filters on
# `dangling=true` **and** `triple-c.managed=true` — so a superseded base image,
# left untagged when a newer build claims `triple-c-sandbox:latest`, could never
# match and was never collected. ~11.9 GB of stranded base images was measured
# on one developer's daemon this way.
#
# `triple-c.managed=true` is what makes them sweepable. Note that Docker merges
# an image's labels into the containers created from it and `docker commit`
# copies a container's labels onto the image, so this value also arrives on
# every container and every snapshot — which is harmless, because
# `create_container` writes the same key explicitly anyway.
#
# `triple-c.base=true` marks *this* image specifically, so a base image can be
# told apart from a project snapshot without parsing repository names.
LABEL triple-c.managed=true
LABEL triple-c.base=true
# Multi-arch: builds for linux/amd64 and linux/arm64 (Apple Silicon)
# Avoid interactive prompts during package install
ENV DEBIAN_FRONTEND=noninteractive
+27 -12
View File
@@ -405,22 +405,37 @@ install_feature_skill pia-vpn "${VPN_SUPPORT_ENABLED:-0}"
unset VPN_SUPPORT_ENABLED
# ── Claude Code settings ────────────────────────────────────────────────────
# Merge Claude Code settings into ~/.claude/settings.json (preserves existing
# keys). Creates the file if it doesn't exist. These control TUI mode, effort
# level, focus mode, thinking summaries, and other CLI behavior.
# Apply the managed Claude Code settings to ~/.claude/settings.json, keeping
# every key the user set inside the container.
#
# `settings.json` lives on the persisted triple-c-claude-config-{id} volume, so
# it outlives the container and a plain `.[0] * .[1]` merge could only ever
# *add*. That is what made every one of these settings one-way: switching one
# off in Triple-C omitted its key, the merge preserved the old on-value, and the
# setting stayed on until a destructive Reset. So the payload from Rust states
# the whole managed key set on every start, and a JSON **null** in it means
# "delete this key" rather than "merge a null" — which is how a setting whose
# neutral state is *unset* (`tui`, `effortLevel`, `viewMode`,
# `awaySummaryEnabled`) is turned back off without pinning a stand-in value.
# See `build_claude_code_settings_json` in app/src-tauri/src/docker/container.rs.
if [ -n "$CLAUDE_CODE_SETTINGS_JSON" ]; then
SETTINGS_FILE="/home/claude/.claude/settings.json"
mkdir -p /home/claude/.claude
if [ -f "$SETTINGS_FILE" ]; then
# Merge: existing settings + new settings (new keys override on conflict)
MERGED=$(jq -s '.[0] * .[1]' "$SETTINGS_FILE" <(printf '%s' "$CLAUDE_CODE_SETTINGS_JSON") 2>/dev/null)
if [ -n "$MERGED" ]; then
printf '%s\n' "$MERGED" > "$SETTINGS_FILE"
else
echo "entrypoint: warning — failed to merge Claude Code settings into $SETTINGS_FILE"
fi
# One code path for "file exists" and "file doesn't": seeding an empty
# object means the null-deleting merge below runs in both cases, so a fresh
# container never gets a settings.json with literal nulls written into it.
[ -f "$SETTINGS_FILE" ] || printf '{}\n' > "$SETTINGS_FILE"
MERGED=$(jq -s '
.[0] as $current
| .[1] as $managed
| ($managed | with_entries(select(.value != null))) as $set
| ($managed | to_entries | map(select(.value == null) | [.key])) as $clear
| ($current * $set) | delpaths($clear)
' "$SETTINGS_FILE" <(printf '%s' "$CLAUDE_CODE_SETTINGS_JSON") 2>/dev/null)
if [ -n "$MERGED" ]; then
printf '%s\n' "$MERGED" > "$SETTINGS_FILE"
else
printf '%s\n' "$CLAUDE_CODE_SETTINGS_JSON" > "$SETTINGS_FILE"
echo "entrypoint: warning — failed to merge Claude Code settings into $SETTINGS_FILE"
fi
chown claude:claude "$SETTINGS_FILE"
chmod 600 "$SETTINGS_FILE"
+22
View File
@@ -20,6 +20,27 @@ generate_id() {
head -c 4 /dev/urandom | od -An -tx1 | tr -d ' \n'
}
# Delete a task's log directory, called wherever a task stops existing.
#
# The task file is the only index of a task, so a log directory that outlives
# it is unreachable — `logs --id` needs an id nothing can hand you any more —
# and it sits on the home volume for the life of the project. The moment of
# removal is the last point at which we still know what to delete.
#
# The `rm -rf` deserves paranoia, so the id is re-validated here rather than
# trusted from the caller: the pattern rejects an empty id (which would expand
# to $LOGS_DIR itself), anything containing `/` or `.` (which could climb out
# of $LOGS_DIR), and a leading `-`. It matches validate_task_id() in
# app/src-tauri/src/commands/inspect_commands.rs. Always one literal path,
# never a glob.
reap_task_logs() {
local id="${1:-}"
[[ "$id" =~ ^[A-Za-z0-9][A-Za-z0-9_-]*$ ]] || return 0
local dir="${LOGS_DIR:?}/${id}"
[ -d "$dir" ] || return 0
rm -rf -- "$dir"
}
# Live run state for a task: prints "pid<TAB>started_epoch<TAB>log" and returns
# 0 when the task is genuinely running, returns 1 otherwise.
#
@@ -292,6 +313,7 @@ cmd_remove() {
local name
name=$(jq -r '.name' "$task_file")
rm -f "$task_file"
reap_task_logs "$id"
rebuild_crontab
echo "Removed task '$name' ($id)"
}
+60
View File
@@ -125,6 +125,37 @@ fi
echo "=== Exit code: $EXIT_CODE ==="
} >> "$LOG_FILE"
# ── Cap the size of this run's log ──────────────────────────────────────────
# `claude -p` output is unbounded — a task told to walk a large tree can emit
# hundreds of megabytes in one run — and the pruning below counts *files*, not
# bytes, so twenty logs of any size are twenty logs. One chatty task can
# therefore fill the home volume, which is also where ~/.claude and the OAuth
# credential live.
#
# The tail is the half worth keeping: `claude -p` writes its answer at the end,
# and the footer just appended carries the exit code that `status` and the app
# both grep for. So an oversize log is rewritten as a marker line plus its last
# MAX_LOG_BYTES rather than being deleted or capped from the front. This runs
# before the notification below so the summary is taken from the capped file.
#
# Best effort throughout: the run's real result is already recorded, so a
# failure here must not change the exit status. Note that `run` may be tailing
# this file — it has already streamed everything up to here, and nothing is
# appended after this point, so replacing the inode is invisible to it.
MAX_LOG_BYTES=$(( 5 * 1024 * 1024 ))
LOG_BYTES=$(wc -c < "$LOG_FILE" 2>/dev/null || echo 0)
if [ "${LOG_BYTES:-0}" -gt "$MAX_LOG_BYTES" ]; then
TRUNC_FILE="${LOG_FILE}.trunc"
if {
echo "=== Log truncated: $(( LOG_BYTES - MAX_LOG_BYTES )) bytes dropped from the start (cap ${MAX_LOG_BYTES} bytes) ==="
tail -c "$MAX_LOG_BYTES" "$LOG_FILE"
} > "$TRUNC_FILE" 2>/dev/null; then
mv -f "$TRUNC_FILE" "$LOG_FILE" 2>/dev/null || rm -f "$TRUNC_FILE"
else
rm -f "$TRUNC_FILE"
fi
fi
# ── Write notification ──────────────────────────────────────────────────────
mkdir -p "$NOTIFICATIONS_DIR"
NOTIFY_FILE="${NOTIFICATIONS_DIR}/${TASK_ID}_${TIMESTAMP}.notify"
@@ -176,6 +207,35 @@ if [ "$LOG_COUNT" -gt 20 ]; then
find "$TASK_LOG_DIR" -name "*.log" -type f | sort | head -n $((LOG_COUNT - 20)) | xargs rm -f
fi
# ── Reap log dirs of tasks that no longer exist ─────────────────────────────
# `triple-c-scheduler remove` deletes a task's log dir with the task, but a
# one-time task deletes its own task file above, so `remove` can never be run
# for it — nothing knows the id any more — and its directory would sit on the
# home volume forever. This is the sweep for that case.
#
# Deliberately delayed rather than done in the cleanup above: the run that just
# finished has only just written the sole record of itself, `run` and the app's
# Automation tab may still be tailing it, and `logs --id` keeps working for a
# task whose file is gone. So a dir is reaped only once nothing in it has been
# touched for LOG_RETENTION_DAYS, and never while a run is publishing state for
# that id. The sweep rides on task runs, so a container whose only task was
# one-time keeps that one directory until something else runs.
#
# Same paranoia as reap_task_logs() in triple-c-scheduler: the id comes from a
# directory name and is re-validated before it is used to build an `rm -rf`
# path, so no empty or path-bearing name can reach beyond $LOGS_DIR.
LOG_RETENTION_DAYS=7
for ORPHAN_DIR in "$LOGS_DIR"/*/; do
[ -d "$ORPHAN_DIR" ] || continue
ORPHAN_ID=$(basename "$ORPHAN_DIR")
[[ "$ORPHAN_ID" =~ ^[A-Za-z0-9][A-Za-z0-9_-]*$ ]] || continue
[ -f "${TASKS_DIR}/${ORPHAN_ID}.json" ] && continue
[ -f "${RUNNING_DIR}/${ORPHAN_ID}.json" ] && continue
# Anything modified inside the window keeps the whole directory.
[ -n "$(find "$ORPHAN_DIR" -mmin "-$(( LOG_RETENTION_DAYS * 1440 ))" -print -quit 2>/dev/null)" ] && continue
rm -rf -- "${LOGS_DIR:?}/${ORPHAN_ID}"
done
# ── Prune old notifications (keep 50 total) ─────────────────────────────────
NOTIFY_COUNT=$(find "$NOTIFICATIONS_DIR" -name "*.notify" -type f 2>/dev/null | wc -l)
if [ "$NOTIFY_COUNT" -gt 50 ]; then