Stop a project id from steering a Docker API DELETE, and stop trusting a store that lost its list
C-2 (critical). `destroy_ownerless_rollback_pin` validated its tag and not its
project id, then interpolated both into `triple-c-snapshot-{id}:{tag}` and handed
the result to bollard. bollard does not percent-encode: `Uri::parse` joins an
absolute path onto the base URL, which replaces the path outright and applies RFC
3986 dot-segment removal. An id of `a/../../v1.47/volumes/<name>?` turns a
"remove image tag" into `DELETE /v1.47/volumes/<name>`. That arm is reached
*because* `find_project` failed, so the id is unconstrained IPC input, and the
typed confirmation is no barrier — it compares the caller's own two strings.
Reproduced against the live daemon, and now a test: with the check removed the
volume is gone and the test fails; with it, the volume survives and a legitimate
ownerless pin still deletes. The reference that reaches `remove_image` is now the
daemon's own repo_tag, matched on the parsed pair, so nothing built from IPC
input addresses the API at all. The same id check now guards the owned arms of
`destroy` and `compact_snapshot`, which build volume names and image references
from a `projects.json` field.
H-1. The ownerless arm decided ownership from the in-memory list alone and then
called `sweep_orphaned_snapshots()`, which deletes the freshly dangling image on
the same pass — so a corrupt `projects.json` could reap a pin whose migration is
still awaiting confirmation, the one thing `pin_is_reapable` orders its
conditions to prevent. It now re-reads the store from disk, runs
`project_store_trust`, refuses an id the store knows, takes the project lock
before reading anything a decision rests on, and checks `has_record`.
H-3. The corrupt-store guard keyed on "empty list + file exists", and
`ProjectsStore::new()` swallows a corrupt file without rewriting it — so the
first `save()`, as little as starting a project, wrote `[{new}]` over it and the
guard passed with every other project's volumes unclaimed. A corrupt load is now
recorded in a sticky `projects.json.corrupt` marker beside the file, and the
existing `.bak` is no longer clobbered by a second corruption. A missing
`projects.json` is refused too: it cannot be told from a moved or partially
restored data directory, and the genuinely fresh case has nothing to find.
Also: the three migration commands surface the lock's real refusal instead of
substituting "a migration is already running"; `note_ownerless_since` re-checks
`has_record` after writing a tombstone, closing the window that could plant one
behind a valid record and reap the pin with zero grace; corrupt migration-record
copies are capped at four; `reconcile_migration` yields to any lock holder, not
only a migration; and a 22-space run in a refusal string is gone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GBq2rGum6GX7xXgsas1fDc
This commit is contained in:
@@ -1,9 +1,96 @@
|
||||
use std::fs;
|
||||
use std::path::PathBuf;
|
||||
use std::path::{Path, PathBuf};
|
||||
use std::sync::Mutex;
|
||||
|
||||
use crate::models::Project;
|
||||
|
||||
/// The sticky marker for `projects.json`: `projects.json.corrupt`, beside it.
|
||||
///
|
||||
/// Derived from the file rather than from `dirs::data_dir()` so the marker
|
||||
/// always lands in the directory the store is actually using — and so the
|
||||
/// writer can be tested against a temp directory.
|
||||
fn corrupt_marker_for(file_path: &Path) -> PathBuf {
|
||||
file_path.with_extension("json.corrupt")
|
||||
}
|
||||
|
||||
/// `<data_dir>/triple-c/projects.json.corrupt`, whether or not it exists.
|
||||
pub fn corrupt_marker_path() -> Option<PathBuf> {
|
||||
dirs::data_dir().map(|d| corrupt_marker_for(&d.join("triple-c").join("projects.json")))
|
||||
}
|
||||
|
||||
/// When this data directory last loaded a `projects.json` it could not parse,
|
||||
/// as the RFC3339 instant recorded in the marker.
|
||||
///
|
||||
/// ## Why this outlives the load that wrote it
|
||||
///
|
||||
/// A corrupt load is *recoverable for the app* — the list starts empty and
|
||||
/// everything keeps working — and that recovery is precisely what makes it
|
||||
/// dangerous for anything that reasons about which projects exist. The
|
||||
/// in-memory symptom does not survive: the first [`ProjectsStore::save`] after
|
||||
/// the failure, which is as little as starting one project (`update_status`),
|
||||
/// writes `[{that one project}]` over the file. From then on `projects.json`
|
||||
/// parses, holds one id, and looks exactly like a user with one project — while
|
||||
/// every *other* project's home and config volume is on the daemon claimed by
|
||||
/// nobody.
|
||||
///
|
||||
/// The guard in `project_store_trust` keyed on "the list is empty and the file
|
||||
/// exists", which that write silently ends. So the fact is recorded on disk
|
||||
/// instead of inferred from the list's shape, and it is **sticky**: nothing in
|
||||
/// this app clears it, because nothing in this app can reconstruct what the
|
||||
/// unreadable file held. The refusal names the marker so a user who has
|
||||
/// restored their list — or accepted the loss — can delete it deliberately.
|
||||
pub fn corrupt_since() -> Option<String> {
|
||||
let raw = fs::read_to_string(corrupt_marker_path()?).ok()?;
|
||||
let trimmed = raw.trim();
|
||||
if trimmed.is_empty() {
|
||||
// The marker's presence is the signal; an empty one still means a
|
||||
// corrupt load happened, it just cannot say when.
|
||||
return Some("an unknown time".to_string());
|
||||
}
|
||||
Some(trimmed.lines().next().unwrap_or(trimmed).to_string())
|
||||
}
|
||||
|
||||
/// Keep the bytes of an unparseable `projects.json`, and record that it
|
||||
/// happened.
|
||||
///
|
||||
/// **The existing `.bak` is never overwritten.** A second corruption used to
|
||||
/// clobber the first, and the first is the valuable one: it was taken before
|
||||
/// the app rewrote the file with whatever it had in memory, so it is the only
|
||||
/// copy that can still hold the full project list. Later ones are copies of an
|
||||
/// already-degraded file and get a timestamped name.
|
||||
fn record_corrupt_load(file_path: &Path, now: &chrono::DateTime<chrono::Utc>) {
|
||||
let first = file_path.with_extension("json.bak");
|
||||
let backup = if first.exists() {
|
||||
file_path.with_extension(format!("json.corrupt-{}.bak", now.format("%Y%m%d-%H%M%S")))
|
||||
} else {
|
||||
first
|
||||
};
|
||||
if !backup.exists() {
|
||||
if let Err(e) = fs::copy(file_path, &backup) {
|
||||
log::error!("Failed to back up corrupted projects.json: {}", e);
|
||||
} else {
|
||||
log::error!(
|
||||
"A copy of the unreadable projects.json was kept at {}",
|
||||
backup.display()
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
let marker = corrupt_marker_for(file_path);
|
||||
if marker.exists() {
|
||||
// Sticky: the *first* corruption is the one that dates the loss.
|
||||
return;
|
||||
}
|
||||
if let Err(e) = fs::write(&marker, now.to_rfc3339()) {
|
||||
log::error!(
|
||||
"Could not record the corrupt projects.json load at {}: {} — orphan detection will \
|
||||
not know the project list is incomplete",
|
||||
marker.display(),
|
||||
e
|
||||
);
|
||||
}
|
||||
}
|
||||
|
||||
pub struct ProjectsStore {
|
||||
projects: Mutex<Vec<Project>>,
|
||||
file_path: PathBuf,
|
||||
@@ -43,20 +130,14 @@ impl ProjectsStore {
|
||||
Ok(parsed) => (parsed, migrated),
|
||||
Err(e) => {
|
||||
log::error!("Failed to parse migrated projects.json: {}. Starting with empty list.", e);
|
||||
let backup = file_path.with_extension("json.bak");
|
||||
if let Err(be) = fs::copy(&file_path, &backup) {
|
||||
log::error!("Failed to back up corrupted projects.json: {}", be);
|
||||
}
|
||||
record_corrupt_load(&file_path, &chrono::Utc::now());
|
||||
(Vec::new(), false)
|
||||
}
|
||||
}
|
||||
}
|
||||
Err(e) => {
|
||||
log::error!("Failed to parse projects.json: {}. Starting with empty list.", e);
|
||||
let backup = file_path.with_extension("json.bak");
|
||||
if let Err(be) = fs::copy(&file_path, &backup) {
|
||||
log::error!("Failed to back up corrupted projects.json: {}", be);
|
||||
}
|
||||
record_corrupt_load(&file_path, &chrono::Utc::now());
|
||||
(Vec::new(), false)
|
||||
}
|
||||
}
|
||||
@@ -203,3 +284,89 @@ impl ProjectsStore {
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
|
||||
fn temp_dir(tag: &str) -> PathBuf {
|
||||
let dir = std::env::temp_dir().join(format!(
|
||||
"triple-c-store-{}-{}",
|
||||
tag,
|
||||
uuid::Uuid::new_v4().simple()
|
||||
));
|
||||
fs::create_dir_all(&dir).expect("temp dir");
|
||||
dir
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_corrupt_load_leaves_a_marker_the_next_write_cannot_erase() {
|
||||
// H-3, the whole chain in one test. `ProjectsStore::new()` swallows an
|
||||
// unparseable file into an empty list *without rewriting it*, and the
|
||||
// first `save()` after that — as little as `update_status()` — writes
|
||||
// `[{one project}]` over it. Everything the old guard keyed on ("the
|
||||
// list is empty and the file exists") is gone at that point, while
|
||||
// every *other* project's volumes are still on the daemon claimed by
|
||||
// nobody.
|
||||
let dir = temp_dir("corrupt");
|
||||
let file = dir.join("projects.json");
|
||||
fs::write(&file, "{ this is not a project list").unwrap();
|
||||
|
||||
let now = chrono::Utc::now();
|
||||
record_corrupt_load(&file, &now);
|
||||
|
||||
let marker = corrupt_marker_for(&file);
|
||||
assert!(marker.exists(), "the corrupt load must be recorded on disk");
|
||||
assert_eq!(fs::read_to_string(&marker).unwrap(), now.to_rfc3339());
|
||||
assert!(
|
||||
dir.join("projects.json.bak").exists(),
|
||||
"the unreadable bytes must be kept"
|
||||
);
|
||||
|
||||
// The write that used to erase the evidence. The marker is a separate
|
||||
// file, so it does not care.
|
||||
fs::write(&file, r#"[{"id":"the-one-project-started-since"}]"#).unwrap();
|
||||
assert!(marker.exists());
|
||||
|
||||
fs::remove_dir_all(&dir).ok();
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn a_second_corruption_keeps_the_first_copy_and_the_first_date() {
|
||||
// The `.bak` used to be a fixed name, so a second corruption clobbered
|
||||
// the first — and the first is the only copy taken before the app
|
||||
// rewrote the file with whatever it had in memory, i.e. the only one
|
||||
// that can still hold the full project list.
|
||||
let dir = temp_dir("second");
|
||||
let file = dir.join("projects.json");
|
||||
fs::write(&file, "original bytes").unwrap();
|
||||
let first = chrono::DateTime::parse_from_rfc3339("2026-01-01T00:00:00Z")
|
||||
.unwrap()
|
||||
.with_timezone(&chrono::Utc);
|
||||
record_corrupt_load(&file, &first);
|
||||
|
||||
fs::write(&file, "degraded bytes").unwrap();
|
||||
let second = chrono::DateTime::parse_from_rfc3339("2026-06-01T00:00:00Z")
|
||||
.unwrap()
|
||||
.with_timezone(&chrono::Utc);
|
||||
record_corrupt_load(&file, &second);
|
||||
|
||||
assert_eq!(
|
||||
fs::read_to_string(dir.join("projects.json.bak")).unwrap(),
|
||||
"original bytes",
|
||||
"the first copy must survive the second corruption"
|
||||
);
|
||||
assert_eq!(
|
||||
fs::read_to_string(dir.join("projects.json.corrupt-20260601-000000.bak")).unwrap(),
|
||||
"degraded bytes"
|
||||
);
|
||||
// And the marker still dates the loss from the first failure, which is
|
||||
// when the project list actually stopped being complete.
|
||||
assert_eq!(
|
||||
fs::read_to_string(corrupt_marker_for(&file)).unwrap(),
|
||||
first.to_rfc3339()
|
||||
);
|
||||
|
||||
fs::remove_dir_all(&dir).ok();
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user