Fix what review found in the skill: five real defects
Adversarial review of #29 found bugs I confirmed by reproducing each one. **Every hand-written error message was unreachable.** `tok=$(curl ...)` is a plain assignment, so `set -e` acts on the command substitution before the following `|| die` can run. A wrong password produced exit 22 and no output at all — the most likely way this gets used wrongly, and the least explained. All four captures now go through a `run` helper that takes a *description* rather than echoing the command, because one of them carries the account password in `-u`. **`up` was not idempotent, and the second run destroyed DNS.** The resolv.conf backup was copied unconditionally, so `up --full` twice overwrote the good backup with PIA's own resolvers; the later `down` then "restored" those and left the container with no working DNS and no way back. `up` now runs `down` first. Verified: two `up --full` runs, then `down`, and the backup still holds the original 192.168.65.7. **An empty gateway produced total connectivity loss, reported as healthy.** `$gw` was never validated and `add_route` swallowed every failure to /dev/null. The two half-routes need no gateway and would succeed, so the tunnel captured everything while the exclusions keeping DNS and the Docker host reachable silently did not exist — and `status` still printed "full tunnel". Routes are now fatal on failure, and a via-less default (`$3` is the literal "eth0") is rejected. **The PIA session token was in the process arguments** — confirmed in `ps` and /proc/*/cmdline, a ~24h bearer credential for the account readable by anything in the container. It now goes to curl on stdin as a config. Verified: 60 polls across a full `up`, zero sightings. **The preflight diagnosed the wrong kernel module.** It checked /dev/net/tun and blamed the tun module, but kernel WireGuard is a netlink interface and does not use it — verified by creating one with NET_ADMIN and no tun device. The check is dropped (the container could not have started without the device anyway) and `ip link add` now reports the real dependency. Also: a full tunnel with no DNS servers from PIA used to warn and carry on, which is a tunnel leaking every lookup while reporting itself healthy — now fatal. `down` validates the backup before restoring it, so a truncated one cannot leave the container with no resolver at all. `wg.priv` is shredded on teardown and created under umask 077, because /run rides `docker commit` into the snapshot image. A mistyped `up --ful` is rejected instead of silently giving a test route. entrypoint: `install_feature_skill` gets `local`, a blank-name guard (the disabled branch would otherwise `rm -rf` the whole skills directory under a persisted volume), `-e`/`-L` so a leftover *file* at the destination is cleaned up, and a chown of the parent so `claude` can still add skills of their own when Mission Control is off. When the base image predates the skill it now says so instead of returning silently — and `/opt/triple-c-skills` joins FEATURE_PROBES so the migration pre-flight reports it. Docs corrected to match: neither half reaches an existing project without a migration. `vpn_env_var` extracted and tested, pinning the property the whole removal path rests on — that the variable is emitted as 0 rather than omitted. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -342,12 +342,20 @@ container is created once by a very long function where a dropped capability is
|
||||
directly avoids both, which is what the skill does.
|
||||
- **The `pia-vpn` skill is installed *and removed* from `VPN_SUPPORT_ENABLED`.** `container/skills/`
|
||||
is baked to `/opt/triple-c-skills` and `install_feature_skill()` in `entrypoint.sh` copies it into
|
||||
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches existing projects, and
|
||||
`rm -rf`'d first, so files dropped from a later version do not linger. The removal branch matters
|
||||
as much as the install: `~/.claude` is a persisted volume, so a skill left behind after the toggle
|
||||
goes off would keep instructing an agent to use a capability the container no longer has. Which is
|
||||
also why the variable is sent as `0` rather than omitted, and why it is in `RESERVED_ENV_EXACT` —
|
||||
a custom env var of that name could otherwise claim the skill without the capability behind it.
|
||||
`~/.claude/skills/` on every start — refreshed each time, so a fix reaches any project whose base
|
||||
image has the source, and `rm -rf`'d first, so files dropped from a later version do not linger.
|
||||
The removal branch matters as much as the install: `~/.claude` is a persisted volume, so a skill
|
||||
left behind after the toggle goes off would keep instructing an agent to use a capability the
|
||||
container no longer has. Which is also why the variable is sent as `0` rather than omitted (see
|
||||
`vpn_env_var`, tested), and why it is in `RESERVED_ENV_EXACT` — a custom env var of that name
|
||||
could otherwise claim the skill without the capability behind it.
|
||||
- **Both halves of that live in the base image, so neither reaches an existing project.** A
|
||||
recreation builds from the project's *own snapshot*, which has no `/opt/triple-c-skills` and no
|
||||
updated `entrypoint.sh`; only a migration or a Reset delivers them. The install path says so out
|
||||
loud rather than returning silently, and `/opt/triple-c-skills` is in `FEATURE_PROBES` so the
|
||||
migration pre-flight lists it as missing. Worth knowing before adding anything else behind an
|
||||
existing toggle: the label fingerprints *the setting*, not the set of things the setting drives,
|
||||
so a project already at `true` gets no recreation at all on upgrade.
|
||||
|
||||
### Container Lifecycle
|
||||
|
||||
|
||||
Reference in New Issue
Block a user