Add llama.cpp + OpenAI backends, URL relay, browser view, and base-image migration #14
Merged
jknapp
merged 10 commits from 2026-08-10 06:23:43 +00:00
feature/model-backends-and-browser into main
10
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
e05156fd0e |
Merge branch 'main' into feature/model-backends-and-browser
Build App / compute-version (pull_request) Successful in 3s
Build Container / build-container (pull_request) Successful in 33s
Build App / build-macos (pull_request) Successful in 2m30s
Build App / build-windows (pull_request) Successful in 5m12s
Build App / build-linux (pull_request) Successful in 7m40s
Build App / create-tag (pull_request) Skipped
Build App / sync-to-github (pull_request) Skipped
|
||
|
|
29fd7de909 |
CI: restore the MSI now that the 32-bit bundlers can resolve their paths
Build App / compute-version (pull_request) Successful in 6s
Build Container / build-container (pull_request) Successful in 1m40s
Build App / build-macos (pull_request) Successful in 2m32s
Build App / build-windows (pull_request) Successful in 5m13s
Build App / build-linux (pull_request) Successful in 6m29s
Build App / create-tag (pull_request) Skipped
Build App / sync-to-github (pull_request) Skipped
Dropping the MSI did not help, because the problem was never WiX. makensis.exe is 32-bit exactly like candle.exe and light.exe, lives in the same SYSTEM-profile cache, and failed the same way — "Unable to start child process, error 0x2" instead of 0x80131700. The cause is WOW64 redirection: a 32-bit process reading C:\Windows\System32 is served C:\Windows\SysWOW64, where the toolset directory does not exist, so the bundlers cannot see their own folder. The build VM now carries two junctions from the SysWOW64 view of systemprofile\AppData\Local\tauri and systemprofile\.cache to the System32 originals. Verified afterwards on the runner: candle.exe reports "WiX Toolset Compiler version 3.14.1.8722" and exits 0, and makensis reports v3.11 and exits 0 — both from the same path that failed before. So both targets build again and the .msi asset comes back. Artifact collection fails if either installer is missing rather than tolerating an empty directory. This is a host-side patch for a runner running as SYSTEM. A runner running as a normal user has a LOCALAPPDATA outside System32 and needs none of it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
98a6c8fd56 |
CI: build NSIS only on Windows, dropping the MSI target
The MSI target needs WiX, whose candle.exe and light.exe are 32-bit. On a runner running as SYSTEM, Tauri caches the WiX toolset under %LOCALAPPDATA% = C:\Windows\system32\config\systemprofile\..., and WOW64 redirection points 32-bit processes at SysWOW64, where that directory does not exist. candle.exe cannot see its own folder, the CLR fails to start, and it exits 0x80131700 — surfaced only as "failed to run candle.exe". Proven by running the identical toolset, as the same SYSTEM identity, from C:\wixtest (exit 0) versus the systemprofile path (0x80131700). Because Tauri aborts the entire bundle when one target fails, the MSI was also suppressing the NSIS installer — so Windows produced no artifact at all. NSIS is what the project already relies on for Windows upgrades, so dropping MSI costs the .msi asset and nothing else. Removes the .NET 3.5 gate, which only existed for WiX. The MSVC step stays: that is what makes the app link, and it works. Artifact collection now fails when no NSIS installer is present instead of tolerating an empty directory with 2>nul, so a silent packaging regression cannot pass as a green build again. To restore the MSI later, run the runner as a normal user — whose LOCALAPPDATA sits outside System32 — and set --bundles msi,nsis. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
c71e54a35f |
Revert: LOCALAPPDATA override does not move Tauri's WiX cache
It looked right and did nothing. Rust's `dirs` crate resolves LOCALAPPDATA on Windows through SHGetKnownFolderPath, which reads the process token rather than the environment, so Tauri still cached the WiX toolset under the SYSTEM profile and 32-bit candle.exe still hit WOW64 redirection. Removing it rather than leaving a plausible-looking non-fix in the workflow. The diagnosis in the previous commit stands; only the remedy was wrong. Running the runner as a normal user, whose token resolves LOCALAPPDATA outside System32, is the actual fix. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
b41077e799 |
CI: keep the WiX toolset out of System32 so 32-bit candle.exe can run
build-windows compiled and linked fine but died at bundling with only "failed to run candle.exe". The real cause was neither .NET nor the runner identity, both of which I chased first and was wrong about. candle.exe and light.exe are 32-bit. A runner running as SYSTEM has %LOCALAPPDATA% = C:\Windows\system32\config\systemprofile\AppData\Local, which is where Tauri caches the WiX toolset. WOW64 redirection sends any 32-bit process reading C:\Windows\System32 to C:\Windows\SysWOW64 — and the WixTools directory exists only in the 64-bit view. So candle.exe could not see its own directory, the CLR failed to start, and the process exited 0x80131700, surfaced in the Application event log as ".NET Runtime version 4.0.30319.0 - This application could not be started." Proven rather than assumed: copying the identical toolset to C:\wixtest and running it as the same SYSTEM identity exits 0, while the systemprofile path exits 0x80131700. Test-Path confirms the WOW64 view of that directory does not exist. Pointing LOCALAPPDATA at a path outside System32 avoids redirection. This fixes it for any runner running as a service or as SYSTEM, without needing a stored user credential, and is a no-op where the runner already runs as a normal user. For the record, two earlier theories were wrong. .NET 3.5 was missing and is now installed from the ISO payload, but candle targets .NET 4.x (its config uses loadFromRemoteSources, a 4.0-only element) so that was never the blocker. Adding explicit supportedRuntime entries changed nothing. Both are documented here so the next person does not repeat them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
704d3b8f79 |
CI: fix the MSVC exit-code check and verify .NET 3.5 before bundling
Two follow-ups to the provisioning step. The exit-code check never worked. %VSEXIT% and %ERRORLEVEL% inside a parenthesised cmd block are substituted when the block is PARSED, not when it runs, so the installer's real result was never read — the log printed "installer failed with " with an empty code, then continued anyway. It happened to be harmless because the install had in fact succeeded, but a genuine failure would have sailed past. Now uses delayed expansion. Added a .NET 3.5 check. WiX 3.x candle.exe is a .NET 2.0/3.5 application, and Tauri aborts the entire bundle when the MSI target fails — so a missing runtime silently costs the NSIS installer too, not just the MSI. Windows 11 ships NetFx3 as DisabledWithPayloadRemoved and Windows Update could not supply the payload on our runner even across a reboot; it needed /Source from a mounted ISO. Rather than guess, the job now fails early with the exact dism command. Verified on the runner: MSVC Build Tools 2022 installed, the Rust build completed in 3m59s and produced triple-c.exe, and NetFx3 is now Enabled with v2.0.50727 present. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
2de00b3c55 |
Fix review findings: secrets in snapshots, URL spoofing, migration data loss
Adversarial review of the branch produced findings across four areas. This addresses them, plus the Windows CI environment. Secrets. commit_container_snapshot baked the container's full env into the per-project snapshot image, so the shared OAuth token — and the AWS keys, git token and gateway master key — outlived revocation and were readable via docker inspect. Verified against Engine 29.6 that a commit body's config merges over the container's: keys cannot be dropped but can be overwritten, so all of them now commit as KEY=. clear_claude_token additionally rewrites images from earlier builds and reports honestly when a tag could not be rewritten. The recommendation to move the token out of env entirely was not taken, with reasoning: apiKeyHelper is a different auth method that outranks CLAUDE_CODE_OAUTH_TOKEN rather than a transport for it, and no file-based delivery exists. The durable exposure — the image — is what is closed here. Separately noted, not fixed: entrypoint.sh captures the token into the scheduler's .env inside the persisted volume. URL spoofing. Three call sites reached openUrl with container-controlled strings, one of which the review missed (the WebLinksAddon handler). The sign-in URL was scraped from container output with a longest-match tie-break and no userinfo check, so claude.ai@evil.tld rendered as "claude.ai…" in a truncating element. There is now one sanitizer in front of every sink — scheme allowlist, no userinfo, C0/C1 and quote rejection, host allowlist for the sign-in case, first-match — and the origin renders un-truncated. The toast is keyed so a changed URL remounts, closing a bait-and-switch where the user read one URL and clicked another. Migration. The rollback pin was best-effort: a tag failure was logged and the migration continued past remove_container, after which the final commit overwrote the only copy of the old system layer. It now aborts before anything destructive and reads the tag back. /var was destroyed while the ordinary recreate path preserves it — making the "safe" alternative to Reset more destructive than Reset's alternative; data-bearing subtrees are now detected and disclosed in the pre-flight rather than copied, since tarring a live database onto a different base's packages is a corruption risk. resume_migration now verifies the migration-state label instead of reporting success for a container that never swapped. dismiss actually resolves the record rather than leaving the feature permanently refusing to migrate. Start and Reset are guarded while a migration is live. Lifecycle. The gateway no longer publishes on 0.0.0.0 — bind address and advertised URL are derived together so they cannot drift. Disabling it now stops it. App exit runs teardown concurrently under a budget with a visible shutting-down state instead of blocking for minutes. Auto-starts retry when Docker is not up yet, and the polling-recovery path now reconciles, so interrupted migrations are still recovered. Auth-bridge forwards are capped, closing a container-driven fd exhaustion. Windows CI. build-windows failed on this branch with "linker link.exe not found". The runner had no MSVC build tools and the workflow assumed a hand-provisioned machine, so a bare runner registers, accepts jobs and fails at link time after downloading the whole crate graph. The job now installs the VC++ workload when vswhere cannot find it, matching how it already conditionally installs Rust and Node. 192 Rust tests, 274 frontend tests, both builds clean, zero warnings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
eb1324cb16 |
Fix pre-auth bypass in the browser-view proxy
read_head computed the head terminator index and discarded it, returning
the whole receive buffer. authorize() then split that buffer on CRLF and
treated every colon-bearing line as a header, last occurrence winning —
so any bytes a client sent after the head were promoted to headers.
A cross-site fetch with a text/plain body is CORS-safelisted and not
preflighted, so a body of
a=x\r\nSec-Fetch-Site: same-origin\r\n
overwrote the real cross-site value and the gate returned Allow. The
same trick overrode Host, defeating the anti-rebinding check too. The
result was unauthenticated mouse, keyboard and CDP control of a browser
running inside a container that has passwordless sudo — reachable from
any page the user happened to visit, with the port range being a fixed
8-wide window that is trivially scanned.
read_head now returns the terminator index and the caller authorizes
against that slice only, while still replaying the full buffer into the
tunnel so pipelined bodies are not lost. Duplicate Host, Origin and
Sec-Fetch-Site headers are now refused outright rather than resolved
last-wins, since that resolution is what turns any smuggling primitive
into a full bypass and no legitimate client sends two.
Three regression tests, including a guard assertion that the untruncated
buffer really was accepted before, so the test cannot quietly stop
testing the bypass.
Found by adversarial review, verified by extracting the real authorize()
and running the attack against it.
148 Rust tests, frontend build clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
||
|
|
d42b741337 |
Migrate a project onto a new base image without losing its volumes
Projects were pinned to the image they were first created from. Both create paths preferred triple-c-snapshot-<id>:latest whenever it existed, and container_needs_recreation compared the container's live image against the triple-c.image label — which create_container wrote from the same image it created from. A tautology that could never fire. The only escape was Reset, which calls remove_project_volumes and destroys the login, skills and transcripts. Measured consequences on this host: real projects are missing socat (so the auth bridge cannot tunnel) and bubblewrap (so sandbox mode does not work), plus Mission Control and triple-c-sso-refresh, and sit 61 packages behind the base including ca-certificates, openssl and curl. Detection. create_container now writes triple-c.base-image-id (the image ID, not RepoDigests, which local-built and custom images do not have) and triple-c.create-image. container_needs_recreation takes the expected create-image and compares against the latter, so the check means something. base-image-id is deliberately NOT compared: a base bump would otherwise silently recreate from the snapshot, consuming the "you should migrate" signal without migrating. Staleness is surfaced, never acted on automatically. Migration keeps the volumes. /home/claude and ~/.claude are volumes and the image's copy is seed-only — permanently masked after first mount — so the login, ~/.claude.json, skills, transcripts, scheduler tasks, SSH keys, cargo, uv, ruff and Claude Code itself re-attach untouched. Only root-level state is rebuilt: apt packages are replayed against the new base rather than copied, so no stale libc is dragged forward, and /usr/local, /opt and the non-bind-mounted parts of /workspace are copied verbatim with tar --skip-old-files so they can never clobber a newer base binary. docker diff is not used: on a snapshot-derived container it reports only changes since the last commit. Raw image-vs-image diffing is filtered through dpkg ownership because it otherwise lies — 8,677 raw path differences on a real project reduced to 2 genuinely user-authored files, both loose /workspace-root files. Crash safety. snapshot:latest keeps pointing at the old image until the final commit, so any crash before it self-heals on next start. Later crashes are caught by reconcile_project_statuses. The rollback pin is a docker tag: 0.057s and 0 bytes. Rollback restores the system layer only — volumes are never touched — and the UI says so rather than implying a time machine. Fixes an infinite recreation loop shipped with the MCP removal. docker commit propagates labels to the image, so a container created from a snapshot inherited its non-empty triple-c.mcp-fingerprint and the one-shot shim recreated it again on every start, forever. Lineage labels are now always written explicitly. Documents the second, separate bug this uncovered: Dockerfile changes under /home/claude never reach an existing project, migration or not, because the volume masks them. Anything that must stay upgradable belongs in /usr/local/bin or /opt, or must be seeded by entrypoint.sh. 145 Rust tests, 227 frontend tests, both builds clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
cc5f691677 |
Add llama.cpp backend, model gateway, URL relay and browser view
Four features, plus a latent bug fix.
llama.cpp backend. Claude Code only ever speaks the Anthropic Messages
API — confirmed empirically by pointing it at a logging server, which
received POST /v1/messages?beta=true. llama-server implements that
natively (verified in its README, alongside --port default 8080), so
this is a plain base-URL backend with no translation shim, the same
shape as Ollama. Its --api-key defaults to none, so the auth token is a
placeholder Claude Code requires and llama-server ignores.
Model alias fix. ANTHROPIC_DEFAULT_HAIKU_MODEL is documented as "also
used for background functionality", and Triple-C set none of the alias
vars. So on every custom-endpoint backend, Claude Code resolved `haiku`
to an Anthropic model id and sent it to a local server that does not
have it — background features failed silently. All four
ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL vars are now pinned to
the backend's configured model, with an optional Haiku override, and
blanked for Anthropic and Bedrock so those keep Claude Code's defaults.
The deprecated ANTHROPIC_SMALL_FAST_MODEL is never emitted. Existing
Ollama and OpenAI-Compatible containers are recreated once so the new
env reaches them; the snapshot is preserved.
Model gateway. Optional LiteLLM sibling container, off by default,
mirroring stt.rs — this is what makes real OpenAI usable, since
api.openai.com has no /v1/messages. Pinned to v1.96.0 by tag and digest:
the 1.82.7/1.82.8 malware was PyPI-only and never affected the official
images, which is precisely why this builds FROM the image rather than
pip-installing, but 1.84.0 is still the floor for proxy CVEs (API-key
SQLi, Host-header auth bypass, MCP auth bypass). Binds 0.0.0.0 because
project containers consume it, and therefore always sets a master_key —
LiteLLM without one accepts any key. The provider key lives in the OS
keychain and is uploaded into a volume, never an image layer or label.
URL relay. A container-side xdg-open/BROWSER shim opens URLs in the
host's browser. Uses an OSC sequence to /dev/tty rather than a printed
sentinel, because the shim usually runs as a grandchild of a process
capturing its children's output. Degrades to printing the URL when no
terminal is attached, so scheduled tasks do not hang. Only http/https,
with control characters rejected before new URL() — which strips
newlines, so java\nscript: would otherwise parse as javascript:. Nothing
auto-opens; the user confirms. The web terminal shows a tap-to-open
banner instead, since that browser may be a phone across a tunnel.
Browser view. A Project Home tab that watches and takes over the browser
Claude drives with Playwright, using Playwright's own dashboard. Zero
image cost — Playwright stays user-installed. It does not reuse the auth
bridge's PortForward, which binds an unauthenticated port: correct for a
throwaway OAuth listener, wrong for mouse and keyboard control of a
browser in a passwordless-sudo container. Instead a token-gated loopback
proxy checks Host, then token or a forbidden-header origin signal,
before a byte reaches the container. Host ports are confined to
47820..=47827 so CSP frame-src can enumerate them rather than widening
to a wildcard, with a test asserting the two agree.
188 frontend tests, 107 Rust tests, both builds clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|