Files
shadow-testandClaude Opus 5 cc5f691677 Add llama.cpp backend, model gateway, URL relay and browser view
Four features, plus a latent bug fix.

llama.cpp backend. Claude Code only ever speaks the Anthropic Messages
API — confirmed empirically by pointing it at a logging server, which
received POST /v1/messages?beta=true. llama-server implements that
natively (verified in its README, alongside --port default 8080), so
this is a plain base-URL backend with no translation shim, the same
shape as Ollama. Its --api-key defaults to none, so the auth token is a
placeholder Claude Code requires and llama-server ignores.

Model alias fix. ANTHROPIC_DEFAULT_HAIKU_MODEL is documented as "also
used for background functionality", and Triple-C set none of the alias
vars. So on every custom-endpoint backend, Claude Code resolved `haiku`
to an Anthropic model id and sent it to a local server that does not
have it — background features failed silently. All four
ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL vars are now pinned to
the backend's configured model, with an optional Haiku override, and
blanked for Anthropic and Bedrock so those keep Claude Code's defaults.
The deprecated ANTHROPIC_SMALL_FAST_MODEL is never emitted. Existing
Ollama and OpenAI-Compatible containers are recreated once so the new
env reaches them; the snapshot is preserved.

Model gateway. Optional LiteLLM sibling container, off by default,
mirroring stt.rs — this is what makes real OpenAI usable, since
api.openai.com has no /v1/messages. Pinned to v1.96.0 by tag and digest:
the 1.82.7/1.82.8 malware was PyPI-only and never affected the official
images, which is precisely why this builds FROM the image rather than
pip-installing, but 1.84.0 is still the floor for proxy CVEs (API-key
SQLi, Host-header auth bypass, MCP auth bypass). Binds 0.0.0.0 because
project containers consume it, and therefore always sets a master_key —
LiteLLM without one accepts any key. The provider key lives in the OS
keychain and is uploaded into a volume, never an image layer or label.

URL relay. A container-side xdg-open/BROWSER shim opens URLs in the
host's browser. Uses an OSC sequence to /dev/tty rather than a printed
sentinel, because the shim usually runs as a grandchild of a process
capturing its children's output. Degrades to printing the URL when no
terminal is attached, so scheduled tasks do not hang. Only http/https,
with control characters rejected before new URL() — which strips
newlines, so java\nscript: would otherwise parse as javascript:. Nothing
auto-opens; the user confirms. The web terminal shows a tap-to-open
banner instead, since that browser may be a phone across a tunnel.

Browser view. A Project Home tab that watches and takes over the browser
Claude drives with Playwright, using Playwright's own dashboard. Zero
image cost — Playwright stays user-installed. It does not reuse the auth
bridge's PortForward, which binds an unauthenticated port: correct for a
throwaway OAuth listener, wrong for mouse and keyboard control of a
browser in a passwordless-sudo container. Instead a token-gated loopback
proxy checks Host, then token or a forbidden-header origin signal,
before a byte reaches the container. Host ports are confined to
47820..=47827 so CSP frame-src can enumerate them rather than widening
to a wildcard, with a test asserting the two agree.

188 frontend tests, 107 Rust tests, both builds clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-09 16:55:28 -07:00

57 lines
3.4 KiB
Docker

# Triple-C model gateway — a LiteLLM proxy that speaks the Anthropic Messages
# API on the front and OpenAI (or any other LiteLLM provider) on the back.
#
# Claude Code only ever talks the Anthropic Messages API: it POSTs to
# `${ANTHROPIC_BASE_URL}/v1/messages`. OpenAI has no such route, so an OpenAI
# key cannot be pointed at Claude Code directly. LiteLLM's proxy exposes
# `/v1/messages` in Anthropic format and translates each request to the
# configured provider, which is what makes first-class OpenAI support possible.
#
# ── Why this image is built FROM the official LiteLLM image ──────────────────
# We deliberately do NOT `pip install litellm` ourselves. The official image is
# built by the LiteLLM maintainers from a locked dependency set and already
# carries the proxy entrypoint (`docker/prod_entrypoint.sh`), the admin UI and
# the Prisma bits. Hand-rolling a pip install would mean re-resolving the whole
# dependency tree on every build — exactly the surface that got poisoned in the
# incident below — and would drift from what upstream tests.
#
# ── Why the version is PINNED and why THIS version ───────────────────────────
# LiteLLM 1.82.7 and 1.82.8 were published to PyPI containing malware. Anything
# that resolves `litellm` at build time (or floats a `latest`/`main-latest` tag)
# can silently land on a compromised build, so the tag here is pinned to an
# exact release and additionally pinned by digest — a tag can be re-pushed, a
# digest cannot.
#
# The malicious wheels never became images — the official images build from a
# pinned requirements.txt and were confirmed unaffected (GHSA-5mg7-485q-xm76,
# https://docs.litellm.ai/blog/security-update-march-2026). The reason to pin is
# to keep it that way: a floating tag re-resolves, and this container ends up
# holding an OpenAI API key.
#
# v1.96.0 is chosen rather than the nearest clean post-incident release (1.83.0)
# because the intervening releases fix a run of *ordinary* proxy CVEs that matter
# for a gateway published on a host port: SQL injection in API-key verification
# (CVE-2026-42208, fixed 1.83.7), auth bypass via Host header injection
# (CVE-2026-49468, fixed 1.84.0) and MCP auth bypass (CVE-2026-59822, fixed
# 1.84.0). 1.84.0 is the floor; v1.96.0 is the newest release with no open OSV
# advisories at the time of writing. Note LiteLLM has retired the `main-*` tag
# family (`main-latest` is no longer updated, `main-stable` is deprecated) in
# favour of plain semver tags, which is why this is `v1.96.0` and not
# `main-v1.96.0-stable`.
#
# Bump deliberately, never automatically, and re-check the digest when you do.
FROM ghcr.io/berriai/litellm:v1.96.0@sha256:90d8de0ea6fbb3cad145d1019d00a0149ae400b1e18e2011a60f1988f143f672
# A placeholder config so the image is runnable on its own. Triple-C overwrites
# /etc/litellm/config.yaml (a named volume) with the generated one before the
# container is started for the first time — the generated file holds the
# provider API key, so it is never baked into an image layer.
COPY config.yaml /etc/litellm/config.yaml
EXPOSE 4000
# The base image's ENTRYPOINT is `docker/prod_entrypoint.sh`, which execs
# `litellm "$@"`. These are the arguments Triple-C also passes explicitly when
# it runs the unmodified upstream image instead of this one.
CMD ["--config", "/etc/litellm/config.yaml", "--host", "0.0.0.0", "--port", "4000"]