Add llama.cpp backend, model gateway, URL relay and browser view
Four features, plus a latent bug fix.
llama.cpp backend. Claude Code only ever speaks the Anthropic Messages
API — confirmed empirically by pointing it at a logging server, which
received POST /v1/messages?beta=true. llama-server implements that
natively (verified in its README, alongside --port default 8080), so
this is a plain base-URL backend with no translation shim, the same
shape as Ollama. Its --api-key defaults to none, so the auth token is a
placeholder Claude Code requires and llama-server ignores.
Model alias fix. ANTHROPIC_DEFAULT_HAIKU_MODEL is documented as "also
used for background functionality", and Triple-C set none of the alias
vars. So on every custom-endpoint backend, Claude Code resolved `haiku`
to an Anthropic model id and sent it to a local server that does not
have it — background features failed silently. All four
ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL vars are now pinned to
the backend's configured model, with an optional Haiku override, and
blanked for Anthropic and Bedrock so those keep Claude Code's defaults.
The deprecated ANTHROPIC_SMALL_FAST_MODEL is never emitted. Existing
Ollama and OpenAI-Compatible containers are recreated once so the new
env reaches them; the snapshot is preserved.
Model gateway. Optional LiteLLM sibling container, off by default,
mirroring stt.rs — this is what makes real OpenAI usable, since
api.openai.com has no /v1/messages. Pinned to v1.96.0 by tag and digest:
the 1.82.7/1.82.8 malware was PyPI-only and never affected the official
images, which is precisely why this builds FROM the image rather than
pip-installing, but 1.84.0 is still the floor for proxy CVEs (API-key
SQLi, Host-header auth bypass, MCP auth bypass). Binds 0.0.0.0 because
project containers consume it, and therefore always sets a master_key —
LiteLLM without one accepts any key. The provider key lives in the OS
keychain and is uploaded into a volume, never an image layer or label.
URL relay. A container-side xdg-open/BROWSER shim opens URLs in the
host's browser. Uses an OSC sequence to /dev/tty rather than a printed
sentinel, because the shim usually runs as a grandchild of a process
capturing its children's output. Degrades to printing the URL when no
terminal is attached, so scheduled tasks do not hang. Only http/https,
with control characters rejected before new URL() — which strips
newlines, so java\nscript: would otherwise parse as javascript:. Nothing
auto-opens; the user confirms. The web terminal shows a tap-to-open
banner instead, since that browser may be a phone across a tunnel.
Browser view. A Project Home tab that watches and takes over the browser
Claude drives with Playwright, using Playwright's own dashboard. Zero
image cost — Playwright stays user-installed. It does not reuse the auth
bridge's PortForward, which binds an unauthenticated port: correct for a
throwaway OAuth listener, wrong for mouse and keyboard control of a
browser in a passwordless-sudo container. Instead a token-gated loopback
proxy checks Host, then token or a forbidden-header origin signal,
before a byte reaches the container. Host ports are confined to
47820..=47827 so CSP frame-src can enumerate them rather than widening
to a wildcard, with a test asserting the two agree.
188 frontend tests, 107 Rust tests, both builds clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,56 @@
|
||||
# Triple-C model gateway — a LiteLLM proxy that speaks the Anthropic Messages
|
||||
# API on the front and OpenAI (or any other LiteLLM provider) on the back.
|
||||
#
|
||||
# Claude Code only ever talks the Anthropic Messages API: it POSTs to
|
||||
# `${ANTHROPIC_BASE_URL}/v1/messages`. OpenAI has no such route, so an OpenAI
|
||||
# key cannot be pointed at Claude Code directly. LiteLLM's proxy exposes
|
||||
# `/v1/messages` in Anthropic format and translates each request to the
|
||||
# configured provider, which is what makes first-class OpenAI support possible.
|
||||
#
|
||||
# ── Why this image is built FROM the official LiteLLM image ──────────────────
|
||||
# We deliberately do NOT `pip install litellm` ourselves. The official image is
|
||||
# built by the LiteLLM maintainers from a locked dependency set and already
|
||||
# carries the proxy entrypoint (`docker/prod_entrypoint.sh`), the admin UI and
|
||||
# the Prisma bits. Hand-rolling a pip install would mean re-resolving the whole
|
||||
# dependency tree on every build — exactly the surface that got poisoned in the
|
||||
# incident below — and would drift from what upstream tests.
|
||||
#
|
||||
# ── Why the version is PINNED and why THIS version ───────────────────────────
|
||||
# LiteLLM 1.82.7 and 1.82.8 were published to PyPI containing malware. Anything
|
||||
# that resolves `litellm` at build time (or floats a `latest`/`main-latest` tag)
|
||||
# can silently land on a compromised build, so the tag here is pinned to an
|
||||
# exact release and additionally pinned by digest — a tag can be re-pushed, a
|
||||
# digest cannot.
|
||||
#
|
||||
# The malicious wheels never became images — the official images build from a
|
||||
# pinned requirements.txt and were confirmed unaffected (GHSA-5mg7-485q-xm76,
|
||||
# https://docs.litellm.ai/blog/security-update-march-2026). The reason to pin is
|
||||
# to keep it that way: a floating tag re-resolves, and this container ends up
|
||||
# holding an OpenAI API key.
|
||||
#
|
||||
# v1.96.0 is chosen rather than the nearest clean post-incident release (1.83.0)
|
||||
# because the intervening releases fix a run of *ordinary* proxy CVEs that matter
|
||||
# for a gateway published on a host port: SQL injection in API-key verification
|
||||
# (CVE-2026-42208, fixed 1.83.7), auth bypass via Host header injection
|
||||
# (CVE-2026-49468, fixed 1.84.0) and MCP auth bypass (CVE-2026-59822, fixed
|
||||
# 1.84.0). 1.84.0 is the floor; v1.96.0 is the newest release with no open OSV
|
||||
# advisories at the time of writing. Note LiteLLM has retired the `main-*` tag
|
||||
# family (`main-latest` is no longer updated, `main-stable` is deprecated) in
|
||||
# favour of plain semver tags, which is why this is `v1.96.0` and not
|
||||
# `main-v1.96.0-stable`.
|
||||
#
|
||||
# Bump deliberately, never automatically, and re-check the digest when you do.
|
||||
FROM ghcr.io/berriai/litellm:v1.96.0@sha256:90d8de0ea6fbb3cad145d1019d00a0149ae400b1e18e2011a60f1988f143f672
|
||||
|
||||
# A placeholder config so the image is runnable on its own. Triple-C overwrites
|
||||
# /etc/litellm/config.yaml (a named volume) with the generated one before the
|
||||
# container is started for the first time — the generated file holds the
|
||||
# provider API key, so it is never baked into an image layer.
|
||||
COPY config.yaml /etc/litellm/config.yaml
|
||||
|
||||
EXPOSE 4000
|
||||
|
||||
# The base image's ENTRYPOINT is `docker/prod_entrypoint.sh`, which execs
|
||||
# `litellm "$@"`. These are the arguments Triple-C also passes explicitly when
|
||||
# it runs the unmodified upstream image instead of this one.
|
||||
CMD ["--config", "/etc/litellm/config.yaml", "--host", "0.0.0.0", "--port", "4000"]
|
||||
@@ -0,0 +1,27 @@
|
||||
# Placeholder LiteLLM config baked into the Triple-C gateway image.
|
||||
#
|
||||
# This file exists only so the image starts on its own. Triple-C generates the
|
||||
# real config from Settings → Model Gateway and uploads it over this path
|
||||
# (/etc/litellm/config.yaml) before the container's first start, because the
|
||||
# generated file contains the provider API key and must not live in an image
|
||||
# layer.
|
||||
#
|
||||
# The generated file has the same shape:
|
||||
#
|
||||
# model_list:
|
||||
# - model_name: <friendly name a project puts in ANTHROPIC_MODEL>
|
||||
# litellm_params:
|
||||
# model: <provider>/<provider-side model id>
|
||||
# api_key: <key read from the OS keychain>
|
||||
# api_base: <optional provider base URL override>
|
||||
# general_settings:
|
||||
# master_key: <generated; projects send it as ANTHROPIC_AUTH_TOKEN>
|
||||
# litellm_settings:
|
||||
# drop_params: true
|
||||
|
||||
model_list: []
|
||||
|
||||
litellm_settings:
|
||||
# Anthropic-format requests carry fields some providers reject outright.
|
||||
# Dropping unsupported params is what makes the translation survive.
|
||||
drop_params: true
|
||||
Reference in New Issue
Block a user