Add llama.cpp backend, model gateway, URL relay and browser view
Four features, plus a latent bug fix.
llama.cpp backend. Claude Code only ever speaks the Anthropic Messages
API — confirmed empirically by pointing it at a logging server, which
received POST /v1/messages?beta=true. llama-server implements that
natively (verified in its README, alongside --port default 8080), so
this is a plain base-URL backend with no translation shim, the same
shape as Ollama. Its --api-key defaults to none, so the auth token is a
placeholder Claude Code requires and llama-server ignores.
Model alias fix. ANTHROPIC_DEFAULT_HAIKU_MODEL is documented as "also
used for background functionality", and Triple-C set none of the alias
vars. So on every custom-endpoint backend, Claude Code resolved `haiku`
to an Anthropic model id and sent it to a local server that does not
have it — background features failed silently. All four
ANTHROPIC_DEFAULT_{OPUS,SONNET,HAIKU,FABLE}_MODEL vars are now pinned to
the backend's configured model, with an optional Haiku override, and
blanked for Anthropic and Bedrock so those keep Claude Code's defaults.
The deprecated ANTHROPIC_SMALL_FAST_MODEL is never emitted. Existing
Ollama and OpenAI-Compatible containers are recreated once so the new
env reaches them; the snapshot is preserved.
Model gateway. Optional LiteLLM sibling container, off by default,
mirroring stt.rs — this is what makes real OpenAI usable, since
api.openai.com has no /v1/messages. Pinned to v1.96.0 by tag and digest:
the 1.82.7/1.82.8 malware was PyPI-only and never affected the official
images, which is precisely why this builds FROM the image rather than
pip-installing, but 1.84.0 is still the floor for proxy CVEs (API-key
SQLi, Host-header auth bypass, MCP auth bypass). Binds 0.0.0.0 because
project containers consume it, and therefore always sets a master_key —
LiteLLM without one accepts any key. The provider key lives in the OS
keychain and is uploaded into a volume, never an image layer or label.
URL relay. A container-side xdg-open/BROWSER shim opens URLs in the
host's browser. Uses an OSC sequence to /dev/tty rather than a printed
sentinel, because the shim usually runs as a grandchild of a process
capturing its children's output. Degrades to printing the URL when no
terminal is attached, so scheduled tasks do not hang. Only http/https,
with control characters rejected before new URL() — which strips
newlines, so java\nscript: would otherwise parse as javascript:. Nothing
auto-opens; the user confirms. The web terminal shows a tap-to-open
banner instead, since that browser may be a phone across a tunnel.
Browser view. A Project Home tab that watches and takes over the browser
Claude drives with Playwright, using Playwright's own dashboard. Zero
image cost — Playwright stays user-installed. It does not reuse the auth
bridge's PortForward, which binds an unauthenticated port: correct for a
throwaway OAuth listener, wrong for mouse and keyboard control of a
browser in a passwordless-sudo container. Instead a token-gated loopback
proxy checks Host, then token or a forbidden-header origin signal,
before a byte reaches the container. Host ports are confined to
47820..=47827 so CSP frame-src can enumerate them rather than widening
to a wildcard, with a test asserting the two agree.
188 frontend tests, 107 Rust tests, both builds clean.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -0,0 +1,99 @@
|
||||
//! Settings and status for the **model gateway** — a LiteLLM proxy container
|
||||
//! Triple-C runs as a sibling of the project containers.
|
||||
//!
|
||||
//! Claude Code speaks only the Anthropic Messages API (`POST
|
||||
//! ${ANTHROPIC_BASE_URL}/v1/messages`). OpenAI has no such route, so an OpenAI
|
||||
//! key cannot drive Claude Code directly. The gateway exposes `/v1/messages`
|
||||
//! in Anthropic format and translates each call to the configured provider,
|
||||
//! which is what turns "OpenAI Compatible" from *bring your own proxy* into
|
||||
//! something Triple-C manages itself.
|
||||
//!
|
||||
//! Nothing secret lives in this module. The provider API key and the gateway's
|
||||
//! own master key are held in the OS keychain (see `storage::secure`); what is
|
||||
//! persisted to `settings.json` is only the non-secret shape of the config.
|
||||
|
||||
use serde::{Deserialize, Serialize};
|
||||
|
||||
/// LiteLLM's own default port, and the one the existing "OpenAI Compatible"
|
||||
/// placeholder text already suggests.
|
||||
pub fn default_gateway_port() -> u16 {
|
||||
4000
|
||||
}
|
||||
|
||||
fn default_gateway_provider() -> String {
|
||||
"openai".to_string()
|
||||
}
|
||||
|
||||
/// One entry of LiteLLM's `model_list`.
|
||||
///
|
||||
/// `name` is the friendly handle a project puts in its model field — it is what
|
||||
/// Claude Code sends as the `model` of a `/v1/messages` request. `model_id` is
|
||||
/// the provider-side id. The gateway config composes them as
|
||||
/// `<provider>/<model_id>`, which is why the shape stays generic across
|
||||
/// providers instead of hard-coding OpenAI.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Default)]
|
||||
pub struct GatewayModel {
|
||||
/// Friendly name projects use (e.g. `gpt-5.1`).
|
||||
pub name: String,
|
||||
/// Provider-side model id (e.g. `gpt-5.1`).
|
||||
pub model_id: String,
|
||||
}
|
||||
|
||||
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)]
|
||||
pub struct GatewaySettings {
|
||||
/// Auto-start the gateway container with the app.
|
||||
#[serde(default)]
|
||||
pub enabled: bool,
|
||||
/// Host port the gateway is published on.
|
||||
#[serde(default = "default_gateway_port")]
|
||||
pub port: u16,
|
||||
/// LiteLLM provider prefix — `openai`, `azure`, `gemini`, `groq`, …
|
||||
#[serde(default = "default_gateway_provider")]
|
||||
pub provider: String,
|
||||
/// Optional provider base URL override (Azure endpoints, proxies, …).
|
||||
#[serde(default)]
|
||||
pub api_base: Option<String>,
|
||||
/// Models the gateway should serve.
|
||||
#[serde(default)]
|
||||
pub models: Vec<GatewayModel>,
|
||||
}
|
||||
|
||||
impl Default for GatewaySettings {
|
||||
fn default() -> Self {
|
||||
Self {
|
||||
enabled: false,
|
||||
port: default_gateway_port(),
|
||||
provider: default_gateway_provider(),
|
||||
api_base: None,
|
||||
models: Vec::new(),
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
impl GatewaySettings {
|
||||
/// Models with both fields filled in. Half-typed rows in the UI must not
|
||||
/// reach the generated YAML.
|
||||
pub fn valid_models(&self) -> Vec<&GatewayModel> {
|
||||
self.models
|
||||
.iter()
|
||||
.filter(|m| !m.name.trim().is_empty() && !m.model_id.trim().is_empty())
|
||||
.collect()
|
||||
}
|
||||
}
|
||||
|
||||
/// What the settings UI needs to know about the gateway. Deliberately carries
|
||||
/// **no** secret: `has_api_key` is a boolean, not the key.
|
||||
#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq)]
|
||||
pub struct GatewayStatus {
|
||||
pub container_exists: bool,
|
||||
pub running: bool,
|
||||
pub port: u16,
|
||||
pub image_exists: bool,
|
||||
/// Number of fully-specified models in the current settings.
|
||||
pub model_count: usize,
|
||||
/// Whether a provider API key is present in the keychain.
|
||||
pub has_api_key: bool,
|
||||
/// The value a project should use for its base URL. See
|
||||
/// `docker::gateway::gateway_base_url`.
|
||||
pub base_url: String,
|
||||
}
|
||||
Reference in New Issue
Block a user