feat: Phase 2 — embeddings + hybrid memory.search
Adds a small embedder sidecar (Xenova/bge-small-en-v1.5, ONNX, CPU-only) that the web app calls inline on memory.write and memory.update, and on demand from the new memory.search tool. memory.search performs three candidate fetches in parallel — pgvector cosine similarity, Postgres full-text via plainto_tsquery + ts_rank_cd, and tag-set overlap — then fuses them with Reciprocal Rank Fusion (k=60). Each result carries its per-source rank so the model can see *why* a memory surfaced. The migrator boot step gained an idempotent embedding backfill: any row with embedding IS NULL is batched (32 at a time) through the embedder after SQL migrations apply. Safe to run on every boot. New tool memory.update fixes the missing edit path; centralises the re-embed-on-content-change rule alongside write. Stack additions: - apps/embedder/ — Fastify server, persistent /data/models volume so the ~30 MB model only downloads once - apps/web/lib/embedder.ts — typed HTTP client with batched embed + health probe - packages/schemas — MemoryUpdateInput, MemorySearchInput - docker-compose — embedder service, healthcheck, app + migrator both depend_on it healthy; EMBEDDER_URL promoted to a required env var Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This commit is contained in:
+2
-2
@@ -48,9 +48,9 @@ POSTGRES_DB=memory
|
||||
# DATABASE_URL=postgres://memory:...@db:5432/memory
|
||||
|
||||
# -----------------------------------------------------------------------------
|
||||
# Embedder sidecar (added in Phase 2; leave EMBEDDER_URL empty in Phase 1)
|
||||
# Embedder sidecar. Default points at the in-compose service.
|
||||
# -----------------------------------------------------------------------------
|
||||
EMBEDDER_URL=
|
||||
EMBEDDER_URL=http://embedder:8080
|
||||
EMBEDDING_MODEL=Xenova/bge-small-en-v1.5
|
||||
EMBEDDING_DIM=384
|
||||
|
||||
|
||||
Reference in New Issue
Block a user