open_notebook/docker-compose.yml
dschlueter 4805fe75ed Pin images by digest; add update, smoke-test and status scripts
Goal: at every session start, run exactly the application the repo says — and be able
to move to a new upstream release deliberately, with a way back.

The old setup (v1-latest + pull_policy: always) was unreliable in both directions.
`docker compose up -d` silently pulled a new application, including an irreversible
SurrealDB schema migration, while a reboot (restart: always) or `docker compose start`
kept running the old image without pulling anything. The running version was effectively
unpredictable.

Note v1-latest tracks releases, not main: the current image (v1.10.0, 2026-06-18) IS the
latest release — the repo being "ahead" is unreleased code, which is explicitly not
wanted here. So this is about guaranteeing and detecting, not catching up.

- docker-compose.yml: both images pinned by digest, pull_policy: missing.
- scripts/check_updates.sh: reports running version/digest vs the latest release and
  registry digest. Changes nothing; meant for the session-start ritual.
- scripts/update_stack.sh: resolve new digest -> stop -> back up surreal_data,
  notebook_data and docker-compose.yml -> repin -> start -> smoke test -> roll back data
  AND compose file if the smoke test fails. The data backup is the point: the app migrates
  the DB on startup and an older app cannot read a migrated DB, so a bad update would
  otherwise be a one-way door. tar runs inside a container because the data dirs are
  root-owned and the host user cannot restore over them.
- scripts/smoke_test.sh: checks what actually breaks here, not just "does it start".
  Every local adaptation leans on upstream internals and can break silently: the
  prompts/podcast directory mount masks the image's directory (a new template upstream
  would be invisible), config/content_core.yaml freezes content-core's defaults (because
  CCORE_CONFIG_PATH replaces rather than merges), and the env-var workarounds depend on
  current content-core/esperanto/podcast_creator behaviour. So it verifies those
  assumptions explicitly, plus API, TTS audio, STT and a real YouTube transcript.

Both scripts read the digest from the lfnovo/open_notebook line specifically. A naive
"first sha256 in the file" grep matches surrealdb (listed first) — caught while testing:
check_updates.sh falsely reported an update, and update_stack.sh would have rewritten the
database image instead of the application.

Verified: smoke test green against the current stack, and red (exit 1) when failures are
injected (dead TTS port, removed template variable). Backup/restore mechanics exercised
separately: 44 MB in 3.3s, restore yields 13 DB files and 91 notebook files. The update
path itself cannot be exercised end to end until a newer image exists.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-11 14:12:31 +02:00

79 lines
4.2 KiB
YAML

# Images sind bewusst auf einen Digest gepinnt, nicht auf einen beweglichen Tag.
#
# Warum: "v1-latest" ist ein wandernder Zeiger. Mit pull_policy: always zog jedes
# `docker compose up -d` unbemerkt eine neue Anwendung — inklusive Schema-Migration der
# SurrealDB, die sich nicht ohne Weiteres zurueckdrehen laesst. Gleichzeitig griff das
# nur bei `up -d`: nach einem Reboot startete `restart: always` weiter das ALTE Image.
# Also beides unzuverlaessig: mal heimlich neu, mal heimlich alt.
#
# Mit Digest-Pin laeuft immer exakt der Stand, der hier im Repo steht. Aktualisiert wird
# bewusst und mit Rueckfallpunkt:
# ./scripts/check_updates.sh -> gibt es ein neues Release? (aendert nichts)
# ./scripts/update_stack.sh -> Backup, Update, Rauchtest, bei Fehler Rollback
services:
surrealdb:
# surrealdb:v2 (Digest-Pin: auch ein Minor-Sprung kann das Storage-Format anfassen)
image: surrealdb/surrealdb:v2@sha256:d653f6c8a89e81f865ee31cd2f587c50f50ace922175e04150b1e385d2f86011
restart: always
pull_policy: missing
command: start --log info --user ${SURREAL_USER:-root} --pass ${SURREAL_PASSWORD:-root} rocksdb:/mydata/mydatabase.db
user: root # Required for bind mounts on Linux
environment:
- SURREAL_EXPERIMENTAL_GRAPHQL=true
ports:
- "127.0.0.1:8000:8000"
volumes:
- ./surreal_data:/mydata
open_notebook:
# Open Notebook v1.10.0 (Release vom 2026-06-18, Image vom selben Tag).
# Digest von scripts/update_stack.sh gepflegt — nicht von Hand aendern.
image: lfnovo/open_notebook:v1-latest@sha256:c8112fbd4b8fee7f2a20d3bdbea24e7d72267acfc990b803fd0ff4de30899b57
restart: always
pull_policy: missing
depends_on:
- surrealdb
ports:
- "127.0.0.1:8502:8502"
- "127.0.0.1:5055:5055"
environment:
- OPEN_NOTEBOOK_ENCRYPTION_KEY=${OPEN_NOTEBOOK_ENCRYPTION_KEY}
- SURREAL_URL=ws://surrealdb:8000/rpc
- SURREAL_USER=${SURREAL_USER:-root}
- SURREAL_PASSWORD=${SURREAL_PASSWORD:-root}
- SURREAL_NAMESPACE=open_notebook
- SURREAL_DATABASE=open_notebook
- OLLAMA_BASE_URL=http://host.docker.internal:11434
- OPENROUTER_API_KEY=${OPENROUTER_API_KEY}
# podcast_creator schickt sonst 5 Audio-Clips gleichzeitig an den TTS-Server.
# Die laufen dort echt parallel auf einer GPU -> Aktivierungsspeicher
# vervielfacht sich -> CUDA OOM (GPU 2 teilt sich TTS, STT und den
# chatterbox-MCP-Dienst). Der TTS-Server serialisiert intern zwar ohnehin,
# aber bei Batch 5 warten 4 Requests im Lock und laufen ins 300s-Timeout.
- TTS_BATCH_SIZE=1
# Kopfraum fuer lange Segmente; Default waere 300s pro Clip.
- ESPERANTO_TTS_TIMEOUT=600
# content-core sucht YouTube-Transkripte sonst nur in en/es/pt — deutsche
# Videos landen dann als LEERE Quelle im Notebook. Siehe config/content_core.yaml.
- CCORE_CONFIG_PATH=/app/config/content_core.yaml
# Audio-Quellen (Upload, Video-Tonspur): Open Notebook reicht sein STT-Modell
# (openai_compatible/faster-whisper-large-v3) an content-core weiter, aber
# content-core baut das Modell OHNE base_url (es gibt nur timeout mit). Esperanto
# weiss dann nicht, wo der lokale Whisper-Server steht, faellt auf OpenAI zurueck
# und scheitert mit "OpenAI API key not found". Diese Variable liefert die URL nach.
- OPENAI_COMPATIBLE_BASE_URL_STT=http://host.docker.internal:8902
extra_hosts:
- "host.docker.internal:host-gateway"
volumes:
- ./notebook_data:/app/data
# Podcast-Prompt-Vorlagen mit expliziter Sprachanweisung ({{ language }})
# und abgeschaltetem Thinking-Modus (/no_think), damit Podcasts in der
# eingestellten Sprache erzeugt werden (statt Englisch) und das JSON bei
# langen Segmenten nicht durch riesige Reasoning-Blöcke abgeschnitten wird.
# Verzeichnis-Mount (nicht Einzeldateien), damit Edits ohne Inode-Probleme
# nach einem Container-Neustart greifen. Der Ordner im Image enthält nur
# genau diese zwei Vorlagen.
- ./prompts/podcast:/app/prompts/podcast:ro
# content-core-Konfiguration (u.a. deutsche YouTube-Transkripte), per
# CCORE_CONFIG_PATH oben referenziert.
- ./config:/app/config:ro