Guard against stopping the stack while a podcast is running

Stopping the stack kills an in-flight podcast job, but its status stays "running"
in the database forever: it never finishes, it blocks prune_podcast_data.py (which
deliberately refuses to run while a job is active), and it cannot even be retried —
/retry only accepts episodes in state "failed".

I walked into this myself: I ran `docker compose down` for the non-root switch
without checking for running jobs, and killed a podcast that had already produced
37 clips.

update_stack.sh now refuses to start when a podcast job is running, and says why.
Verified against a genuinely running job: it aborts before touching the stack.

BEDIENUNGSANLEITUNG documents how to recover an existing zombie. The status does not
live on the episode but on the linked `command` record (the episode's own job_status
is null), so the fix is to set that record to 'failed' via SurrealDB, after which the
normal retry works.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Dieter Schlüter 2026-07-11 15:29:14 +02:00
commit 5b389fae58
2 changed files with 54 additions and 0 deletions

View file

@ -543,6 +543,42 @@ for t in YouTubeTranscriptApi().list('VIDEO_ID'): print(t.language_code, t.is_ge
Kommt hier nichts zurück, hat das Video keine Untertitel — dann hilft nur, das Video separat Kommt hier nichts zurück, hat das Video keine Untertitel — dann hilft nur, das Video separat
herunterzuladen und über die STT-Funktion (Audio-Upload) zu transkribieren. herunterzuladen und über die STT-Funktion (Audio-Upload) zu transkribieren.
### Podcast hängt ewig auf „running" (Karteileiche nach Container-Neustart)
Wird der Stack gestoppt (`docker compose down`, Update, Reboot), während ein Podcast läuft, ist der
Job weg — aber in der Datenbank steht er **weiterhin auf `running`**. Er wird nie fertig, blockiert
`prune_podcast_data.py` (das bricht bei laufenden Jobs absichtlich ab) und lässt sich auch nicht
wiederholen: `/retry` akzeptiert nur Episoden im Status `failed`.
Erkennen: keine neuen Clips im Episodenordner, GPU 2 im Leerlauf (`nvidia-smi`), nichts im Log
(`docker compose logs --since 5m open_notebook`).
Der Status hängt nicht an der Episode, sondern am verknüpften `command`-Datensatz. Ihn auf
`failed` setzen — danach funktioniert der normale Retry-Knopf:
```bash
set -a; . ./.env; set +a
# 1. Command-ID der Episode holen
curl -s -X POST http://127.0.0.1:8000/sql -u "$SURREAL_USER:$SURREAL_PASSWORD" \
-H "surreal-ns: open_notebook" -H "surreal-db: open_notebook" -H "Accept: application/json" \
-d "SELECT command FROM episode:<ID>;"
# 2. Auf failed setzen
curl -s -X POST http://127.0.0.1:8000/sql -u "$SURREAL_USER:$SURREAL_PASSWORD" \
-H "surreal-ns: open_notebook" -H "surreal-db: open_notebook" -H "Accept: application/json" \
-d "UPDATE command:<COMMAND-ID> SET status = 'failed', error_message = 'Worker beim Neustart beendet';"
# 3. Wiederholen
curl -X POST http://127.0.0.1:5055/api/podcasts/episodes/<ID>/retry
```
**Vorbeugen:** vor `docker compose down` oder `./scripts/update_stack.sh` kurz prüfen, ob ein
Podcast läuft — die Episodenliste in der UI zeigt es, oder:
```bash
curl -s http://127.0.0.1:5055/api/podcasts/episodes | python3 -c "
import json,sys
print([e['name'] for e in json.load(sys.stdin) if e['job_status']=='running'] or 'kein Job läuft')"
```
### Podcast schlägt beim Vertonen fehl („Failed to generate speech") ### Podcast schlägt beim Vertonen fehl („Failed to generate speech")
Text und Transkript sind fertig, erst die Sprachausgabe scheitert. Die genaue Fehlermeldung Text und Transkript sind fertig, erst die Sprachausgabe scheitert. Die genaue Fehlermeldung

View file

@ -46,6 +46,24 @@ if [ "$DRY_RUN" = 1 ]; then
exit 0 exit 0
fi fi
# Ein laufender Podcast ueberlebt das Stoppen nicht: der Job ist weg, sein Status bleibt
# aber auf "running" stehen. Er wird nie fertig, blockiert prune_podcast_data.py und laesst
# sich nicht per /retry wiederholen (das will "failed"). Also vorher fragen, nicht nachher
# aufraeumen.
running_jobs="$(curl -sf -m 10 http://127.0.0.1:5055/api/podcasts/episodes 2>/dev/null |
python3 -c "
import json,sys
try: eps = json.load(sys.stdin)
except Exception: sys.exit(0)
print('; '.join(e['name'] for e in eps if e.get('job_status') in ('running','pending','submitted')))" 2>/dev/null)"
if [ -n "${running_jobs// /}" ]; then
echo
echo "ABBRUCH: Es laeuft gerade ein Podcast-Job: $running_jobs"
echo " Das Stoppen wuerde ihn abschiessen und als Karteileiche zuruecklassen."
echo " Bitte warten, bis er fertig ist, und danach erneut starten."
exit 1
fi
ts="$(date +%Y%m%d-%H%M%S)" ts="$(date +%Y%m%d-%H%M%S)"
backup="$BACKUP_DIR/$ts" backup="$BACKUP_DIR/$ts"
mkdir -p "$backup" mkdir -p "$backup"