my_voice_assistant_v3/BEDIENUNGSANLEITUNG.md

223 lines
6 KiB
Markdown
Raw Normal View History

# Bedienungsanleitung — Voice Assistant Gateway
Diese Anleitung führt Schritt für Schritt durch Installation, Start, Konfiguration
und Fehlerbehebung. Technische Hintergründe stehen im
[Architektur-Dokument](Docs/voice-assistant-architecture.md), eine kompakte
Übersicht im [README](README.md).
---
## 1. Voraussetzungen
- **Python 3.11 oder neuer** (`python3 --version`)
- Ein **OpenRouter-API-Key** — nur nötig, wenn ein Profil entfernte KI nutzt
(`hybrid`, `cloud`). Für rein lokalen Betrieb (`local-dev`) nicht erforderlich.
- Optional: Docker, falls im Container betrieben.
---
## 2. Installation
```bash
cd voice-assistant-scaffold
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install -e .[test]
```
Danach die zentrale Konfigurationsdatei anlegen:
```bash
cp config/voice-assistant.example.toml config/voice-assistant.toml
```
---
## 3. API-Key hinterlegen (für Cloud/Hybrid)
Der Schlüssel wird **aus der Umgebung** gelesen und gehört **nicht** in eine Datei.
Dauerhaft am besten in `~/.bashrc`:
```bash
echo 'export OPENROUTER_API_KEY=sk-or-v1-DEIN_KEY' >> ~/.bashrc
chmod 600 ~/.bashrc
source ~/.bashrc
```
Prüfen, ob er ankommt:
```bash
echo ${OPENROUTER_API_KEY:0:8} # zeigt nur den Anfang
```
> **Sicherheit:** Den Key niemals in `.env` oder `config/*.toml` schreiben. Wird ein
> Key versehentlich öffentlich, im OpenRouter-Dashboard löschen (= widerrufen) und
> neu erzeugen.
---
## 4. Betriebsart (Profil) wählen
Profile bestimmen, welche KI-Module genutzt werden:
| Profil | Bedeutung | Key nötig? |
|-------------|--------------------------------------------|------------|
| `local-dev` | alles lokal (eigene KI/Hardware) | nein |
| `hybrid` | STT/TTS über Cloud, Haupt-LLM lokal | ja |
| `cloud` | alles über OpenRouter (Standardbetrieb) | ja |
Profil **einmalig** für einen Start:
```bash
VA_PROFILE=cloud make run
```
Profil **dauerhaft** — in `.env` eintragen:
```
VA_PROFILE=cloud
```
> Hinweis: Stehen in `.env` noch `DEFAULT_STT_PROVIDER` / `DEFAULT_LLM_PROVIDER` /
> `DEFAULT_TTS_PROVIDER`, überschreiben diese das Profil. Für profilbasiertes
> Umschalten sollten sie auskommentiert sein.
---
## 5. Starten und Stoppen
```bash
make run
```
Standard-Adresse: `http://localhost:8080` (Port änderbar, siehe Abschnitt 8).
Beenden mit **Strg + C**.
Schnelltest in einem zweiten Terminal:
```bash
curl http://localhost:8080/health
# {"status":"ok"}
curl http://localhost:8080/api/config
# zeigt aktives Profil und die aufgelöste Standard-Route
```
---
## 6. Tägliche Bedienung — typische Aufgaben
### a) Text sprechen lassen (`/api/speak`)
```bash
curl -X POST http://localhost:8080/api/speak \
-H 'Content-Type: application/json' \
-d '{"text":"Guten Morgen, wie geht es Ihnen?"}' \
--output antwort.pcm
```
### b) Chatten (Text rein, gesprochene Antwort raus) (`/api/chat`)
Nur den Trace als JSON ansehen (ohne Audio):
```bash
curl -X POST "http://localhost:8080/api/chat?debug=true" \
-H 'Content-Type: application/json' \
-d '{"text":"Wie wird das Wetter morgen?"}'
```
Komfortabler mit dem mitgelieferten Client (spielt die Antwort ab):
```bash
python chat_client.py "Erzähl mir einen guten Morgen-Spruch"
```
> `chat_client.py` erwartet den Dienst auf Port **8003** — bei Bedarf im Skript
> `GATEWAY_URL` anpassen oder den Dienst mit `PORT=8003 make run` starten.
### c) Audio transkribieren (`/api/transcribe`)
```bash
curl -X POST http://localhost:8080/api/transcribe \
-F "file=@aufnahme.wav" -F "language=de"
```
### d) Gerät oder Provider einmalig umstellen (pro Aufruf)
```bash
curl -X POST http://localhost:8080/api/speak \
-H 'Content-Type: application/json' \
-d '{"text":"Test","tts_provider":"piper","output_endpoint":"loopback"}'
```
### e) Präferenzen für eine Session festlegen
```bash
# einmal setzen
curl -X POST http://localhost:8080/api/sessions/oma-anna/route \
-H 'Content-Type: application/json' \
-d '{"llm_provider":"openrouter","language":"de"}'
# danach mit dieser Session nutzen
curl -X POST "http://localhost:8080/api/chat?session_id=oma-anna&debug=true" \
-H 'Content-Type: application/json' -d '{"text":"Hallo!"}'
```
---
## 7. Verfügbare Geräte und Bausteine ansehen
```bash
curl http://localhost:8080/api/devices # Audio-Endpunkte mit Fähigkeiten
curl http://localhost:8080/api/config # Profil, Route, Provider, Endpunkte
```
---
## 8. Port ändern
```bash
PORT=8003 make run # einmalig
sed -i 's/^PORT=.*/PORT=8003/' .env # dauerhaft
```
---
## 9. Mit Docker betreiben
```bash
export OPENROUTER_API_KEY=sk-or-v1-...
docker compose up --build
```
Der Key wird aus der Shell in den Container durchgereicht; fehlt er, bricht der
Start mit klarer Meldung ab.
---
## 10. Fehlerbehebung
| Symptom | Ursache | Lösung |
|---|---|---|
| `OPENROUTER_API_KEY is empty` / 401 | Key nicht in der Umgebung | `export OPENROUTER_API_KEY=…`, neues Terminal / `source ~/.bashrc` |
| HTTP **422** „Unbekannter …-Provider/Endpunkt" | Tippfehler in `*_provider` / `*_endpoint` | gültige Werte via `GET /api/config` prüfen |
| `VA_PROFILE` wirkt nicht | `DEFAULT_*_PROVIDER` in `.env` überschreibt es | diese Zeilen in `.env` auskommentieren |
| LLM-Timeout / Connection refused (lokal) | lokaler LLM-Server (Port 11434) läuft nicht | LLM-Server starten oder Profil `cloud` wählen |
| `Address already in use` | Port belegt | anderen `PORT` setzen (Abschnitt 8) |
| `chat_client.py` bekommt keine Antwort | Client nutzt Port 8003 | Dienst mit `PORT=8003` starten oder `GATEWAY_URL` anpassen |
| Profil greift nicht / Standardwerte | `config/voice-assistant.toml` fehlt | Datei aus `*.example.toml` kopieren (Abschnitt 2) |
Logs erscheinen im Terminal, in dem `make run` läuft. Für mehr Details
`LOG_LEVEL=debug` in `.env` setzen.
---
## 11. Tests ausführen
```bash
make test
```
Alle Tests sollten grün sein. Schlägt etwas fehl, gibt die Ausgabe den genauen
Testnamen und die Ursache an.