Keep only README, LICENSE, and CHANGELOG as prose at the repo root
(tooling/convention) and move the remaining manuals under docs/ next to
SECURITY_AND_OPERATIONS.md.
- git mv the two files into docs/ (history preserved).
- Update all cross-references: README doc-index links, the internal
SECURITY_AND_OPERATIONS link and the §12 pointer list in
BEDIENUNGSANLEITUNG, and the repo inventory in SECURITY_AND_OPERATIONS §1.
- build_archive.py: point REQUIRED_FILES at docs/ and drop the now-redundant
INCLUDE_FILES entries (the docs/ dir is included wholesale).
- build_archive.py: drop the obsolete requirements.txt<->pyproject dependency
mirror check, which broke once requirements.txt was reduced to `.`
(pyproject is the single source of truth). Verified by building and
re-opening the archive.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Improve orientation across the doc set without duplicating content:
- README: add a "Dokumentation" table mapping each doc (README,
BEDIENUNGSANLEITUNG, INSTALL_FROM_ARCHIVE, SECURITY_AND_OPERATIONS, man
page, CHANGELOG) to who it is for and what it covers.
- SECURITY_AND_OPERATIONS §1: extend the module map with a repo-level
inventory (tests, config template, packaging, build/CI/hook scripts,
docs) so "which file does what" is answered beyond the src/ modules.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Clean up documentation redundancy:
- Delete docs/Archiv_fertig_-_Check_und_Installation.md: an AI hand-off
transcript (first-person, stale test count, malformed markdown) whose
useful content is already in INSTALL_FROM_ARCHIVE.md.
- Delete docs/How_to_use.md after moving its one unique asset -- the
old-script -> new-command migration table -- into the README; drop it
from the build_archive manifest.
- Fix INSTALL_FROM_ARCHIVE.md drift: list LICENSE/CHANGELOG.md in the
archive allowlist and use `python -m pytest`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Forgejo Actions is not enabled on the instance, so run the same checks
locally before pushing:
- scripts/check.sh runs ruff, mypy, and pytest (mirrors the CI workflow).
- .githooks/pre-push invokes it; enable per clone with
`git config core.hooksPath .githooks`. Bypass with `git push --no-verify`.
- Fix the pytest invocation in the CI workflow (and document it): use
`python -m pytest` so the repo root is on sys.path, otherwise the test
modules fail to `import tests.*`.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
- Add CHANGELOG.md (Keep a Changelog format) with a 0.1.0 entry summarizing
the initial release plus the recent healthcheck fix, image pin, and
profile-naming changes.
- Add LICENSE and CHANGELOG.md to the build_archive.py manifest so the
distributable tarball actually contains the license and history.
- Expand the README license section to the full MIT + copyright line.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
requirements.txt / requirements-dev.txt used to duplicate the dependency
pins from pyproject.toml by hand, which can silently drift. Point them at
the package itself (`.` and `-e .[dev]`) so the pins live in exactly one
place; `pip install -r requirements*.txt` keeps working.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Back the shipped py.typed promise with an enforced type check and a linter:
- Add ruff + mypy (+ types-requests) to the dev extras and dev requirements,
with [tool.ruff]/[tool.mypy] config in pyproject.toml (mypy checks the
package, not the tests).
- Add a lint job to the Forgejo workflow running ruff check + mypy.
- Fix the issues this surfaced: type FileLock.fd as TextIO | None, add a
targeted type: ignore for the intentional socket.getaddrinfo monkeypatch,
and drop an unused import.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The :server-cuda tag is a moving target, so a fresh pull could silently
change server behaviour (flags, the baked-in healthcheck, ...). Pin the
default image to the current digest for reproducibility; overriding `image`
in the config or via --image still works. A comment documents how to bump
the pin.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Public-repo hygiene now that the project is hosted on Forgejo:
- LICENSE: add the MIT text that pyproject.toml already declares.
- .forgejo/workflows/ci.yml: run the pytest suite on push/PR against
Python 3.10 and 3.12, so regressions (e.g. the healthcheck port bug)
get caught automatically.
- Untrack llama.cpp.config and gitignore it: it is a machine-specific
runtime config that may later hold an api_key. The tracked template
remains llama.cpp.config.example (copy it to get started).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Rename the size-numbered profiles so the digits read as parameter count,
not Qwen version: qwen35 -> qwen35b (llama.cpp.config.example, build_archive
smoke test, SECURITY_AND_OPERATIONS.md) and qwen27 -> qwen27b
(llama.cpp.config).
Drop the misleading `--profile qwen35` from the README/man/install examples:
that profile only exists in the .example file, so pasted commands failed
against the real config. Primary examples now omit --profile (using the
[default] section, which always resolves), with one example plus a note
showing how to select a real [model.<name>] profile.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The ghcr.io llama.cpp image bakes in a HEALTHCHECK that curls port 8080
(the llama.cpp default). When the server runs on a different --port (here
8000), that check always fails and Docker reports the container as
"unhealthy" even though it serves fine. Override the healthcheck in the
docker run command to target the configured container_port/health_endpoint,
with a 300s start-period so large-context model loads don't flap.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Replace the stale example profiles (qwen35->8002, deepseek->8003, pointing at
non-existent paths) with profiles for the actually installed models. They now
override ONLY model_path and inherit host_port=8001, container_name and alias
from [default] -> one model at a time on the standard port; --start/--change
swaps it. Distinct port/container is only needed for concurrent operation.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
llama.cpp sends the streaming response as text/event-stream WITHOUT a charset;
requests then does not decode as UTF-8, so iter_lines(decode_unicode=True)
mangled multibyte characters (German Umlaute) into double-encoded garbage
(e.g. "schön" -> "schön"). The non-streaming path via resp.json() was fine.
Set resp.encoding = "utf-8" before iter_lines. Verified live: streamed bytes for
"schön" are now c3 b6 (correct UTF-8), file detected as UTF-8.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Vollständige deutschsprachige Bedienungsanleitung (Installation, Konfiguration,
alle fünf Aktionen, Prompt-Quellen, Antwort-/Reasoning-Steuerung, Netzwerk/
Sicherheit, GPU/Kontext, Exit-Codes, Fehlersuche, Tests) sowie eine Anleitung
zur Installation aus dem .tar.gz-Archiv.
build_archive.py: beide Anleitungen in INCLUDE_FILES + REQUIRED_FILES (werden
gepackt und verifiziert).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Die docs/ waren gegenüber den neuen Features veraltet.
- SECURITY_AND_OPERATIONS.md: Netzwerk-Exposition (Loopback-Default,
--expose/--api-key inkl. 401-Verhalten der Endpunkte), DNS-Pinning,
--start/--change unter Lock + --force, --change validate-before-remove,
Chat-Parameter (max_tokens/chat_temperature/--stream), ${ENV}-Expansion,
--check-Exit-Codes, smoke.sh, ChatReply; Fehler-Tabelle erweitert
(401, Trunkierung, Docker-health vs. HTTP-OK).
- How_to_use.md: Abschnitt zu den neuen Optionen.
- Archiv-Report: Testzahl (130) und Dateiliste aktualisiert.
- build_archive.py: scripts/ in INCLUDE_DIRS (sonst fehlt smoke.sh im Archiv),
test_actions.py/How_to_use.md/smoke.sh in REQUIRED_FILES aufgenommen.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
llama.cpp exempts /health and /v1/models from --api-key; only
/v1/chat/completions is protected. The smoke test probed /v1/models and thus
saw 200 without a key. Probe /v1/chat/completions instead (no key -> 401,
key -> 200). Verified live end-to-end: 9 PASS, 0 FAIL.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Sieben Verbesserungen; die Dateien überschneiden sich thematisch, daher ein
Commit (jeder Commit bleibt grün: 115 Tests).
- #1 --check ist scriptbar: Exit 0 wenn Container läuft und erreichbar,
sonst 5 (check_exit_code / CheckResult).
- #2 --force implementiert: Bypass eines belegten Locks mit Warnung
(_container_lock) und stop_container(force=…) schluckt Inkonsistenzen.
- #3 stille Trunkierung behoben: chat_completion_text liefert ChatReply
(content + finish_reason); bei finish_reason=length Hinweis auf stderr,
--chat gibt Exit 1 bei leerem Content zurück.
- #4 keine vermeidbare Downtime: --change validiert den Modellpfad VOR dem
Entfernen des laufenden Containers.
- #5 Netzwerk dicht: Port-Publish standardmäßig nur auf 127.0.0.1
(--expose/expose für alle Interfaces), optionaler --api-key/api_key
(Server --api-key + Bearer-Token auf allen Requests).
- #6 --stream: Chat-Reply token-weise via SSE auf stdout (stream_chat).
- #7 tests/test_actions.py: Orchestrierungs-Ebene (dry-run-Nebenwirkungen,
Lock, validate-before-remove, chat/stream/exit-codes).
Doku aktualisiert (Manpage, README, llama.cpp.config.example).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Steuert einen llama.cpp-Server als Docker-Container: --start/--check/--stop/
--change/--chat, INI-Konfiguration (builtin defaults -> [default] ->
[model.<profile>] -> CLI), SSRF-gehärtete Prompt-Eingabe (Datei/HTTPS-URL),
File-Locking für --start/--change und ein OpenAI-kompatibler HTTP-Layer.
Enthält u. a.:
- Env-Var-Expansion in hf_home (hf_home = ${HF_HOME})
- konfigurierbares Chat-Antwortbudget (max_tokens/chat_temperature,
CLI: --max-tokens/--chat-temp); temperature defer an Server-Default
- DNS-Pinning gegen DNS-Rebinding bei URL-Quellen
- dry-run als nebenwirkungsfreie Vorschau (kein Lock/Removal/Modell-Check)
- 98 Tests (pytest)
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>