# Changelog All notable changes to this project are documented here. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). ## [Unreleased] ### Added - Per-profile chat streaming and chat-request timeouts: `stream`, `read_timeout` and `connect_timeout` are now resolvable in `[default]`/`[model.]` (CLI `--stream`/`--read-timeout`/`--connect-timeout` still override). - Multimodal support: `mmproj` and `mmproj_offload` config keys (CLI `--mmproj`, `--no-mmproj-offload`) pass a vision projector to the server. The path resolves under `hf_home` like `model_path`, and a configured-but-missing projector now fails with exit code 3 — for `--change` before the running container is removed. - `docs/KI_TOOLS_PROFILES.md`: which local model suits which task, plus mmproj compatibility per base model. - `docs/EVAL_RUBRIC.md` and `scripts/eval_prompt_tests.py`: scoring rubric and an evaluator for the manual prompt-test archive that measures length compliance and **executes** the model-generated code against its own pytest suite. - `scripts/run_prompt_suite.sh`: runs a prompt domain against a running profile and stores the outputs under the archive's naming convention. - `[model.qwen35base]` profile: the non-abliterated Qwen3.6-35B-A3B (bartowski imatrix Q4_K_M) as a control for the abliteration comparison. Inherits sampling, `ctx_size` and cache types from `[default]`; only model, container, port and GPU differ, so it runs alongside the production server. ### Changed - The chat request now honors the configured `(connect, read)` timeout instead of a hard-coded 30 s; the default `read_timeout` is 600 s so slow reasoning models are no longer cut off mid-generation. - **Breaking:** `--print-effective-config` is now an action in the mutually exclusive action group instead of a flag, and it skips the `docker_available()` check. `--print-effective-config --start` is therefore an argparse error (exit 2) rather than a config dump followed by a real container start. ### Fixed - `--print-effective-config` combined with an action silently performed that action. Documented as a diagnostic command (README, installation guide, manual), `--print-effective-config --config … --start` printed the resolved config and then ran `do_start()`, replacing a running container with the `[default]` model. ## [0.1.0] - 2026-07-07 Initial release. Replaces the previous collection of shell scripts (`start-llm-server.sh`, `status-llm-server.sh`, `switch-llm.sh`, …) with a single Python package. ### Added - CLI to start, check, stop, and switch (`--change`) `llama.cpp` OpenAI-compatible Docker server instances, plus `--chat` and `--dry-run`. - Layered configuration (builtin defaults → `[default]` → `[model.]` → CLI overrides) with `[prompt.]` fallback system prompts and `${ENV}`/`~` expansion. - Security-hardened prompt input layer for `--system-*`/`--prompt-*`: file reads limited to regular UTF-8 files with a size cap and symlink rejection; URL reads are HTTPS-only with SSRF defenses (IP-literal and private/loopback/link-local/ metadata blocking, DNS pinning against rebinding, no-credentials URLs, redirect re-validation, content-type allowlist, streaming size limit). - Loopback-only port publishing by default; `--expose` + `--api-key` for LAN use. - Readiness check via a real chat completion (not just an open port) and an exclusive non-blocking file lock for `--change`. - MIT `LICENSE`, man page, and an opt-in end-to-end smoke test (`scripts/smoke.sh`). - Continuous integration on Forgejo: `pytest` (Python 3.10 and 3.12), `ruff`, and `mypy`. ### Changed - Pin the default `llama.cpp` Docker image by digest for reproducibility instead of the moving `:server-cuda` tag. - Model-size profile names carry a `b` suffix (`qwen35b`, `qwen27b`) so the digits read as parameter count, not Qwen version; documentation examples run against the `[default]` section so they work by copy & paste. ### Fixed - Override the image's baked-in Docker `HEALTHCHECK` (hardcoded to port 8080) so it targets the configured `container_port`, ending the false "unhealthy" status when the server runs on a different port. [Unreleased]: https://kitux.de/forgejo/dschlueter/llamacppctl/compare/v0.1.0...HEAD [0.1.0]: https://kitux.de/forgejo/dschlueter/llamacppctl/releases/tag/v0.1.0