docs: add model profiles, evaluation rubric, and CLAUDE.md
KI_TOOLS_PROFILES.md records what each local GGUF is actually for, based on the model cards: ornith is a purpose-built agentic coding model, Carnice targets agent runtimes, Qwopus is reasoning plus vision, and the two HauhauCS models are abliterated — an axis about refusals, not literary quality. It also maps which models the local mmproj.gguf fits (the Qwen3.6-35B-A3B based ones, not ornith). Capability and benchmark claims are attributed to their authors, not asserted. EVAL_RUBRIC.md separates what a script can measure from what needs reading, and warns that a passing test suite only proves the code satisfies its own tests. CHANGELOG covers the mmproj feature, the --print-effective-config fix, and the new tooling. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
6f2f8aff6c
commit
bee6a71de0
4 changed files with 406 additions and 0 deletions
25
CHANGELOG.md
25
CHANGELOG.md
|
|
@ -11,11 +11,36 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
|
|||
- Per-profile chat streaming and chat-request timeouts: `stream`, `read_timeout`
|
||||
and `connect_timeout` are now resolvable in `[default]`/`[model.<profile>]`
|
||||
(CLI `--stream`/`--read-timeout`/`--connect-timeout` still override).
|
||||
- Multimodal support: `mmproj` and `mmproj_offload` config keys (CLI `--mmproj`,
|
||||
`--no-mmproj-offload`) pass a vision projector to the server. The path resolves
|
||||
under `hf_home` like `model_path`, and a configured-but-missing projector now
|
||||
fails with exit code 3 — for `--change` before the running container is removed.
|
||||
- `docs/KI_TOOLS_PROFILES.md`: which local model suits which task, plus mmproj
|
||||
compatibility per base model.
|
||||
- `docs/EVAL_RUBRIC.md` and `scripts/eval_prompt_tests.py`: scoring rubric and an
|
||||
evaluator for the manual prompt-test archive that measures length compliance and
|
||||
**executes** the model-generated code against its own pytest suite.
|
||||
- `scripts/run_prompt_suite.sh`: runs a prompt domain against a running profile and
|
||||
stores the outputs under the archive's naming convention.
|
||||
- `[model.qwen35base]` profile: the non-abliterated Qwen3.6-35B-A3B (bartowski
|
||||
imatrix Q4_K_M) as a control for the abliteration comparison. Inherits sampling,
|
||||
`ctx_size` and cache types from `[default]`; only model, container, port and GPU
|
||||
differ, so it runs alongside the production server.
|
||||
|
||||
### Changed
|
||||
- The chat request now honors the configured `(connect, read)` timeout instead
|
||||
of a hard-coded 30 s; the default `read_timeout` is 600 s so slow reasoning
|
||||
models are no longer cut off mid-generation.
|
||||
- **Breaking:** `--print-effective-config` is now an action in the mutually
|
||||
exclusive action group instead of a flag, and it skips the `docker_available()`
|
||||
check. `--print-effective-config --start` is therefore an argparse error
|
||||
(exit 2) rather than a config dump followed by a real container start.
|
||||
|
||||
### Fixed
|
||||
- `--print-effective-config` combined with an action silently performed that
|
||||
action. Documented as a diagnostic command (README, installation guide, manual),
|
||||
`--print-effective-config --config … --start` printed the resolved config and
|
||||
then ran `do_start()`, replacing a running container with the `[default]` model.
|
||||
|
||||
## [0.1.0] - 2026-07-07
|
||||
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue