llamacppctl/docs
dschlueter b0c215abbb feat(scripts): detect how a model evades, not just whether
classify_refusal() separates refusal (head), preamble, appended disclaimer
(tail) and moralising insert (inline). Only the first three are meaningful, and
only where a prompt forbids them: a preamble on a coding answer is normal, so it
no longer counts as evasion.

Two sources of false positives had to be removed first, both found by checking
the detector's hits against the archive rather than trusting them:

  * llamacppctl's own truncation warning on stderr had leaked into an archived
    output and was read as a model disclaimer. It also inflated that run's word
    count, so count_words() strips it too.
  * An AI character saying "Bitte beachten Sie:" inside a dystopian story is
    plot, not distancing. Quoted speech is removed before markers are matched.

run_prompt_suite.sh gains a CASES filter so a single prompt can be re-run.

Measured result, recorded in KI_TOOLS_PROFILES.md: across seven prose prompts
neither the abliterated model nor the aligned base model evaded once. On
literary prose the abliteration buys nothing measurable.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-10 17:10:39 +02:00
..
BEDIENUNGSANLEITUNG.md fix(cli): stop --print-effective-config from starting a container 2026-07-10 16:25:07 +02:00
EVAL_RUBRIC.md feat(scripts): detect how a model evades, not just whether 2026-07-10 17:10:39 +02:00
INSTALL_FROM_ARCHIVE.md fix(cli): stop --print-effective-config from starting a container 2026-07-10 16:25:07 +02:00
KI_TOOLS_PROFILES.md feat(scripts): detect how a model evades, not just whether 2026-07-10 17:10:39 +02:00
SECURITY_AND_OPERATIONS.md fix(cli): stop --print-effective-config from starting a container 2026-07-10 16:25:07 +02:00