feat(scripts): add prompt-test evaluator and suite runner

eval_prompt_tests.py measures the objective half of docs/EVAL_RUBRIC.md over
the manual test archive: word count against the target stated in each prompt,
truncation suspicion, and — for the coding domain — it writes the generated
module and tests to a temp dir and actually runs pytest against them.

Deriving the module's filename is the delicate part: a name taken from a test's
`import sqlite3` would shadow the stdlib and fail the run for a reason the model
is not responsible for. Names now come from the last *.py mention before the
block, then from `from X import`, and anything in sys.stdlib_module_names is
rejected. A module that no test imports is reported as such, since that is a
finding about test quality rather than a guess the runner got wrong.

run_prompt_suite.sh drives one prompt domain against a running profile and
stores the outputs under the archive's naming convention. Both scripts join the
ruff gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Dieter Schlüter 2026-07-10 16:25:26 +02:00
commit 6f2f8aff6c
4 changed files with 556 additions and 2 deletions

View file

@ -27,7 +27,7 @@ run() {
fi
}
run "ruff (lint)" "${BIN}ruff" check src/ tests/ build_archive.py
run "ruff (lint)" "${BIN}ruff" check src/ tests/ build_archive.py scripts/eval_prompt_tests.py
run "mypy (types)" "${BIN}mypy"
# Use `python -m pytest` (not the pytest console script) so the repo root is on
# sys.path — the test modules import `from tests.test_docker_ops import ...`.