feat(scripts): add prompt-test evaluator and suite runner

eval_prompt_tests.py measures the objective half of docs/EVAL_RUBRIC.md over
the manual test archive: word count against the target stated in each prompt,
truncation suspicion, and — for the coding domain — it writes the generated
module and tests to a temp dir and actually runs pytest against them.

Deriving the module's filename is the delicate part: a name taken from a test's
`import sqlite3` would shadow the stdlib and fail the run for a reason the model
is not responsible for. Names now come from the last *.py mention before the
block, then from `from X import`, and anything in sys.stdlib_module_names is
rejected. A module that no test imports is reported as such, since that is a
finding about test quality rather than a guess the runner got wrong.

run_prompt_suite.sh drives one prompt domain against a running profile and
stores the outputs under the archive's naming convention. Both scripts join the
ruff gate.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
Dieter Schlüter 2026-07-10 16:25:26 +02:00
commit 6f2f8aff6c
4 changed files with 556 additions and 2 deletions

View file

@ -19,7 +19,7 @@ jobs:
run: pip install -e ".[dev]"
- name: Ruff (lint)
run: ruff check src/ tests/ build_archive.py
run: ruff check src/ tests/ build_archive.py scripts/eval_prompt_tests.py
- name: Mypy (type check)
run: mypy