feat(scripts): add prompt-test evaluator and suite runner
eval_prompt_tests.py measures the objective half of docs/EVAL_RUBRIC.md over the manual test archive: word count against the target stated in each prompt, truncation suspicion, and — for the coding domain — it writes the generated module and tests to a temp dir and actually runs pytest against them. Deriving the module's filename is the delicate part: a name taken from a test's `import sqlite3` would shadow the stdlib and fail the run for a reason the model is not responsible for. Names now come from the last *.py mention before the block, then from `from X import`, and anything in sys.stdlib_module_names is rejected. A module that no test imports is reported as such, since that is a finding about test quality rather than a guess the runner got wrong. run_prompt_suite.sh drives one prompt domain against a running profile and stores the outputs under the archive's naming convention. Both scripts join the ruff gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
3e3dd1c150
commit
6f2f8aff6c
4 changed files with 556 additions and 2 deletions
|
|
@ -27,7 +27,7 @@ run() {
|
|||
fi
|
||||
}
|
||||
|
||||
run "ruff (lint)" "${BIN}ruff" check src/ tests/ build_archive.py
|
||||
run "ruff (lint)" "${BIN}ruff" check src/ tests/ build_archive.py scripts/eval_prompt_tests.py
|
||||
run "mypy (types)" "${BIN}mypy"
|
||||
# Use `python -m pytest` (not the pytest console script) so the repo root is on
|
||||
# sys.path — the test modules import `from tests.test_docker_ops import ...`.
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue