feat(scripts): add prompt-test evaluator and suite runner
eval_prompt_tests.py measures the objective half of docs/EVAL_RUBRIC.md over the manual test archive: word count against the target stated in each prompt, truncation suspicion, and — for the coding domain — it writes the generated module and tests to a temp dir and actually runs pytest against them. Deriving the module's filename is the delicate part: a name taken from a test's `import sqlite3` would shadow the stdlib and fail the run for a reason the model is not responsible for. Names now come from the last *.py mention before the block, then from `from X import`, and anything in sys.stdlib_module_names is rejected. A module that no test imports is reported as such, since that is a finding about test quality rather than a guess the runner got wrong. run_prompt_suite.sh drives one prompt domain against a running profile and stores the outputs under the archive's naming convention. Both scripts join the ruff gate. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
This commit is contained in:
parent
3e3dd1c150
commit
6f2f8aff6c
4 changed files with 556 additions and 2 deletions
|
|
@ -19,7 +19,7 @@ jobs:
|
|||
run: pip install -e ".[dev]"
|
||||
|
||||
- name: Ruff (lint)
|
||||
run: ruff check src/ tests/ build_archive.py
|
||||
run: ruff check src/ tests/ build_archive.py scripts/eval_prompt_tests.py
|
||||
|
||||
- name: Mypy (type check)
|
||||
run: mypy
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue