Python-Package zur unabhängigen Überwachung statischer Websites gegen SEO-Spam und unbefugte Manipulationen. Läuft außerhalb des Hosters. Kernfunktionen: - Reiner Python-Crawler (bs4+lxml), kein wget - Baseline-Management mit manuellem Freigabe-Workflow (niemals automatisch) - Text-/Link-/Meta-Diff mit Risiko-Scoring (grün/gelb/rot) - Hidden-Content-Erkennung (CSS inline, noscript, Kommentar-Links) - Whitelist für externe Domains und interne Pfade - E-Mail + Webhook Alarmierung - CLI: init, crawl, check, scan, approve, report, status - 79 Unit-Tests (pytest), alle grün Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
83 lines
4.2 KiB
Markdown
83 lines
4.2 KiB
Markdown
# CLAUDE.md
|
|
|
|
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
|
|
|
|
## Commands
|
|
|
|
```bash
|
|
# Install dependencies
|
|
pip install -r requirements.txt
|
|
|
|
# Run all tests
|
|
pytest tests/ -v
|
|
|
|
# Run a single test file
|
|
pytest tests/test_extractor.py -v
|
|
|
|
# Run a single test
|
|
pytest tests/test_differ.py::TestScoreDiff::test_green_for_no_changes -v
|
|
|
|
# CLI usage (always run from project root where config.yaml lives)
|
|
python -m scanner status
|
|
python -m scanner init
|
|
python -m scanner scan
|
|
python -m scanner approve --all --note "Reason"
|
|
python -m scanner approve --url https://bredelar.info/page/
|
|
python -m scanner report
|
|
python -m scanner --verbose scan
|
|
```
|
|
|
|
## Architecture
|
|
|
|
The scanner is a Python package (`scanner/`) invoked via `python -m scanner`. All subcommands share one data flow:
|
|
|
|
```
|
|
Crawler → Extractor → BaselineManager (snapshot)
|
|
↓
|
|
differ.compare_snapshots
|
|
↓
|
|
differ.score_diff → alerter.send_alert
|
|
↓
|
|
__main__._write_report (JSON + MD)
|
|
```
|
|
|
|
**Key design constraint**: The baseline is immutable except through an explicit `approve` call. `compare_snapshots` and `score_diff` are pure functions — they never write anything.
|
|
|
|
### Module responsibilities
|
|
|
|
- **`config.py`** — loads `config.yaml`, deep-merges onto `DEFAULT_CONFIG`. All paths (data_dir, reports_dir, logs_dir, config_dir) are resolved to absolute paths by `resolve_paths()` before being passed to other modules.
|
|
- **`crawler.py`** — breadth-first crawl using `requests`. `normalize_url()` is the canonical URL form used everywhere (strips fragment, lowercases host, removes noise query params, sorts remaining params). `should_crawl()` filters out binary assets and tracking URLs.
|
|
- **`extractor.py`** — parses HTML with bs4+lxml. Returns a single dict per page with keys: `text`, `links`, `metadata`, `hidden_content`, `inline_scripts`, `jsonld`, `comment_links`. The `links` dict has typed sub-lists (`a`, `link_rel`, `script`, `img`, `iframe`, `form`, `meta_refresh`, `canonical`).
|
|
- **`baseline.py`** — `BaselineManager` handles two separate stores: `data/snapshots/YYYYMMDD_HHMMSS/` (raw crawl output) and `data/baseline/` (approved reference). Each page is stored as `<url_key(url)>.json`. `approve_*` methods are the only path into the baseline.
|
|
- **`differ.py`** — `compare_snapshots(baseline_pages, current_pages, cfg)` returns a flat diff dict. `score_diff(diff, cfg)` maps it to `{score, level, reasons, exit_code}`. Thresholds: yellow ≥ 20, red ≥ 60 (configurable). A single new external domain = 50 pts (yellow); combined with hidden content = 90 pts (red).
|
|
- **`checker.py`** — stateless helper functions for whitelist enforcement, security-header auditing, canonical hijack detection. Called from `__main__.cmd_check`, not from `differ`.
|
|
- **`alerter.py`** — sends alert only when `level >= min_level`. SMTP password is read from the env var named in `smtp_password_env` (default: `SCANNER_SMTP_PASSWORD`).
|
|
- **`__main__.py`** — wires everything together. `cmd_scan` = `cmd_crawl` (crawl + save snapshot) + `cmd_check` (load baseline + snapshot, diff, score, report, alert). Exit codes: 0=green, 1=yellow, 2=red.
|
|
|
|
### Data layout
|
|
|
|
```
|
|
data/baseline/manifest.json # approved_at, approved_by, note, urls[]
|
|
data/baseline/pages/<16hex>.json # one file per approved URL
|
|
data/snapshots/YYYYMMDD_HHMMSS/ # one dir per crawl run
|
|
reports/YYYYMMDD_HHMMSS/report.{json,md}
|
|
config/allowed_external.yaml # set of permitted external domains
|
|
config/allowed_paths.yaml # set of permitted internal paths (optional)
|
|
config/ignore_rules.yaml # normalization selectors + regex patterns
|
|
```
|
|
|
|
### Scoring weights (defaults, all configurable in `config.yaml`)
|
|
|
|
| Event | Points |
|
|
|---|---|
|
|
| New external domain | 50 |
|
|
| Hidden content (CSS-hidden element with links/text) | 40 |
|
|
| Meta-Refresh | 40 |
|
|
| Unexpected canonical | 35 |
|
|
| Unexpected JSON-LD @type | 35 |
|
|
| New internal URL | 30 |
|
|
| New suspicious inline script | 30 |
|
|
| Comment-embedded link | 25 |
|
|
| Large text addition (>200 chars) | 20 |
|
|
| Missing internal URL | 20 |
|
|
| Missing security header | 5 |
|