integrity_scanner_fuer_stat.../CLAUDE.md
Your Name 68243a97b9 fix: Bugfixes, toten Code entfernt, Cache-Normalisierung & Session
Echte Bugs:
- extractor: JSON-LD-Crash bei Nicht-Dict-Items (isinstance-Guard)
- baseline/__main__: Crawl-Fehler ins Snapshot-Manifest -> erscheinen im Report
- __main__: Whitelist-/Header-/Webshell-Checks nur noch auf status==200
- crawler: Noise-Param-Regex auf Voll-Key (view=/value= nicht mehr verworfen)
- differ/__main__: unerwartetes JSON-LD via eigenem Kanal, auch auf
  unveraenderten Seiten erkannt, kein Re-Alert auf bereits genehmigte Typen

Aufgeraeumt:
- checker: check_internal_paths, find_broken_pages, canonical_is_hijacked,
  check_jsonld_types entfernt (nicht verdrahtet)
- allowed_paths.yaml-Logik und tote Scoring-Keys entfernt
- tote Imports entfernt

Neuer aktiver Schutz:
- Webshell-/Backdoor-Dateinamen-Check verdrahtet (suspicious_filename: 60).
  Regex wortgrenzen-verankert gegen Fehlalarm auf PSEMailerAntispam.js

Effizienz/Struktur:
- crawler nutzt requests.Session (keep-alive)
- Cache-Buster-Normalisierung an einer Stelle (Extraktion) -> stabile
  Snapshots, differ wieder reiner Set-Diff

Tests: 111 gruen (neu: test_checker.py + Regressionstests)
CLAUDE.md aktualisiert

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-12 03:20:32 +02:00

5.1 KiB

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

Commands

# Install dependencies
pip install -r requirements.txt

# Run all tests
pytest tests/ -v

# Run a single test file
pytest tests/test_extractor.py -v

# Run a single test
pytest tests/test_differ.py::TestScoreDiff::test_green_for_no_changes -v

# CLI usage (always run from project root where config.yaml lives)
python -m scanner status
python -m scanner init
python -m scanner scan
python -m scanner approve --all --note "Reason"
python -m scanner approve --url https://bredelar.info/page/
python -m scanner report
python -m scanner --verbose scan

Architecture

The scanner is a Python package (scanner/) invoked via python -m scanner. All subcommands share one data flow:

Crawler → Extractor → BaselineManager (snapshot)
                              ↓
                    differ.compare_snapshots
                              ↓
                    differ.score_diff  →  alerter.send_alert
                              ↓
                    __main__._write_report (JSON + MD)

Key design constraint: The baseline is immutable except through an explicit approve call. compare_snapshots and score_diff are pure functions — they never write anything.

Module responsibilities

  • config.py — loads config.yaml, deep-merges onto DEFAULT_CONFIG. All paths (data_dir, reports_dir, logs_dir, config_dir) are resolved to absolute paths by resolve_paths() before being passed to other modules.
  • crawler.py — breadth-first crawl using a shared requests.Session (keep-alive). normalize_url() is the canonical URL form used everywhere (strips fragment, lowercases host, removes noise query params via whole-key matching in _is_noise_param, sorts remaining params). should_crawl() filters out binary assets and tracking URLs.
  • extractor.py — parses HTML with bs4+lxml. Returns a single dict per page with keys: text, links, metadata, hidden_content, inline_scripts, jsonld, comment_links. The links dict has typed sub-lists (a, link_rel, script, img, iframe, form, meta_refresh, canonical). strip_cache_params() removes cache-buster query params (e.g. style.css?v=123) from asset URLs (link_rel/script/img) at extraction time, so snapshots store stable URLs — this is the single source of truth for cache-buster normalisation.
  • baseline.pyBaselineManager handles two separate stores: data/snapshots/YYYYMMDD_HHMMSS/ (raw crawl output) and data/baseline/ (approved reference). Each page is stored as <url_key(url)>.json. Snapshot manifest.json also records crawl errors. approve_* methods are the only path into the baseline.
  • differ.pycompare_snapshots(baseline_pages, current_pages, cfg) returns a flat diff dict (incl. unexpected_jsonld, new_broken_urls, known_broken_urls). score_diff(diff, cfg) maps it to {score, level, reasons, exit_code}. Thresholds: yellow ≥ 20, red ≥ 60 (configurable). Unexpected JSON-LD @type is detected independently of page diffs (flags only types not allowed AND not already in the baseline for that URL). A single new external domain = 50 pts (yellow); combined with hidden content = 90 pts (red).
  • checker.py — stateless helpers for whitelist enforcement (check_links_against_whitelist), security-header auditing (check_security_headers, informational only — not scored), and webshell/backdoor filename detection (is_suspicious_url, word-boundary anchored to avoid false positives). Called from __main__.cmd_check only on status == 200 pages.
  • alerter.py — sends alert only when level >= min_level. SMTP password is read from the env var named in smtp_password_env (default: SCANNER_SMTP_PASSWORD).
  • __main__.py — wires everything together. cmd_scan = cmd_crawl (crawl + save snapshot) + cmd_check (load baseline + snapshot, diff, score, report, alert). Exit codes: 0=green, 1=yellow, 2=red.

Data layout

data/baseline/manifest.json          # approved_at, approved_by, note, urls[]
data/baseline/pages/<16hex>.json     # one file per approved URL
data/snapshots/YYYYMMDD_HHMMSS/      # one dir per crawl run
reports/YYYYMMDD_HHMMSS/report.{json,md}
config/allowed_external.yaml         # set of permitted external domains
config/ignore_rules.yaml             # normalization selectors + regex patterns

Scoring weights (defaults, all configurable in config.yaml)

Event Points
Suspicious filename (webshell/backdoor link) 60
New external domain 50
Hidden content (CSS-hidden element with links/text) 40
Meta-Refresh 40
Unexpected canonical 35
Unexpected JSON-LD @type 35
New internal URL 30
New suspicious inline script 30
Comment-embedded link 25
Large text addition (>200 chars) 20
Missing internal URL 20
New broken link (4xx/5xx, not in baseline) 20

Security headers are audited and reported but not scored. Known broken links (already in the baseline) are reported but not scored.