Reads robots.txt Sitemap directive and fetches sitemap.xml / sitemap_index.xml before crawling. All listed URLs are added to the crawl queue so pages without any inbound links are still scanned. Configurable via crawl.sitemap. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| __init__.py | ||
| __main__.py | ||
| alerter.py | ||
| baseline.py | ||
| checker.py | ||
| config.py | ||
| crawler.py | ||
| differ.py | ||
| extractor.py | ||
| plain.py | ||