Reads robots.txt Sitemap directive and fetches sitemap.xml / sitemap_index.xml before crawling. All listed URLs are added to the crawl queue so pages without any inbound links are still scanned. Configurable via crawl.sitemap. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> |
||
|---|---|---|
| .. | ||
| config | ||
| config.yaml | ||