Configure and run a crawler
Create a crawler
Section titled “Create a crawler”- Enter a name and the full public XML sitemap URL.
- Choose Disabled, Daily, Weekly, or Monthly scheduling.
- Choose the reference-ID extraction rule.
- Set chunk size, overlap, and score threshold.
- Save.
Nested sitemap indexes and compressed .xml.gz sitemaps are supported.
Private or password-protected sitemaps are not.
Doublewire markers only is the safest reference option. Use schema.org product-ID fallback only when those product IDs or SKUs are stable identifiers in your source system.
Start a crawl
Section titled “Start a crawl”Open the crawler and select Crawl Now. Doublewire reads the sitemap, discovers allowed pages, refreshes known URLs, extracts main content, splits it, and builds the searchable index.
If failed pages exist, choose whether the run should also retry them. Use Retry Failed when you want only failed pages retried.
Stop or reset
Section titled “Stop or reset”Stop Crawl cancels unfinished work. Already indexed pages remain available, while unfinished pages become Cancelled and can be picked up by a later crawl.
If the active page count is above the current allowed limit, Doublewire offers Reset and Crawl instead of an invalid normal run. It verifies the sitemap, removes the old active page list from assistant search, and starts again within the current limit.
Schedule notifications
Section titled “Schedule notifications”Choose crawler-run notification preferences under Profile > Notification preferences. Each available delivery surface can be set to all results, failures only, or disabled.