Control what your bot crawls

When you crawl a website, you rarely want every page. Your cart, checkout and tag pages add noise, not answers. The crawl settings, just below your list of sites, let you skip those and fine-tune what the bot learns.

Skip pages you don't want

In Ignore URLs when indexing, list the pages to leave out, one pattern per line. Use ** as a wildcard. These pages are skipped on every crawl, so cart and checkout stay out of your bot's answers for good.

The Ignore URLs field with cart, checkout and tag patterns highlighted.
One pattern per line; ** matches anything.

Read the text in your images

Turn on Include images in crawl and the bot reads the alt text and captions on your pages, so it can answer questions about a photo or a diagram.

The Include images in crawl setting with its toggle highlighted.
Pull alt text and captions into the knowledge base.

Save and re-crawl

Click Save crawl settings and the rules apply from the next crawl. Sites re-crawl automatically once a week, so anything you change on your pages flows into the bot without you lifting a finger.

The Save crawl settings button highlighted.
Save, and the rules apply on the next crawl.
Start with cart and checkout

The usual suspects are /cart, /checkout/** and tag or filter pages. Skipping them keeps your bot focused on the pages that actually answer questions.

Just adding a site?

If you only need to add or remove a website, start with Add or remove a website. Come back here when you want to fine-tune the crawl.

Still need a hand?Email our team and we answer within one business day.
Control what your bot crawls | ovellan