About Crawlers Webmaster tools API Support
Add your page

Crawler documentation

The DiscoverBot.

Finding new paths. WebAtlasDiscoverBot explores the web to feed our global indexing queue.

How it works

#

Identification

WebAtlasDiscoverBot identifies itself with the following User-Agent.

WebAtlasDiscoverBot/2.1 (+https://webatlasindex.pl/discover-bot.php)

↝

Purpose & logic

A lightweight crawler focused purely on link discovery. It respects robots.txt rules under its own name, downloads at most 512 KB per connection, and honors Crawl-delay to stay low-impact on your server.

⫶

Filtering layers

Every discovered link passes through three filters before it's queued:

  • Blocklisted ad/social domains.
  • Technical paths (e.g. /admin, /wp-login, /api).
  • Heuristic ad/tracking parameters (gclid, utm_source, clickid).
⌫

Webmaster control

To block WebAtlasDiscoverBot entirely, add this to your robots.txt:

User-agent: WebAtlasDiscoverBot
Disallow: /