About Crawlers Webmaster tools API Support
Add your page

Crawler documentation

The RssSitemapBot.

Reads RSS, Atom, and sitemap files to feed our global indexing queue.

How it works

#

Identification

WebAtlasRssSitemapBot identifies itself with the following User-Agent.

WebAtlasRssSitemapBot/1.1 (+https://webatlasindex.pl/rsssitemap-bot.php)

⟲

Purpose & logic

A lightweight crawler whose job is to read submitted RSS/Atom feeds and sitemap files, then extract the URLs they contain.

⫶

Filtering layers

Extracted URLs are automatically filtered through three layers:

  • Blocklisted ad/social domains.
  • Technical paths (e.g. /admin, /wp-login, /api).
  • Heuristic ad/tracking parameters (gclid, utm_source, clickid).
⌫

Webmaster control

To block WebAtlasRssSitemapBot entirely, add this to your robots.txt:

User-agent: WebAtlasRssSitemapBot
Disallow: /