About GrafStackBot

If you found GrafStackBot in your server logs, this page explains why we visited and how to opt out.

Why we crawled your site

GrafStack studies how news sites structure content for AI search, licensing, and agents. GrafStackBot reads a sample of public pages from news publishers to measure:

Your site appeared on our list of news publishers, compiled from public directories of news outlets and press-association membership lists.

  • Structured data and schema markup
  • robots.txt rules and directives for AI crawlers
  • Rights declarations, such as RSL
  • Content structure at the page, section, and paragraph level
  • Answer-engine visibility — whether AI assistants cite your reporting

What GrafStackBot does not do

  • Train AI models on your content
  • Republish, sell, or share your content
  • Access paywalled or login-gated pages
  • Bypass any access control
  • Ignore robots.txt

How we crawl

User agentGrafStackBot/1.0 (+https://grafstack.com/bot)
From headerEvery request also sends From: crawler@grafstack.com
robots.txtWe check and follow your robots.txt before every crawl, including our own reads.
Rate limitNo more than 30 requests per minute per domain (one request every two seconds), one connection at a time.
ScopeA sample of fewer than 40 public pages per domain — about 20 articles plus a few well-known files — not a full-site crawl.
FrequencyEvent-driven. We do not run scheduled bulk re-crawls; typically no more than once per quarter per domain, or on request.
Source IPsA static list of source IP ranges will be published here before the first live crawl.

How we use the results

  • We build aggregate benchmarks across news publishers.
  • We do not publish scores or rankings naming individual publishers without their consent.
  • You get your own results at no cost. Email us to request your report.

Data retention: We keep raw page snapshots only as long as needed to compute scores, then delete our copies. We retain structural scores for benchmarking.

How to opt out

Add these lines to your robots.txt:

User-agent: GrafStackBot
Disallow: /

GrafStackBot checks robots.txt before every crawl, so the change takes effect on our next visit. To remove existing data about your site, email us and we will delete your records within 48 hours.

Contact

Questions, report requests, or removal requests:

crawler@grafstack.com