User-agent: * Allow: / # ── Verified SEO crawlers (explicit allow; prevents edge WAF misclassification) ── User-agent: SemrushBot Allow: / User-agent: AhrefsBot Allow: / User-agent: AhrefsSiteAudit Allow: / User-agent: MJ12bot Allow: / User-agent: DotBot Allow: / User-agent: rogerbot Allow: / User-agent: BLEXBot Allow: / User-agent: Screaming Frog SEO Spider Allow: / User-agent: SiteAuditBot Allow: / User-agent: DataForSeoBot Allow: / User-agent: Mediatoolkitbot Allow: / # ── Public SEO landing pages (explicit allows for clarity) ── Allow: /username-search Allow: /reverse-image-search Allow: /reverse-username-search Allow: /find-dating-profiles Allow: /email-breach-check Allow: /phone-lookup Allow: /profile-picture-search Allow: /find-people-online Allow: /face-intelligence Allow: /scan-my-online-presence Allow: /search-engines-to-find-people Allow: /compare/ Allow: /developers Allow: /intelligence-graph Allow: /account-finder Allow: /exposure-intelligence Allow: /safe-osint-for-ai Allow: /learn Allow: /learn/ # ── Public dynamic SEO prefixes (curated; thin variants noindex via X-Robots-Tag) ── Allow: /risk/ Allow: /for/ Allow: /atlas/ # /username/ entity pages — all canonical to /username-search and noindex. # Disallowed (was Allow) on 2026-06-26: Google had 330k of these stuck in # "Crawled - currently not indexed" because Allow + noindex created a recrawl # loop. They're already de-indexed; blocking crawl drops them from the report # over ~4-6 weeks and stops dragging whole-site quality signals. Internal links # are wrapped in so they render as