For webmasters

About our crawler

You reached this page from a user agent string in your server log. This page says what that visit was, what we kept from it, and how to end it. This page is not legal advice.

User agent

ConformTrailBot/1.0 (+https://conformtrail.com/bot)

Last updated: 2026-09-07.

Why your site was read

ConformTrail records what public web pages displayed and watches them for changes, so that the operator of a site has a dated record of its own pages. Your site was read because it is a public business website of the kind this service is built for. Nothing about it has been published: no public page of ours names your firm, and our own robots.txt keeps the addresses under which a report is served out of every index.

If a report about your site exists, it sits at a private address that we send to the firm itself. That address opens without a sign in for whoever holds it, and the pages it shows carry screenshots of what your pages displayed on the date we read them. Write to [email protected] with the domain if you want to know whether such a report exists; how to have what we hold deleted is under How to block it below.

There are two occasions on which our crawler reaches your server, and your log looks the same for both.

  • We read the site ourselves. A site that is not a customer is read once. We read it again only if the first read did not finish, or if someone at your firm asks us to. A second look at what we already read does not reach your server again. Customers who order monitoring are re-read on a schedule, because that is the thing they ordered.
  • A visitor typed your address into the free scan on our own website. Anybody can type any domain into that form, so this visit was started by a person at their own keyboard, not by us. It reads the home page and a few pages linked from it, under the rules below: robots.txt first, at least one second between two page addresses, and the user agent at the top of this page. It takes no screenshot and keeps no copy: one log line with the domain and a number is all that stays with us.

What it does

  • Opens public page addresses of a website in a browser and records what is displayed on them: the text, the element it sits in, and the date it was seen.
  • Reads robots.txt first and checks every page address against it. A page address you disallow is not opened. The files of a page that is open take a different route: the browser loads what that page needs, the way any browser does, images, stylesheets, scripts and fonts, from wherever the page points them. To keep us away from all of it, disallow everything for our user agent, as shown under How to block it below: then no page address of your site is opened.
  • Keeps at least one second between two page addresses of the same host, and takes a longer Crawl-delay from your robots.txt up to ten seconds. Ten seconds is the longest spacing we take from that line: a higher number is read as ten. The files a page loads are not spaced out; they arrive with the page, as they would for a visitor.
  • Sends the user agent above on every request, the page and the files it loads alike, so what stands in your log is what stands on this page.
  • Lists the documents that public pages link to, PDF and office files, and asks each of those addresses for its type, its size and its date, so an inventory can say where each document comes from.
  • Takes screenshots of public pages to show what was displayed at that place on that date.

What we keep of your pages

  • A copy of each page as we read it, with what that load measured. It is stored with the other data of that scan and is never published, and a later look at it does not reach your server again.
  • Extracts of the text we read go to the service provider the privacy notice names. Nothing is sent there to train a model, and nothing goes there when that service is switched off.
  • That copy has no fixed deletion date today: we keep it until you write to [email protected] with the domain and ask us to delete it, which we then do. What else we hold, and for how long, is on the privacy notice.

What it never does

  • Submit a form. Search boxes, contact forms and filters are left alone. Only GET and HEAD leave the browser; anything else is stopped before it is sent.
  • Log in, or try to. No credentials, no password fields, no member portals.
  • Upload anything anywhere.
  • Click cookie banners or consent dialogs to get past them.
  • Ignore robots.txt for a page address, a rate limit, or a block.
  • Type into a field, so nothing on a screenshot was put there by us.
  • Attempt to bypass a paywall, a login wall or any access control.

How to block it

This is the way that works without us. Add it to your robots.txt:

User-agent: ConformTrailBot
Disallow: /

When we read a site ourselves, that read runs on its own and fetches your robots.txt before it opens a page address, so the line above holds from the next such read on. The free scan on our website is answered by a service that stays running and keeps the robots.txt it has already read of a host, so a line added after a free scan of your domain holds there once that service is restarted. We do not promise a date for that restart.

You can also write to [email protected] with the domain. We stop reading the site, delete what we hold about it, and confirm that to you in writing within 30 days. There is no automatic list behind that address: a person reads the mail and acts on it.

Contact

Write to [email protected]. The operator, the postal address and the applicable law are at the bottom of this page and on the operator page. See also the privacy notice.