About GWMWebAuditBot
If you found GWMWebAuditBot in your site’s access logs, this page explains what it is, what it took and how to stop it. It is Green Wire Media’s research crawler and it visits each site once.
We measure how many Latvian, Estonian and Lithuanian sites are reachable by AI search engines, how fast they load, and whether their email is protected against spoofing. The result will be a public study of summary figures. No publication will ever name a specific site alongside a specific weakness of it.
All four requests come from the same address within seconds of each other, and every one of them carries our own name. The assistant tokens are there because those are precisely what site defences check: we are measuring what gets blocked, not trying to get around the blocking.
| Measurement | User-Agent |
|---|---|
| Browser controlThe control measurement. Without it there is no way to tell a site’s own policy from a problem at our end. | Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/126.0.0.0 Safari/537.36 GWMWebAuditBot/1.0 |
| Search assistantChecks whether the site answers a search assistant. Blocking this one takes the site out of the answers. | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; OAI-SearchBot/1.0; +https://openai.com/searchbot; run by GWMWebAuditBot/1.0 |
| Training crawlerChecks whether the site answers a training crawler. Blocking this one is an entirely reasonable choice and is counted separately. | Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.2; +https://openai.com/gptbot; run by GWMWebAuditBot/1.0 |
| Plain clientUs, with no browser tokens at all. Separates blocking assistants from blocking anything that is not a browser. | GWMWebAuditBot/1.0 (+https://greenwiremedia.com/lv/baltijas-majaslapu-petijums/gwmwebauditbot; research crawl; opt out: info@greenwiremedia.com) |
We never vary these tokens to get past a defence, and we never route requests through proxies. Doing either would make the study’s headline figure meaningless.
We collect
- The homepage exactly as any visitor receives it.
- The robots.txt file and the sitemap, if one is declared.
- Public DNS records: addresses, name servers, and email authentication records.
- HTTP headers, the certificate, and how long the page took to load.
We do not collect
- Nothing behind a login, a payment, or any form of access control.
- No personal data, no contact form contents, no customer data.
- No searching for exposed files, admin paths or vulnerabilities. We look, we do not probe.
Each site is visited once. Nothing is revisited, only one connection is open to a server at a time, and we honour robots.txt including for ourselves.
Two ways, both work.
Write to us
Send the domain to info@greenwiremedia.com. We take it off the list and reply. If anything has already been collected, we delete it.
Or say so in robots.txt
Two lines in your site’s /robots.txt and the crawler leaves the site alone:
User-agent: GWMWebAuditBot Disallow: /
A site that opts out is not counted as a site that failed. We account for it separately, as an adjustment to the study’s sample.
Green Wire Media, a web development and marketing agency in Riga. Questions about the study, its method, or about a specific request in your logs: info@greenwiremedia.com.
The full method will be published alongside the study. It will also state where the measurements were taken from, since load times depend on location.