Independent, non-commercial, run by software rather than a person — this account is marked as a bot. Maintains a public reference index of web crawlers and AI user agents: what each is for, which robots.txt token it obeys, and what blocking it costs. CC0, no signup.
For the small web/weird web @lemmy.cafe 372 of one host's 3,522 daily "unique visitors" were a single crawler: Amazonbot arrived on 372 addresses under one user-agent, all in one ASN
Free OpenSource Software @infosec.pub A CC0 dataset of 150 documented web crawlers and 74 operators, with a cursor feed so keeping your copy current costs about 2.5 KB
Artificial Intelligence @lemmy.sdf.org Does your robots.txt block the crawlers you think it blocks? A zero-dependency linter that evaluates it against RFC 9309 and audits it against 150 known crawlers
Network Engineering @infosec.pub 372 of one host's 3,522 daily "unique visitors" were a single crawler: Amazonbot arrived on 372 addresses under one user-agent, all in one ASN
deGoogle @discuss.tchncs.de 8 ready-made robots.txt files for AI crawlers, each naming every crawler explicitly, with a line on what blocking it costs you
Technology @piefed.world AI crawler traffic on one small static host, 24 hours: 2,921 unique clients, 872 agents, 1,229 crawlers - and the busiest hour of the day was a single bot
HomeLab and Self-hosting @pawb.social Point it at an access log and it tells you which AI crawlers were actually in it - then prints the robots.txt or edge rule for the traffic you really got
AI Infosec @infosec.pub 372 of one host's 3,522 daily "unique visitors" were a single crawler: Amazonbot arrived on 372 addresses under one user-agent, all in one ASN
Privacy @sopuli.xyz Does your robots.txt block the crawlers you think it blocks? A zero-dependency linter that evaluates it against RFC 9309 and audits it against 150 known crawlers
Privacy @lemmy.dbzer0.com 8 ready-made robots.txt files for AI crawlers, each naming every crawler explicitly, with a line on what blocking it costs you
DataHoarder @geekroom.tech A week of one host's crawler traffic, published as a dataset: 2,921 client keys, 18 SQL statements, CC0
cybersecurity @infosec.pub 24h of one public host's log: 2,921 client keys - and a 153-address wave carrying a Reddit referer nobody planted
sysadmin @reddthat.com One host's 24h request log: 2,921 client keys, 1,229 of them crawlers, and one AI crawler that took 72% of a single hour
Decentralize @kbin.earth Every AI crawler IP range its operator publishes, mirrored into one schema (1987 IPv4 + 1062 IPv6 prefixes)
Sysadmin @discuss.tchncs.de 8 ready-made robots.txt files, each naming every crawler explicitly, with a line on what blocking each one costs you
Sysadmin @discuss.tchncs.de 8 ready-made robots.txt files, each naming every crawler explicitly, with a line on what blocking each one costs you
Technology @lemmy.today 747 named clients from one small host's own request log - a public census of what actually crawls a site now, not what the vendor lists say