Stay discoverable in search while disallowing AI training
Posted by djfergus 9 hours ago
Comments
Comment by DharmaPolice 6 hours ago
Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page".
If you don't want your content to end up in some database/archive don't publish it for the whole world to see.
Comment by 1vuio0pswjnm7 7 hours ago
Is that really true
CF classifies anyone not using a popular browser with Javascript enabled as a "bot"
CF fingerprints www users
As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused
IA = Internet Archive
Comment by userbinator 7 hours ago
This is my biggest complaint about CF. They are implicitly supporting user-agent discrimination in favour of Big Browser, instead of discriminating on actual behaviour.
...and of course there are already companies running tons of VMs with "officially sanctioned" browser + OS stacks, that can get past all these "protections", for a fee.
"AI bots" is the newest boogeyman they came up with to take away freedom.
Comment by devmor 7 hours ago
Comment by basilikum 1 hour ago
Comment by RobotToaster 3 hours ago
Comment by actionfromafar 2 hours ago
Comment by skybrian 8 hours ago
https://blog.google/innovation-and-ai/products/an-update-on-...
Comment by nirmeetimthebes 7 hours ago
Comment by userbinator 6 hours ago
[1] Analog hole and other workarounds aside, naturally.
Comment by qsbuilder 5 hours ago
Comment by VBprogrammer 1 hour ago
If I was a big AI company I'd certainly be tempted to make sure that anyone who excluded themselves from "AI training" also got themselves excluded from AI results.
Comment by tchalla 3 hours ago
Comment by OroPla 3 hours ago
I could totally see a future where Google just stops showing you links to actual web pages altogether and just gives you their chatbot.
Comment by nullbio 5 hours ago
Comment by kinduff 8 hours ago
I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email:
"Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."
Comment by mskalski 8 hours ago
[1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth
Comment by AnonC 8 hours ago
> We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search.
I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.
Comment by shark1 3 hours ago
Specified Time Frame ;)
Comment by RobotToaster 3 hours ago
Comment by gdiamos 6 hours ago
Comment by nicolodev 4 hours ago
Comment by fwlr 4 hours ago
Comment by GetSMS 6 hours ago
Comment by dzhiurgis 7 hours ago
Comment by userbinator 6 hours ago
Comment by DaSHacka 7 hours ago
Comment by gchamonlive 9 hours ago
Comment by arm32 9 hours ago
17 years
Fun lava lamp story, though
Comment by zergrush 9 hours ago
theres no way my index.html page with nothing is getting 10000 hits a day
wtf?
Comment by itake 9 hours ago
I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).
Comment by JoshTriplett 8 hours ago
Comment by DaSHacka 7 hours ago
I just want basically captive.apple.com with a shorter domain and the webserver not even listening on port 443 at all.
Surprised someone hasn't made this yet, it only requires one spare public IP.
Comment by itake 7 hours ago
Comment by jesterson 6 hours ago
For a site with proven visitors 100,000 per month CF tell me they saved 90,000 over 1,000,000 visitors (rough numbers).
Sure they know large numbers make people feel good.
Comment by gleezard 9 hours ago
There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken that away. New business models, perhaps actually decent ones are required.
Death of ads everywhere? Good fucking riddance.
Posted from LibreWolf.
Comment by tracerbulletx 8 hours ago
Comment by lostmsu 12 minutes ago
Comment by octoberfranklin 7 hours ago
You aren't independent. You work for the BigTech company that serves ads on your site.
Comment by gleezard 8 hours ago
Comment by antonvs 3 hours ago
Comment by India_InfraNote 7 hours ago
Comment by aaron695 7 hours ago