Tell HN: Amazonbot aggressively scraping my website and ignoring robots.txt
Posted by pera 1 day ago
At the beginning of the year I decided to set up a scraping and LLM honeypot on one of my personal websites which included a fake git repo with code containing fake HTTP endpoints. The address to this repo was hidden in a public page inside a comment.
About three weeks ago IP addresses from Amazon Searchbot attempted to make requests to the fake endpoints included inside a shell script.
My robots.txt explicitly includes Amazonbot.
I am honestly surprised that this is coming from Amazon. Is this kind of behavior legal?
Comments
Comment by schappim 1 day ago
Comment by pera 1 day ago
Another detail: the scraper did not attempt to access the endpoints immediately (as it did for hrefs in htmls) but it did it on the day after, twice.
Comment by schappim 1 day ago
1. https://developer.amazon.com/amazonbot/searchbot-ip-addresse...
Comment by pera 1 day ago
They also show up in AbuseIPDB with multiple reports.
Comment by cyanydeez 16 hours ago
Just make sure you have a good lawyer.
Comment by moomoo11 16 hours ago
because idk i could be wrong but some small project vs a 2.5T market cap company is gonna need more than “a good lawyer”
Comment by ipaddr 1 day ago
Comment by KellyCriterion 1 day ago
I hosted a small website with some newfeeds and it got killed by all the AI scrapers in the end.
Comment by Bender 1 day ago
fetch_aws()
{
curl -s \
-H "Accept-Charset: UTF-8" \
-H "Accept-Encoding: gzip, deflate" \
-H "Accept: text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8" \
-H "Accept-Language: en" \
-H "User-Agent: Mozilla/5.0 (Windows NT 11.1; rv:102.0; ) Gecko/20100101" \
-H "Cache-Control: 0" \
-H "Host: ip-ranges.amazonaws.com" \
-H "Connection: keep-alive" \
--referer https://ec2.amazon.com/ \
-o /dev/shm/aws.json \
--url "https://ip-ranges.amazonaws.com/ip-ranges.json"
grep ip_prefix /dev/shm/aws.json | awk -F "\"" '{print $4}' | sort -n | uniq
rm -f /dev/shm/aws.json;
};
Make sure that gives you CIDR blocks, then add it to a script and use a for loop to for AWS in $(fetch_aws);do ip route add blackhole "${AWS}" 2>/dev/null
Google and Cloudflare if you need them: get_google()
{
for line in $(dig +short txt _cloud-netblocks.googleusercontent.com | tr " " "\n" | grep include | cut -f 2 -d :)
do
dig +short txt "${line}"
done | tr " " "\n" | grep ip4 | cut -f 2 -d : | sort -n | uniq
}
get_cloudflare()
{
curl -A Mozilla "https://api.cloudflare.com/local-ip-ranges.csv"|grep -Ev "::|/32"|awk -F "," '{print $1}'|sort | uniq
}
This assumes you have no need to connect to or get connections from Amazon on your web server. If your server is an instance in Amazon that will break DNS resolution unless you are using something outside of AWS for DNS. The gateway should still work just fine. Obviously test from an out of band console if that is an option. This is easier to maintain if you remove DNS records for IPv6 and eventually disable IPv6 listeners.If you paste a couple of lines of the bots from your access logs I can offer more suggestions in the event they try from outside of Amazon.
If you want to have some fun, add a hidden link only the bot will see that points to http://cpanel.yourdomain.tld/ after adding a DNS record for cpanel that points to 169.254.169.254 so they start scraping the AWS cloud-init IP.
Comment by pera 1 day ago
https://developer.amazon.com/amazonbot/searchbot-ip-addresse...
By the way, they used false user-agents.
Comment by az09mugen 1 day ago
Comment by Era24UK 1 day ago
Comment by toomuchtodo 1 day ago
You could go down the rabbit hole and try to find someone at Amazon to tell their crawler to chill. Don’t expect success but would make a fun blog post.
https://www.cloudflare.com/learning/ai/how-to-block-ai-crawl...
Comment by pera 1 day ago
Comment by toomuchtodo 1 day ago
It would be different if they were attempting to brute force credentials to access an endpoint, but they aren’t.
Comment by rovr138 13 hours ago
Weev went to jail for accessing public api's, https://en.wikipedia.org/wiki/Weev#AT&T_data_breach
> The flaw was part of a publicly-accessible URL, which allowed the group to collect the e-mails without having to break into AT&T's system.
It was argued that he didn't circumvent, but it didn't stop them from putting him in jail initially.
Comment by toomuchtodo 7 hours ago
Comment by bellowsgulch 1 day ago
That is, I haven't personally seen a bad actors list that gets used in a fail2ban-like setup.
Comment by bruce_xu 19 hours ago
Comment by jvondev 3 hours ago
Comment by deadcatfound 1 day ago