An update on Wayback Machine access

520 points245 comments15 hours ago
simonw

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running.

I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.

In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.

show comments
basilikum

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access.

The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.

If you got some money to spare, consider donating to them. They need it.

show comments
robotmay

Unrelated, but this week I've been on a memory binge with the Wayback Machine, trying to find old content of mine from the early 2000s. Took me a while but I've finally put together a good bit of info about myself at the time that I'd completely forgotten, and it's all thanks to the Internet Archive storing my little gaming review website from when I was 16. I could barely remember any of the other stuff, it's been genuinely surprising figuring out what I'd forgotten. I couldn't even remember most domains I owned aside from one, which I used as the starting point.

Still can't remember what my Tripod site address was, but that might be lost to time.

Thank you, Archive.org.

BeetleB

Wow, but I wonder if there's more to it.

I've not been able to access web.archive.org from my work computer - I always get the 429 error.

But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.

show comments
userbinator

The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic

Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.

I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.

timpera

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them.

Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.

emaro

It's shame that the AI arms race causes such collateral damage. Free resources were always exploited, but the stakes ($T) and capabilities around AI allow unprecedented abuse. I wish we could go back... :/

I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.

show comments
CqtGLRGcukpy

> We’re getting better at telling abusive bots apart from the people who depend on the Wayback Machine every day. If you think you were blocked in error, email info@archive.org with your operating system, browser, and IP address, and we’ll look into it.

roughly

Bonus points for anyone who’d like to guess how the tragedy of the commons was resolved in the times before the enclosure movement.

show comments
delis-thumbs-7e

I recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.

I really so through some money their way, they do wonderful work.

1vuio0pswjnm7

"Here's what's going on."

Thank you

https://news.ycombinator.com/item?id=49571448

I had a feeling it was due to "AI" companies and developers using "agents"

Not surprised

pelican0

Is it established that the scraping scourge of late is primarily driven by AI companies? Anyone aware of any relevant studies?

Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.

show comments
sicktriple

Anyone else feel like making a new internet and starting over

show comments
thimabi

I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.

show comments
xacky

The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.

show comments
petterroea

I'd be happy to pay a 5$/month donation to get a higher rate limit/more lenient filter put on me

show comments
ilamont

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website?

My blogs are getting slammed and there are issues with cloudflare or captchas.

show comments
potato-peeler

Wayback can’t be accessed through vpn, atleast on proton. Heck, most sites simply block you for using vpn.

xbar

Thank you for the Wayback Machine. It is immensely powerful for good.

msephton

Why can't they capture OS, Browser, and IP address at the time of error?

All that information is available at the point of failure, the user should not need to email it in.

show comments
tech234a

I wonder if they'll end up behind Anubis at some point. I'm surprised it hasn't happened already.

show comments
vlyan

unrelated: if a website gets hit with "This URL has been excluded from the Wayback Machine", do existing snapshots get purged or may they still be preserved somewhere?

show comments
int32_64

Are any AI companies using residential proxies to scrape?

show comments
tgtweak

Can't wayback machine just offer direct access to the archive for a premium and in doing so, pay for the service?

show comments
MattCruikshank

There was a feature on Amazon Web Services for a while, and I wish it was still there...

Downloader pays.

I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.

I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.

brador

The only solution is to make visitors do compute. Compressing files for the archive to access other files would be perfect for this.

Cross verify hashes to prevent cheating.

Ez.

hubraumhugo

There is a HN article on abusive AI crawlers on the front page almost every week, but we rarely talk about the path forward. Web scraping has been around for as long as the internet, and it was fine because we had established best practices (rate limiting, self-identification, robots.txt, etc.) that the industry agreed upon. Now we have AI labs and their crawlers that don't care about any of this gentlemen's agreement. So how do we go from here? Is adding more difficult Anubis and Cloudflare bot protection really the solution? How many millions of human hours and billions in infra costs are we willing to spend on this arms race?

Some approaches that I think are promising:

- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).

- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.

- what else?

[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...

[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...

show comments
UltraSane

Why not put it in S3 with downloader pays?

show comments
lousken

AI companies should pay billions to wayback machine for access

show comments
Onavo

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon.

It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.

I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

show comments
msephton

I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

show comments