548 points ChrisArchitect 16 hours ago 274 comments
Onavo 16 hours ago | parent
It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs.
I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.
croes 16 hours ago | parent
Onavo 16 hours ago | parent
simonw 16 hours ago | parent
quotemstr 13 hours ago | parent
xp84 16 hours ago | parent
celsoazevedo 15 hours ago | parent
faefox 16 hours ago | parent
bonoboTP 15 hours ago | parent
Using the information for training purposes is not the same thing. Not legally the same and otherwise.
drdexebtjl 16 hours ago | parent
xp84 16 hours ago | parent
This is a major "this is why we can't have nice things" situation in my opinion. IA is one of the most valuable gems of the Internet. The only thing that even comes close to preserving our shared history. The damage being caused (both by the effective DDOSing and by the knock-on impact that abuse has in encouraging publishers to remove their content from the archive) is incredibly serious.
katatue 5 hours ago | parent
imglorp 15 hours ago | parent
Content creators could charge by page instead of depending on malware/ad/surveillance revenue. Spam is cut if there's a charge per mail. Scraping abuse goes away, along with a bunch of DDOS garbage.
The impact is a few cents per page or mail, negligible for a human. But if you're consuming a trillion pages per day, you'd reconsider.
novok 15 hours ago | parent
Analemma_ 15 hours ago | parent
Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact that users despise micropayments and we'd rather blame credit card companies for the lack of adoption.
mindcandy 14 hours ago | parent
imglorp 13 hours ago | parent
My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever.
The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.
landgenoot 3 hours ago | parent
You visit website A,A,A,B,C,D,A,A
At the end of the month, you send your entire 20$ randomly to one of the websites you visited.
This will level out everyone's contribution and reward websites with lots of traffic. It eliminates the need for micropayments.
KPGv2 15 hours ago | parent
Because then you're definitely violating US copyright law. There are four prongs of fair use analysis, and one of them is the "nature of the use." In this case, you'd be turning into a commercial use.
Ajedi32 15 hours ago | parent
Onavo 15 hours ago | parent
oasisbob 12 hours ago | parent
The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.
On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.
When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.
"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."
No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.
simonw 16 hours ago | parent
I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior.
In addition to the load it puts on this vital non-profit piece of Internet infrastructure, we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
packetslave 16 hours ago | parent
bsimpson 16 hours ago | parent
gambiting 15 hours ago | parent
ValentineC 15 hours ago | parent
petcat 15 hours ago | parent
organsnyder 14 hours ago | parent
petcat 14 hours ago | parent
Hence, distinction without a difference.
fluffybucktsnek 14 hours ago | parent
petcat 14 hours ago | parent
So Internet Archive, Archive.today, Archive.ph, etc. are all just means to the same end.
HDBaseT 11 hours ago | parent
The internet archive is not designed to circumvent anything. It is not designed to "grant access without having your own access".
fluffybucktsnek 10 hours ago | parent
publlus_enigma 10 hours ago | parent
Archive.org exists to preserve historical snapshots of the public parts of websites, and not to bypass subscriptions or pay walls.
celsoazevedo 14 hours ago | parent
I think it's a distinction worth making.
Not to mention that the Wayback Machine itself isn't exactly a good tool to bypass paywalls as most paid sites don't let them archive paywalled content anyway.
rpdillon 14 hours ago | parent
DaSHacka 11 hours ago | parent
archive.org is the more straight-laced archive that doesn't circumvent sites that try to block it, and removes content they deem 'problematic' even if not illegal or requested by the site owner.
Meanwhile archive.today/ph/is/etc is the guerrilla alternative run by a die-hard datahoarder that seeks to archive the information itself, bypassing whatever blockers/login pages/whathaveyou to achieve the result.
It's nice to have both options. When I archive a site, I usually use both for added resiliency.
sandcat_ 14 hours ago | parent
petcat 14 hours ago | parent
sandcat_ 14 hours ago | parent
petcat 14 hours ago | parent
The end result is exactly the same.
DaSHacka 11 hours ago | parent
sippingabonedry 11 hours ago | parent
fc417fc802 8 hours ago | parent
Substitute almost any disruptive public service to see the issue with your line of reasoning. For example - you seem to think that [ bulldozing private property ] to "construct an emergency fire break" is somehow different than [ bulldozing private property ] for any other reason.
Never mind that the sort of scraping being objected to is actually harmful to service health while what the wayback machine does is almost entirely unnoticeable.
eek2121 12 hours ago | parent
DaSHacka 11 hours ago | parent
normie3000 11 hours ago | parent
I use them. I haven't ever heard mention that the content is edited. Do you have a source?
uqers 10 hours ago | parent
https://arstechnica.com/tech-policy/2026/02/wikipedia-might-...
https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan...?
Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.
Mogzol 10 hours ago | parent
They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.
sam345 9 hours ago | parent
Mogzol 5 hours ago | parent
There's also a bunch of previous hackernews discussions about it:
- https://news.ycombinator.com/item?id=47474255
- https://news.ycombinator.com/item?id=46624740
- https://news.ycombinator.com/item?id=47092006
- https://news.ycombinator.com/item?id=46843805
And you obviously have no reason to believe me, but I was following this when it was happening at the start of this year and can confirm that the DDoS script and archive text replacements really did happen.
fn-mote 9 hours ago | parent
Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.
opello 7 hours ago | parent
Meneth 20 minutes ago | parent
Because there's no working alternative.
koolala 7 hours ago | parent
sam_lowry_ 7 hours ago | parent
Why being shy in the era of stealing AI?
LoganDark 5 hours ago | parent
schnebbau 5 hours ago | parent
You have to include more than that, this isn't common knowledge.
LoganDark 5 hours ago | parent
x______________ 4 hours ago | parent
Wikipedia deprecates Archive.today, starts removing archive links (arstechnica.com) 616 points by nobody9999 6 months ago | hide | past | favorite | 368 comments
jimmydorry 2 hours ago | parent
1. https://news.ycombinator.com/item?id=46843805
luckylion 16 hours ago | parent
Very understandable, you can't store all 15000 pages of any random website and update them etc etc, but that makes them pretty useless for indirect scraping because you usually don't want a tiny taste, you want everything.
toomuchtodo 16 hours ago | parent
https://en.wikipedia.org/wiki/Tragedy_of_the_commons
(no affiliation)
ronsor 15 hours ago | parent
On the other hand, the Internet Archive is a non-profit offering a free public resource.
toomuchtodo 15 hours ago | parent
itintheory 15 hours ago | parent
The cheapest solution is to require a login and rate limit by API key. I also have strong feelings about the tragedy of the commons.
toomuchtodo 14 hours ago | parent
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
userbinator 5 hours ago | parent
bradly 15 hours ago | parent
> Hacker News and the Rails forum are blocking the text fetcher, so I'm using the browser workflow to inspect the pages directly
TeMPOraL 15 hours ago | parent
Unfortunately I sometimes have to browbeat Claude into acting like an agent of the user is supposed to. Usually it works, though last time it refused to recognize my moral argument (on the grounds that it's not bound to my interests exclusively and needs to protect the interests of its maker too).
bradly 15 hours ago | parent
aaron_m04 15 hours ago | parent
bradly 15 hours ago | parent
xena 14 hours ago | parent
ghaff 12 hours ago | parent
recursive 12 hours ago | parent
TeMPOraL 1 hour ago | parent
I wouldn't want them to. The whole point of using agents to do stuff on the web for me, is for them to do the stuff on the web for me.
This is the reverse of "do not track" case. It'll not be effective because every service will set it to DISALLOW by default anyway, because it costs them nothing, and for most services, it actually is what they want anyway - most of businesses on the web are making money on wasting people's time, and for that, they need to force themselves on people; end-user automation defeats that, so they actively fight it (and complain a lot).
Dylan16807 11 hours ago | parent
bityard 12 hours ago | parent
dhx 3 hours ago | parent
What you suggest is explicitly not a purpose of robots.txt per RFC9309[1]:
"These rules are not a form of access authorization."
HTTP 429 and HTTP 403 are what servers are meant to return to clients to slow them down or tell them to stop doing something without having first gained authorisation.
Analemma_ 14 hours ago | parent
compiler-guy 13 hours ago | parent
cruffle_duffle 13 hours ago | parent
fineIllregister 13 hours ago | parent
compiler-guy 13 hours ago | parent
TeMPOraL 1 hour ago | parent
Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason for the tension is simply that most websites and businesses on-line rely on users wasting their time, and cannot abide any form of end-user automation. Their business plans hinge on their ability to force themselves on you.
daveoc64 12 hours ago | parent
e.g. a prompt of "fetch <article URL> and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.
compiler-guy 11 hours ago | parent
The problem is that it is hard to distinguish your one off (which seems perfectly fine) from the tidal wave of bad actors.
kelnos 5 hours ago | parent
That's the scale argument.
TeMPOraL 1 hour ago | parent
Same with browsing HN, btw. I have a row of 9 HN tabs open, all of them opened at the same time, as I scrolled the front page and middle-clicked on thread link to anything interesting.
cruffle_duffle 13 hours ago | parent
ryandrake 12 hours ago | parent
pantsforbirds 15 hours ago | parent
Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021 when we were doing this.
subarctic 15 hours ago | parent
msephton 14 hours ago | parent
carlosjobim 13 hours ago | parent
msephton 13 hours ago | parent
But anyway, no, I wouldn't keep finding reasons. I donate to them every year already. Somebody asked if I would be willing to pay and my answer was "yes, but".
It would need to be improved because certain aspects of it suck right now, not only the error this post is about. They only need go as far as their forums and github repos to see the community feedback.
carlosjobim 10 hours ago | parent
msephton 10 hours ago | parent
bonestamp2 14 hours ago | parent
DaSHacka 11 hours ago | parent
usr1106 6 hours ago | parent
bonestamp2 1 hour ago | parent
progval 3 hours ago | parent
bee_rider 14 hours ago | parent
cloakley 13 hours ago | parent
pantsforbirds 6 hours ago | parent
pantsforbirds 5 hours ago | parent
with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?
RobotToaster 15 hours ago | parent
echelon 15 hours ago | parent
QuantumNomad_ 14 hours ago | parent
However, that experiment ended. They mention there were some learnings and they then say:
> The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to a number of projects that are still in use.
https://wiki.archiveteam.org/index.php/INTERNETARCHIVE.BAK
I would really like to know if any sort of thing like that is still ongoing and if it’s accessible to people in general. Would be nice to mirror some data from IA to my local drives, for example via BitTorrent or IPFS, to have it for offline exploration and personal archive.
I know that individual items have torrents. And I’ve downloaded a few that way but always it ends up only using the “web seed” (i.e. the BitTorrent client is retrieving the files from IA via HTTP) because there are no one seeding some random single item I found. Plus, those torrents are unreliable sometimes because they include meta data files that were since updated but the torrent was not updated and so the web seed is giving the updated files that don’t match what the torrent says their hashes should be. So then you have to jump through some extra hoops to fix that and then resume the download, and all the while the HTTP connections to IA servers time out because their servers are overloaded. So when I say I wonder about possibilities of using BitTorrent I mean to retrieve whole collections of many items instead of individual ones, and with actual other peers instead of just having it put load on IA HTTP servers.
giantrobot 12 hours ago | parent
alightsoul 11 hours ago | parent
jader201 15 hours ago | parent
Appalling, yes. But also expected. I'm surprised they haven't been the target of scrapers for years. But sites putting their content behind login walls and other anti-bot mechanisms has certainly exacerbated this. But again, this isn't at all a surprising progression.
> we've also already seen some sites opt out of the Wayback Machine to prevent their content from being scraped via this alternative route.
To be fair, another big motivation was likely users on sites like HN using archive.org (and similar sites) to get around their paywalls. In fact, I'd be surprised if this wasn't a big motivator.
Again, it sucks, but it's not at all surprising to see it progress like this. I wouldn't be surprised to see similar blocks on other archive sites eventually.
autoexec 12 hours ago | parent
eek2121 12 hours ago | parent
A popular tech news site blocked my phone because of Apple Private Relay. That didn't last long because their traffic fell off a cliff when that happened.
Many sites are throwing more captchas at the problem, without understanding that captchas don't actually help with LLMs, they just hinder normal users and primitive scripts. LLMs solve captchas just fine.
Some big sites have put up improved paywalls. I'm fine with subscribing to a quality site, however, WSJ and all the other big media sites routinely spit out regurgitated garbage that can be had for free elsewhere (and due to political spin, their garbage is less valuable than the free versions of said content).
Some folks are declaring the internet dead. I wouldn't go that far, however, I will say that a reckoning is going to happen, especially when advertisers figure out that most ads served on basically every website are no longer viewed by humans.
matt_heimer 12 hours ago | parent
Open access doesn't seem sustainable.
But I might just grumpy about spending another hour this week adjusting rules to prevent bots.
intrasight 10 hours ago | parent
TZubiri 11 hours ago | parent
Kodiack 10 hours ago | parent
However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a small donation their way. They provide an incredibly valuable service and I love the benefit that I get from them just for personal side projects.
mkatx 9 hours ago | parent
sippingabonedry 9 hours ago | parent
Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
fc417fc802 9 hours ago | parent
usr1106 6 hours ago | parent
Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.
What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.
grumbelbart2 5 hours ago | parent
https://developers.google.com/search/blog/2006/09/how-to-ver...
dewey 3 hours ago | parent
4thguy 3 hours ago | parent
ezekiel68 10 hours ago | parent
e40 8 hours ago | parent
hedora 8 hours ago | parent
I often cannot get past captchas, and archive.org is one of the fallbacks I try.
However, archive.is, etc are more reliable.
I wish the internet archive acted more like a library system, where multiple organizations could mirror the content.
They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives down yet.
throwawayk7h 7 hours ago | parent
simonjgreen 4 hours ago | parent
CqtGLRGcukpy 16 hours ago | parent
tech234a 16 hours ago | parent
stickfigure 15 hours ago | parent
autoexec 12 hours ago | parent
phendrenad2 12 hours ago | parent
danbolt 12 hours ago | parent
I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit?
[1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...
stickfigure 10 hours ago | parent
Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate.
You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.
murderfs 7 hours ago | parent
If you waste your user's time with something that would take a full minute to run on a datacenter core, you're costing the scraper something like $0.000005: 360 W TDP on a 128-core EPYC 9754 * $0.10/kWh. In reality, it'll be substantially less than that, because CPUs don't use 0W at idle.
The only way this would make any sense is if there were many more scrapers than users and scrapers cared more about latency than real users, but that's the exact opposite of reality. The entire endeavor is so fundamentally misguided that it almost seems like a psyop.
BeetleB 16 hours ago | parent
I've not been able to access web.archive.org from my work computer - I always get the 429 error.
But I then pull out my phone and can access it just fine. All along I was assuming my company was blocking it. Still weird that it happens every time from my work PC and never from my home one.
dotmanish 16 hours ago | parent
flexagoon 16 hours ago | parent
BeetleB 14 hours ago | parent
Wonder who the bad actors in my company are...
flexagoon 13 hours ago | parent
You can try emailing the address mentioned in their post so they adjust their filters to match just the bot networks more precisely
ButlerianJihad 13 hours ago | parent
BeetleB 10 hours ago | parent
iamacyborg 12 hours ago | parent
jcrawfordor 14 hours ago | parent
BeetleB 14 hours ago | parent
hedora 8 hours ago | parent
novok 15 hours ago | parent
Try making a vpn via digital ocean for example and you'll see similar patterns.
giantrobot 12 hours ago | parent
GetSMS 5 hours ago | parent
timpera 16 hours ago | parent
Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you often can't access the WM at all. I hope they'll find a way to relax those.
lousken 15 hours ago | parent
KPGv2 15 hours ago | parent
roblh 15 hours ago | parent
jMyles 15 hours ago | parent
autoexec 12 hours ago | parent
jMyles 11 hours ago | parent
* If you train AI on it, you have to afford public access to it.
* Nobody can exact violence against anybody else in response to that person providing public access to any data anymore (ie, all bytestrings are public domain).
That's the world I'd like to try in the coming years.
autoexec 12 hours ago | parent
0xDEAFBEAD 9 hours ago | parent
ignoramous 15 hours ago | parent
MattCruikshank 15 hours ago | parent
Downloader pays.
I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc.
I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LLM tokens today. But this just feels like such a useful thing that it baffles me that it doesn't exist already.
vlyan 15 hours ago | parent
tech234a 14 hours ago | parent
Note that Archive Team is separate from the Internet Archive.
basilikum 15 hours ago | parent
The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gatekeeper showing me the middle finger.
If you got some money to spare, consider donating to them. They need it.
superxpro12 15 hours ago | parent
The future is bleak :\
mrguyorama 14 hours ago | parent
It's going to suddenly be extremely valuable that wikipedia didn't settle for having a small rainy day fund and instead ceaselessly grabbed every fucking donation they could for two decades so they can fight such a legal battle.
pvab3 13 hours ago | parent
nephihaha 13 hours ago | parent
jasonfarnon 11 hours ago | parent
khafra 11 minutes ago | parent
niuzeta 10 hours ago | parent
hubraumhugo 15 hours ago | parent
Some approaches that I think are promising:
- A robots.txt V2[0] as a standard way for website owners to state how bots and AI crawlers can use their online content and where to go (e.g. distinguish search from AI training use cases, point to a downloadable file instead of crawling everything, etc.).
- Something like Web Bot Auth[1] as a non-centralized standard for self-identifying bots and agents cryptographically. This would allow websites to allow or deny bots very precisely.
- what else?
[0] https://datatracker.ietf.org/doc/draft-vaughan-machine-reada...
[1] https://datatracker.ietf.org/doc/html/draft-meunier-http-mes...
maxrev17 14 hours ago | parent
edelbitter 12 hours ago | parent
msephton 15 hours ago | parent
jolmg 14 hours ago | parent
It's surely to serve as data to help tell humans apart from bots.
> Changes made by IA shouldn't become my responsibility.
They're a free service. It's ultimately not their responsibility to service you either.
msephton 13 hours ago | parent
IA have broken it and have no real idea how to make it better so they are going to whitelist IPs or browsers or entire operating systems? Wild.
jolmg 12 hours ago | parent
kjs3 12 hours ago | parent
Wild.
msephton 10 hours ago | parent
jolmg 8 hours ago | parent
You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that info?
> They created a problem
No, they're dealing with a problem, and compromised that some human users may unfortunately get blocked.
> and now users have to pay for the inconvenience
You don't have to anything. You can just not use them. They don't owe you their service.
Somebody is handing out free apple lollipops, they ran out, compromised on giving grape ones, and now you're complaining you're being forced to eat a grape one and you don't like grape. Don't eat it.
msephton 7 hours ago | parent
HDBaseT 10 hours ago | parent
The fact you can even access the Internet Archive for free is a result of tens thousands of human hours striving for one goal. Digital Preservation. If you rely so much on IA, you should consider donating.
msephton 10 hours ago | parent
doctor_radium 9 hours ago | parent
pelican0 14 hours ago | parent
Beginning to think that the difficulty to browse most websites nowadays due to throttling, is yet another negative externality of AI development that society is forced to bear.
userbinator 5 hours ago | parent
emaro 14 hours ago | parent
I really don't see any solution to this; the scrapers probably wouldn't even mind destroying sources like IA too much, which would leave them as the only "authorative" source of knowledge in the end. Best way is likely regulation incl. hefty (!) fines, but politics are too slow and too fragmented to be effective. So... Enjoy it while it lasts, I guess.
zdragnar 14 hours ago | parent
- charging (news / journalist services)
- gate-keeping (X forcing log-ins)
- enshittifying (lots of ads and degraded service)
The fact that the way back machine is incredibly useful but most people didn't know about it or use it very much doesn't change the fact that it has basically become very popular... only with LLM agents rather than humans. Ads alone aren't enough to support human traffic for many sites with human traffic.
thimabi 14 hours ago | parent
extralongdivisi 14 hours ago | parent
hamandcheese 12 hours ago | parent
extralongdivisi 11 hours ago | parent
That's the point. The solution should avoid information not being available. Requiring login will incentivize bots to create spam accounts and move the battle to a new frontier, hurting real people in the process.
TechSquidTV 6 hours ago | parent
int32_64 14 hours ago | parent
xena 14 hours ago | parent
oasisbob 13 hours ago | parent
brador 13 hours ago | parent
Cross verify hashes to prevent cheating.
Ez.
ilamont 13 hours ago | parent
My blogs are getting slammed and there are issues with cloudflare or captchas.
hamboomger 1 hour ago | parent
But maybe the problem is that they can't serve the data from the other websites like this, if they use it commercially. Right now they have non-commercial use, from what I understand.
tgtweak 12 hours ago | parent
edelbitter 12 hours ago | parent
robotmay 12 hours ago | parent
Still can't remember what my Tripod site address was, but that might be lost to time.
Thank you, Archive.org.
msephton 10 hours ago | parent
All that information is available at the point of failure, the user should not need to email it in.
ericpauley 10 hours ago | parent
1vuio0pswjnm7 7 hours ago | parent
Thank you
https://news.ycombinator.com/item?id=49571448
I had a feeling it was due to "AI" companies and developers using "agents"
Not surprised
potato-peeler 7 hours ago | parent
delis-thumbs-7e 6 hours ago | parent
I really so through some money their way, they do wonderful work.
xbar 5 hours ago | parent
userbinator 5 hours ago | parent
Thank you for not immediately blaming it on "AI bots". I suspect there's some entity manufacturing consent for strong identity/age verification/sanctioned-browser-OS "walled garden" Internet, and these random DDoSes are part of that.
I knew something was up when a few alternative YouTube front-ends I use suddenly put up the 'nubis and complained about the high volumes of traffic they were getting flooded with; of course someone actually going after that data would be aiming their "AI bots" at YouTube directly instead of trying to suck it through a tiny little-known proxy-site, so it really strained the credibility of the argument.