Jump to content

Talk:List of companies that block Internet Archive

Add topic
From Consumer Rights Wiki
Latest comment: Sunday at 12:41 by Mr Pollo in topic Infeasible?

moved

[edit source]

w the advent of more archival services being used perhaps more of them might end up getting blocked in which case I feel we could migrate to a general list. SinexTitan (talk) 18:29, 5 April 2026 (UTC)Reply

Companies not specifically excluded but sort of are

[edit source]

Adding information here that doesn't quite fit the topic on the article page, if anyone chooses to make use of it in some manner.


  • Instagram — Only the front-facing website home page, and even that is only partially since at least 1 January 2026. Attempting to archive a post/reel results in a "limitations" error. [1]
    • Facebook too, unsurprisingly, since both are subsidiaries under the Meta name.
  • X — Same problem as Instagram above, though for years longer (since at least its rebranding from Twitter). [2]
  • Reddit — Bit of a different beast here. Most of the time, the Wayback Machine will archive the main post. However if you're wanting the comment replies included, you'll need to use the sub-domain for old Reddit.


There's many websites that hide the user content behind account walls, but those are the biggest offenders I can think of so far. — Sojourna (talk) 00:09, 6 April 2026 (UTC)Reply

More complete(-ish) list

[edit source]

I was looking up Bambu Lab's exclusion of web archiving on Bing and one of the results returned was for a wiki page on archiveteam.org: List of websites excluded from the Wayback Machine.

It is not a dynamic page — sites are added (or removed) when the error is encountered by a registered user, which is why Life360 is not on their list. — Sojourna (talk) 22:59, 4 May 2026 (UTC)Reply

Infeasible?

[edit source]

Is it really feasible to have such a list? I think this article suggests not:

https://amjohnphilip.medium.com/the-internet-archive-just-got-blocked-by-big-news-f74fe59bbd2f "Twenty-three major outlets, including The New York Times and USA Today have updated their robots.txt files to explicitly block the Internet Archive’s Wayback Machine crawler." list isn't in the publicly viewable version of the article. Correct for NYT.

Unless there's a way for IA to have the content but not make it available to train AI, a huge proportion of sites will be blocking it. CrookKilla (talk) 23:04, 11 August 2026 (UTC)Reply

It's better to have than not to have. We won't have EVERY website on here, but enough so people have an idea of how widespread this is. Mr Pollo (talk) 12:41, 16 August 2026 (UTC)Reply