Briefs

The half-life of a web page

Half the hyperlinks the U.S. Supreme Court cites as legal authority no longer lead to the material they cite, a 2014 Harvard study found. A decade-wide Pew Research sample shows why that is not a legal-citation quirk: roughly a fifth of a given year's web disappears within two years, and well over a third within ten. Durable pages are the exception, not the default.

50%
of hyperlinks in U.S. Supreme Court opinions no longer lead to the material they cite
Zittrain, Albert & Lessig, Harvard Law Review Forum — spot-checked citation sample through October Term 2011, published 2014

Half of the links a Supreme Court justice cites as legal authority now lead nowhere. That figure, published in the Harvard Law Review Forum in 2014, was not a one-off curiosity. It was an early measurement of a decay rate that Pew Research Center confirmed a decade later across the wider web: a quarter of nearly one million sampled pages from 2013 to 2023 were unreachable by October 2023, and the oldest cohort in that sample had lost 38% of its pages entirely. Web content does not fail all at once. It has a decay curve, and for several common classes of page — court citations, news links, Wikipedia references — that curve is steeper than publishers tend to assume.

The court that cites the internet, and the internet that forgets

Jonathan Zittrain, Kendra Albert and Lawrence Lessig examined hyperlinks in the footnotes of three Harvard journals and in every published U.S. Supreme Court opinion, then ran a two-step check: first an automated HTTP status test, then a manual read of each “working” link to confirm it still showed the material the footnote claimed. The automated pass alone understated the problem — of links that returned a normal 200 status in their Supreme Court sample, only 76% actually still led to the cited material. Combining both tests, the authors found that 50% of Supreme Court citation URLs and more than 70% of the URLs in the Harvard Law Review, Harvard Journal of Law and Technology and Harvard Human Rights Journal no longer produced the information originally cited.

That 50% figure runs well above an earlier, related estimate: Raizel Liebler and June Liebert’s study of Supreme Court citations from 1996–2010 had put the “invalid” share at 29%. The gap is a methodology story, not a contradiction — Liebler and Liebert counted only outright error pages, while Zittrain, Albert and Lessig also caught links that resolve normally but silently point to changed or replaced content. Both are real forms of failure; they just measure different things.

A decade of pages, tracked year by year

Pew Research Center’s 2024 analysis took a broader, non-legal sample: roughly 90,000 pages per year pulled from the Common Crawl archive for each year from 2013 through 2023, just under one million pages in total, rechecked in October 2023. A quarter had gone dark — 16 percentage points from individual pages vanishing off otherwise-working sites, 9 points from entire domains disappearing. Age was the dominant variable, and the year-by-year pattern reads like a half-life curve in progress.

Year page was collected Approx. age by Oct 2023 Share no longer accessible
2013 ~10 years 38%
2014 ~9 years 35%
2015 ~8 years 31%
2016 ~7 years 30%
2017 ~6 years 26%
2018 ~5 years 31%
2019 ~4 years 32%
2020 ~3 years 27%
2021 ~2 years 22%
2022 ~1 year 15%
2023 <1 year 8%

The curve is not perfectly smooth — 2018 and 2019 pages show slightly more loss than 2016 and 2017 — but the direction holds: pages collected in 2021 had already lost roughly one-in-five members within two years, and by the ten-year mark, well over a third of a given year’s web is gone.

Institutions rot at different rates

Pew’s second pass looked not at whether pages themselves survive, but at whether the links inside surviving pages still work. Sampling government, news and Wikipedia pages as of spring 2023: 23% of news webpages contained at least one broken link, as did 21% of government webpages — with city-level government sites worst hit (29% of pages with at least one broken link, 13% of all links checked) and state government sites least affected (15% of pages, 4% of links). Wikipedia’s “References” sections fared worse than either: 54% of Wikipedia pages contained at least one dead reference link. On X (formerly Twitter), the decay is faster still — nearly one-in-five tweets sampled in spring 2023 were no longer publicly visible within three months, three-fifths of that loss coming from suspended, deleted or privatized accounts rather than single deleted posts.

Read together, the two studies bracket the same phenomenon from opposite ends: courts and journals cite the web as if it were print, while Pew’s decade-wide sample shows the underlying material has roughly a ten-year half-life at best. A citation, a reference list or a backlink is only as durable as the page it points to — which is why archiving services built specifically for citation permanence, such as Perma.cc for legal materials and the Internet Archive’s Wayback Machine more broadly, exist as direct institutional responses to these numbers rather than general-purpose conveniences.

Methodology

Sources
Jonathan Zittrain, Kendra Albert & Lawrence Lessig, “Perma: Scoping and Addressing the Problem of Link and Reference Rot in Legal Citations,” originally 127 Harv. L. Rev. F. 176 (2014), republished in Legal Information Management 14 (2014), pp. 88–99 — footnote hyperlinks pulled from the Harvard Law Review, Harvard Journal of Law and Technology and Harvard Human Rights Journal (collection date September 7, 2012) and from all published U.S. Supreme Court opinions through October Term 2011, tested by automated HTTP status check plus manual content verification on a 5%-margin-of-error spot-check sample. Raizel Liebler & June Liebert, “Something Rotten in the State of Legal Citation,” Yale Journal of Law & Technology, cited within the above paper, covering Supreme Court citations 1996–2010. Pew Research Center, “When Online Content Disappears,” published May 17, 2024 (authors: Athena Chapekis, Samuel Bestvater, Emma Remy, Gonzalo Rivero) — a random sample of 999,989 URLs from the Common Crawl web repository (~90,000 per year, 2013–2023), rechecked October 2023; a separate sample of roughly 500,000 U.S. government webpages (Common Crawl March/April 2023 crawl, ~4.2 million links analyzed) and news webpages identified via comScore domain data; Wikipedia “References” links sampled via the Wikimedia Foundation archive; and a real-time sample of X/Twitter posts collected spring 2023 and tracked for three months via the Twitter Search API.
Scope limits
The Harvard study’s Supreme Court and journal figures reflect citations collected in 2012, not current practice — courts and journals have since adopted permanent-link tools such as Perma.cc that this data predates and cannot measure. Its 50% Supreme Court figure and Liebler & Liebert’s 29% figure are not directly comparable, since they use different definitions of link failure (content-mismatch inclusive vs. HTTP-error-only). Pew’s government and news samples are U.S.-domain-specific (.gov and comScore-tracked outlets); the study does not report equivalent figures for UK, EU or other national government or news domains. Pew’s “inaccessible” definition is deliberately conservative — pages that still resolve but have changed content, or that are unreadable for accessibility reasons, are not counted as failures, so true content decay is likely understated rather than overstated by these numbers.
Updates
This brief will be revised as new data is published. Corrections are handled under our corrections policy.