Digital cultureDigital archives

The great digital black hole: whoever controls the archives can decide which history we will remember

The great digital black hole: whoever controls the archives can decide which history we will remember

We are producing more information than any previous civilisation. But we have entrusted it to servers, platforms and databases controlled by a few players. To erase the past, tomorrow, it may not be necessary to burn a single book.

For centuries history survived thanks to an apparently trivial property of paper: once printed, it was extremely hard to call it back.

A book could be censored.

A newspaper could be seized.

A library could be burned.

A regime could order the destruction of certain texts.

But once thousands of copies had ended up in homes, cellars, archives, universities, libraries and private collections, erasing them all became an almost impossible undertaking.

This is also why today we can reconstruct history.

Historians open books printed hundreds of years ago, compare newspapers, letters, registers, photographs, diaries. They find contradictions. They recover forgotten documents. They compare what a government declared in public with what others were writing at the very same moment.

History did not survive because someone kept one perfect copy.

It survived because there were too many copies for anyone to control them all.

Today we are doing exactly the opposite.

We have turned the largest part of our civilisation's information output into something that can be edited, updated, hidden or deleted by acting on a database.

And we keep calling it an archive.

We have never produced so much history. And it may never have been so vulnerable

Every day we produce a quantity of testimony that a historian a hundred years ago could not even have imagined.

Video. Photographs. Articles. Posts. Podcasts. Discussions. Interviews. Amateur footage. Local news broadcasts. Public records. Entire political debates.

Probably no previous generation has documented its own existence with such an obsession.

And yet there is a paradox.

We could be at once the generation that produces the most documents in history and the one that will preserve the fewest.

This is not just a philosophical hypothesis.

In 2024 the Pew Research Center analysed nearly a million web pages collected between 2013 and 2023.

The result is striking: a quarter of the pages that existed at some point in that period were no longer accessible by October 2023.

Looking only at the pages from 2013, the figure rises to 38%.

In just ten years, more than one page in three had already vanished.

And the phenomenon touches the news as well: Pew found that 23% of the news site pages it examined contained at least one link that no longer worked.

We are not talking about the year 1524.

We are talking about 2013.

The problem is not that the Internet forgets. It is who can decide what it should forget

The phenomenon known as link rot, the gradual death of Internet links, is already enough to worry archivists.

But there is an even more delicate problem.

A paper page can be destroyed.

A digital page can be rewritten.

An article published ten years ago can be edited today.

A headline can change. A paragraph can disappear. A photograph can be swapped. A video can be made private. An account can be deleted. An entire platform can shut down.

And without an independent system that keeps the earlier versions, the reader of the future might not even have a way to know that anything changed.

It is no accident that professional archives treat this problem with the utmost seriousness.

The United States National Archives and Records Administration requires, for example, file integrity checks, checksums, logging of the operations performed on documents and periodic controls, precisely so it can verify that a digital document has not been altered over time.

In other words: a trustworthy digital archive is not simply a folder full of files.

It has to be able to prove that those files are authentic and that no one has changed them.

UNESCO said as much back in 2003 in its Charter on the Preservation of the Digital Heritage: digital heritage is at risk of being lost, and technical and legal tools are needed to guarantee its authenticity and to protect it from intentional manipulation and alteration.

Twenty-three years later, the problem has become immensely larger.

Because a growing part of our collective memory is not kept in public archives.

It is kept by private companies.

What if tomorrow YouTube decided the past costs too much?

Let us run a thought experiment.

YouTube decides that keeping certain videos older than fifteen years online is no longer worthwhile.

Not all of them. Only the ones with few views. The ones that produce no advertising. The ones nobody watches.

From an economic standpoint it might even look rational.

But who decides today which videos will be historically important in 2070?

The speech of a US president watched by fifty million people would probably survive. The World Cup final would probably survive. The great music videos would probably survive.

But the problem is all the others.

The small channel that filmed a local protest. The account of a resident during a war. An interview watched by two thousand people. A regional news bulletin. Amateur footage of a square. A conference that has long been forgotten.

A video that seems insignificant today and that in fifty years could prove that a certain thing was already known, discussed or contested back then.

This is precisely the work of the historian: to discover afterwards what nobody had understood was important before.

The US Library of Congress indeed describes websites as ephemeral, at-risk content: addresses change, content is modified and sites can disappear entirely. This is why specific web archiving programs exist.

But not even these archives can capture everything.

The Library of Congress itself explains that archiving happens through captures taken at specific moments, and that technical limits can prevent the complete preservation of some sites and content.

The Internet Archive, the national libraries and the other preservation projects are therefore an extraordinary line of defence.

But the very fact that they are necessary should make us realise how fragile the thing we call digital memory really is.

Ukraine: what will a historian read in 2070?

Let us take a subject on which, today, propaganda, mutual accusations and the selection of information carry enormous weight: the war between Russia and Ukraine.

In the contemporary public account, 24 February 2022 is inevitably a watershed.

The United Nations General Assembly called the Russian action an aggression against Ukraine and, on 2 March 2022, demanded with 141 votes in favour the withdrawal of Russian forces.

This is part of the historical record.

But it is not the only document a historian will have to consult in order to understand how 2022 was reached.

You have to go back at least to 2014.

To the Russian annexation of Crimea, not recognised by the United Nations General Assembly, which on 27 March 2014 reaffirmed the territorial integrity of Ukraine.

You have to study the birth of the separatist republics. The Russian involvement. The Ukrainian military response. The Minsk agreements. The shelling. The civilian deaths. The opposing versions.

And above all you have to be able to read what was being written at the time, not only what we say about it now.

In 2014 the OSCE directly documented shelling in residential areas of Donetsk and civilian casualties.

On 23 August of that year its observers entered a residential area heavily hit by artillery and found three bodies that, according to residents, belonged to a father, a mother and their son.

Other OSCE reports documented damaged hospitals and civilian buildings, markets that had been hit and further casualties. In several cases the observers could verify the shelling but not attribute responsibility with certainty.

The international press was reporting these episodes too.

In August 2014 the Associated Press, carried by the Guardian, reported civilians killed in the shelling of Donetsk and gathered accounts from residents who believed some rounds came from Ukrainian forces; the same article also noted the presence of separatist rocket launchers in the affected areas and the whole context of the conflict.

Another report from the same period carried accounts that Ukrainian government forces were firing artillery from near a village south of Donetsk, while also documenting accusations aimed at the separatists.

These documents do not prove that one of the two contemporary political narratives is "the true one".

They prove something far more important for our discussion: the past was complicated even while it was happening.

And it is exactly that complexity we have to preserve.

Imagine erasing half of it

Now let us run the most unsettling experiment.

Suppose that in twenty years only the documents about the 2022 invasion remain easily available.

A historian could reconstruct a particular sequence of events.

Now imagine the opposite.

The international archives about the annexation of Crimea and the Russian involvement gradually disappear, while hundreds of accounts about the shelling suffered by the people of the Donbas remain.

The historical perception could change radically.

There is no need to falsify a single document.

It is enough to choose which documents survive.

This is perhaps the most sophisticated form of memory manipulation.

You do not have to invent the false.

You have to remove a part of the true.

And what remains will build the narrative on its own.

Perfect censorship does not delete a sentence. It deletes the context

When we think of censorship we imagine someone taking a black pen and blacking out a line.

But power over information can be far more refined.

If I hold an archive of a million documents and I systematically make a hundred thousand of them disappear, the nine hundred thousand that remain can all be perfectly authentic.

And yet the archive as a whole can tell a completely different story.

This is where the digital changes radically the relationship between memory and power.

The ability to alter documents existed before too.

Governments have always censored. Publishers have always chosen what to print. Historians have always interpreted. Propaganda was not invented by the Internet.

But paper had a property that today we risk underestimating: its decentralisation was built into the medium itself.

Once ten thousand copies of a newspaper had been printed, those ten thousand copies materially stopped being controllable by the publisher.

In the digital world, instead, ten million people can read the exact same file kept on a server controlled by a single organisation.

We have multiplied the readers.

We have not necessarily multiplied the independent copies.

And that is an enormous difference.

Amazon has already shown us what it means

In 2009 something happened that today reads almost as if it had been written on purpose for this debate.

Amazon remotely deleted, from some Kindle devices, digital copies of two books that users had already bought.

The titles?

1984 and Animal Farm, by George Orwell.

The copies had been distributed by a party that did not hold the necessary rights, and Amazon refunded the users. The company later acknowledged that the remote deletion had been a very bad decision, and Jeff Bezos apologised.

It was not an act of political censorship.

But it demonstrated something that until then many consumers had not really grasped: someone on the other side of the network could reach into their digital library and make a book disappear.

Try transferring the same idea into the physical world.

You buy 1984 in a bookshop. You take it home. You put it on the nightstand.

A week later the shop discovers it sold you an edition it was not allowed to sell.

At night someone enters your flat, takes the book, carries it away and leaves on the table the money you had paid.

Technically you have been refunded.

But the conversation about ownership would probably take on a very different tone straight away.

In the digital world we have accepted powers of control that in the physical world we would consider surreal.

The digital is not the problem. Centralised digital is

Here, though, we have to avoid a conclusion that is too simple.

The digital is not necessarily more fragile than paper.

Potentially it is the exact opposite.

A digital document can be copied millions of times without degrading. It can be kept in a hundred countries at once. It can be cryptographically signed. You can compute a hash that lets you check whether even a single bit has been changed. It can be distributed across universities, libraries, private individuals and independent archives.

A system designed this way could produce an incredibly resilient historical memory.

So the problem is not digital versus paper.

The real problem is centralised digital versus distributed memory.

When the original document, the publishing system, the archive and even the mechanism we use to search for it all belong to the same party, our memory inevitably depends on that party's decisions.

Economic. Political. Legal. Technical. Or simply commercial.

We have privatised the memory of our civilisation

And this is probably the question we should start asking ourselves seriously.

YouTube is not a national archive. Google is not a public library. Facebook is not a museum. Amazon is not the universal repository of literature. TikTok is not a historical institute. X is not a state archive.

They are companies.

They can of course decide to keep enormous amounts of data. And they can do an extraordinary job of making it accessible.

But their primary function is not to guarantee the historians of 2126 the ability to reconstruct the society of 2026 faithfully.

Their function is to run a business.

We have therefore carried out, almost without noticing, an unprecedented experiment: we have handed an enormous part of our civilisation's memory to entities that have no historical obligation to preserve it forever.

And we may have made the mistake of confusing two completely different things.

Being online.

Being archived.

They are not the same thing.

The Ministry of Truth might not need to exist

George Orwell imagined a system in which the past was continuously corrected so that it stayed compatible with the present.

It took men. It took offices. It took documents to destroy and reprint.

In the digital world the mechanism could be infinitely simpler.

Edit a page. Delete a video. Close an account. Change a database. Make a copy disappear from search. Let a domain die. Wait.

You do not even need to stop someone from speaking.

It can be enough to make sure that, twenty years later, nobody can find what they said any more.

And this is perhaps the most disturbing difference between traditional censorship and the possible censorship of the future.

Traditional censorship tries to stop people from knowing something today.

Control over digital memory can stop people from knowing, tomorrow, that the thing ever existed.

To burn a million books you need men, trucks, fuel and time.

To erase a million files a single query may be enough.

And this is why the greatest cultural battle of the digital age may not be deciding who has the right to publish.

It may be deciding who has the right to delete.

Because whoever controls information controls the present.

But whoever controls the archives holds something even more powerful: the ability to decide which present will become history.

Sources

Geschrieben von Claudio