• The Breakdown
  • Posts
  • đŸŸȘ Does the internet need to be liberated?

đŸŸȘ Does the internet need to be liberated?

AI (and crypto) is putting an information paradox to the test.

“Information wants to be free.”
— Stewart Brand

Does the internet need to be liberated?

Stewart Brand founded the Whole Earth Catalog in 1968 on the idea that everyone should have the information required to live a self-sovereign life. The Catalog was a bi-annual directory of everything you'd need to build your own house, grow your own food, generate your own power, and find your own inspiration.

Steve Jobs later called it “Google in paperback form.”

In 1985, Brand co-founded the Whole Earth 'Lectronic Link, a dial-up internet forum that brought the Catalog’s ethos of information-sharing online. Importantly, the forum enabled a growing community of technologists and counterculturists to exchange ideas, knowledge, and practical know-how, peer-to-peer.

“Information wants to be free,” Brand told an in-person gathering of hackers, who took the words rather literally. The comment became an axiom in the cypherpunk community that later coalesced around an uncompromising commitment to the free flow of information.

The absolutists may have missed some of the nuance, though.

“Information wants to be expensive,” Brand told that same group of hackers.

He did not think he was contradicting himself.

Information “wants” to be free in the sense that “it’s so easy to copy, send and transform that the price tag gets left far behind,” he explained in the LA Times.

But it also wants to be expensive in the sense that “the right piece of news, data or advice at the right time can be beyond price.”

In both senses, he was right. Since the advent of the internet, the course of technology and media has been shaped in large part by the tension between information being 1) immensely valuable and 2) free to copy and move.

Linux vs. the software industry, for example. The Web vs. newspapers. Napster vs. the music industry. BitTorrent vs. the movie industry.

In each case, the ability of valuable information to move freely has shaped the business of creating and distributing it. The result was software-as-a-service, paid blogs, Spotify, and Netflix.

So far.

“That tension will not go away,” Brand predicted way back in 1988. “It leads to endless wrenching debate about price, copyright, ‘intellectual property,’ and the moral rightness of casual distribution, because each round of new devices makes the tension worse.”

He was right about that, too, because the tension is currently reaching record heights now that the “new device” of large language models wants to consume all of the world’s information.

Or all the information that’s publicly available, at least. Because new LLMs need a lot of it.

“AI models require 240 trillion tokens of training data,” Grass founder Adrej Radonjic wrote in a recent blog post. “To put this size into perspective, it is more than 100 times more information than every single published book in human history combined.”

The only place to find that much information is the internet, of course. So Radonjic built the Grass protocol to help AI labs collect it. “Grass is a DePIN node application that routes publicly accessible web data through the idle portion of your internet bandwidth,” the website explains.

In other words, you give Grass access to your home internet, and it borrows your bandwidth on behalf of companies that need to retrieve data from the Web. 

If that sounds to you a bit like a FedEx driver asking to borrow your car to make a delivery while his truck is parked right in front of your house, I get it.

Don’t AI labs have their own access to the internet??

They do, of course — and faster than yours, surely. The problem is that much of the Web is unavailable to AI labs because people don’t want bots scraping their websites.

Publishers, content creators, retailers, and other bot-resistors have developed business models that make information freely available to humans while capturing value elsewhere — business models that may be undermined when the same information is harvested by AI labs.

As a result, the labs increasingly find their bots blocked from collecting the data they need to train new models.

This has enabled Grass to build a booming business of disguising the bots by rerouting them through residential IP addresses (which are never blocked).

It’s sneaky, yes.

But Radonjic frames his project in high-minded terms — his blog post ends with a mission statement straight out of the cypherpunk handbook: “Let’s free the internet.”

That’s not exactly what Stewart Brand had in mind, I don’t think. Brand only said information “wants” to be free, not that it should be. (Not always, at least.)

Either way, the law is currently on Radonjic’s side, because scraping the publicly available Web with bots appears to be legal. 

On the other hand, blocking bots from scraping your website is legal, too (everyone does it). 

And on the other other hand, using a third-party residential IP address, like Grass does, to avoid being blocked is also legal (see, Meta v. Bright Data).

All this suggests that, as Brand predicted, the tension between information being both valuable and free to copy will continue to escalate.

But perhaps we’re getting to a breaking point?

The tension is now so acute that some are forced to argue both sides of the issue. Major AI labs claim the freedom to train their models on information scraped from the Web — but also that competitor labs should not be free to use information from those same models to train their own.

They might be right!

Information wants to be free.

And expensive.

Brought to you by:

Meridian 2026 is Stellar's annual gathering for the institutions, fintechs, and developers putting financial infrastructure onchain.

Join them October 28-29 at Convento do Beato in Lisbon for two days on tokenization, payments, and what it takes to run them in production.

Register today with Blockworks10 for 10% off.