[Head's up... this one's an angry post that's probably not totally coherent. But it's from the heart, at least...]
"I walk the corner to the rubble / That used to be a library, line up to the mind cemetery now / What we don't know keeps the contracts alive and movin' / They don't gotta burn the books, they just remove 'em" -- Rage Against the Machine
"The document describes how the scanning company’s 'hydraulic powered cutting machine' would 'neatly cut' books, whose pages would later be 'scanned on high speed, high quality, production level scanners.' Finally, it notes, the scanning company will 'schedule with the recycling company to pick up the completed books.'” -- Washington Post, 29 July 2026.
The reporting from the Washington Post that Anthropic built a massive library of books only to destroy them after scanning their pages hit me pretty hard.
I still have a paperback copy of Steinbeck's Cannery Row, one that fit in my pocket, from my time as a young man two decades ago. I carried that book when I lived on Beaver Island in Lake Michigan, working nights as a bass player in a resort cover band and spending days reading and hiking.
I didn't make much money, but paperback books were cheap, and I started falling in love with writing by reading those books. I later became an academic and the author of several books. I also have built a career teaching university students, with a core pedagogy based on reading and writing.
Music, reading, writing... all feels to me to have been shredded, pulped, in favor of consuming content.
Train-and-Pulp
I wish I could say Anthropic's shredding of hundreds of thousands of books was the single most trenchant example of what I'm calling the contentification of culture. It's not, though. There's too many. But the Anthropic story is hitting a nerve out there, so let's look at it.
As the Post story reports, Anthropic and its competitors (Meta, Google, OpenAI) realized that training their large language models on "internet speak" would result in less-than-compelling synthetic prose.
What to train on, then? Books from the before times (that is, before AI slop) were the answer.
But these companies aren't really interested in books the way I have been, the way so many have. This is not about reading. They're interested in content -- nothing more. The Post story reports Anthropic's efforts as a "content acquisition" project.
Contentification: the reduction of creative activity -- making music, making movies, writing books, creating games -- to content. Load up hundreds of thousands of books into a machine that slices them up, digitizes them, and shreds the paper. That's contentification.
No one will read any of those books. They'll consume synthetic content, instead.
A writer struggles to write a book. A book becomes content. A book becomes pulp. The content gets summarized for the harried consumer. The content gets generated -- a synthetic, semantically ablated text emerges. Someone comments on it -- more content. That content is mined. More content.
It's like the canneries in Cannery Row -- LLMs dip their tails into the creative output of humans. "The figure is advisedly chosen, for if the canneries dipped their mouths into the bay the canned sardines which emerge from the other end would be metaphorically, at least, even more horrifying."
The Inversion
Like I said, I'm not surprised, just sad. I guess I saw this coming.
In 2015 I published an article in the European Journal of Cultural Studies about how Big Data would bring about a reversal. At that time, the mid-2010s, the common trope about data was that data is ubiquitous, but the skill to analyze it -- data science -- was rare. I argued that this would invert: data would become highly protected, and the ability to work with it would become ubiquitous and cheap. The dreams of wealth and power accruing to data scientists would become a nightmare of domination by corporations.
That article came to mind when I read that the train-and-pulp would help Claude write better:
In a January 2023 document, one Anthropic co-founder theorized that training AI models on books could teach them “how to write well” instead of mimicking “low quality internet speak.”
Teach them to write well? How about we teach kids to write well? But no, we're teaching them to prompt well -- at best. At worst, the kids will just get a flow of AI content.
Writing, coding, art -- the making of these has been cheapened. The ability to generate content has never been easier.
Meanwhile, I recall living in a town that refused to fund its local library. Access to books just keeps getting tougher -- not just due to underfunded libraries but also due to attention being undermined.
Try to stop these companies from stealing all the "content" or buying it up cheap, hoarding it, locking us out of the means of creativity. Yeah, I know, they're so rich they'll just settle the lawsuits for billions of dollars, admit no wrongdoing, and plow more money into data centres. Their dream is for us to rent computation from them to speed up the incineration of the planet.
Try to hold on to any creative work prior to the rise of these companies. Try to make something beyond what these companies say you can. Try to think of yourself as anything other than a "content creator." Be a writer. Be a singer. Be a friend. Be a lover.
In the 2000s, corporations commodified our sociality. Now in the 2020s, they are pulping our creativity.
Comments
For each of these posts, I will also post to Mastodon. If you have a fediverse account and reply to my Mastodon post, that shows up as a comment on this blog unless you change your privacy settings to followers-only or DM. Content warnings will work. You can delete your comment by deleting it through Mastodon.
Don't have a fediverse account and you want one? Ask me how! robertwgehl AT protonmail . com
Reply through Fediverse