AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop (404media.co) 72
An anonymous reader quotes a report from 404 Media: As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on its site. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."
Rainbows End (Score:5, Informative)
Yeah, it's happening.
Not the best idea (Score:4, Insightful)
Re: (Score:1)
The idea is to have both. And the training is (partially) split into knowledge (or generally base training) and style training in the post-training step. So you can have a LLM describe a modern computer UI without saying delve into and "why it matters" a hundred times if you train the style on AI style free writing.
Re: (Score:3)
That doesn't make much sense either, the further back, the difference in the very way English is spoken/written.
Honestly this, again, means I get to raise a point that I never got a good reply to at the beginning of this LLM BS: what is the incentive for creating new, clean, correct, works at this point? Because "People always want to argue about shit on Reddit" isn't going to get us new information on the majority of topics with the possible exception of politics. And let's be honest, that's not a thing wh
Re: (Score:2)
"keep the stuff AI is killing"
Like what, exactly?
- Reddit, which is a minefield of misinformation and crazy ideologues
- Wikipedia, which has an official stance of ideological filtering, having banned the founder
- StackExchange - hasn't been relevant in years, likely the source of much of the AI coding hallucination and bad agent tooling
- Google - hasn't been a useful search engine since at least 2018, mostly just an ad delivery system now
- not sure what else
Meanwhile, I can still read books published before
Re: (Score:1)
If I had a huge pile of money to burn, and wanted to be more of a AI goblin than a tech bro, I'd scan only golden and silver age Science Fiction novels, novellas, short stories, etc, and feed it into a LLM... to do something. Maybe try to get it licensed and implemented into work software (or would it be "platforms"? big tech love their "platforms" these days)
Isaac Asimov's "The Feeling of Power" has felt so much more realistic these days. A couple of decades ago I would have gone that the scenario was very
Re: (Score:2)
Yep. In fact a pretty bad idea. But I think they are getting desperate.
Fact Check (Score:1)
Companies are offering books for sale.
Book sellers are selling books.
The sellers seem to have no idea if the buyers are using it for AI training, just that the sales volumes have increased.
Anecdote: I see an increase in book stores. AI also see in increase in youths reading/collecting physical books at a higher rate than recent years. YMMV
Re: Pre-2019 for AI, pre-2010 for Woke (Score:1, Flamebait)
Then they made it law.
Re: (Score:2)
That is not true. I did actually read the old encyclopedia in my parent's book shelf, well not the whole book, but the technology parts there as I was very interested in it. Since then I have found much better source for the information and that has been youtube. Yes, youtube is mostly unnecessary things, but as long as you can find it, it contains incredibly detailed information on niche things in a format that is much more easy to understand than it books. Books usually leave a lot of things out and same
$3000 per book (Score:5, Insightful)
https://techcrunch.com/2026/07... [techcrunch.com]
"The payout will deliver $3,000 per work across an estimated 500,000 works, shared among the authors and publishers who hold rights to them. While the settlement is believed to be the largest in the history of U.S. copyright law, many authors and creators still donâ(TM)t view it as a win.
Thatâ(TM)s because of how the legal question was resolved. Alsup sided with Anthropic on the core issue. He ruled that training an AI model on copyrighted text counts as fair use â" a decision widely seen as a turning point for the AI industry. But the ruling didnâ(TM)t excuse how Anthropic obtained the books in the first place. Anthropic had built its training library from two sources: books it purchased and scanned (fine), and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. Alsup found the second method illegal on its own terms and said that piracy question could go to trial; Anthropic agreed to a settlement soon after to avoid a trial and whatever damages a jury might have awarded."
So... it's legal to scan books you own and then use them to train LLMs, but it's not legal to use scans that someone else made (I'll assume in this case, they didn't own the books in question.) Hence... the perverse incentive to buy and re-scan books that might already have been scanned... and the cheapest way of doing it is to chop the spine off.
"Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, donâ(TM)t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both âoehigh volume destructive and non-destructive book scanningâ services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropicâ(TM)s creation of digital copies of the books was legal specifically because the books were destroyed.
âoeHere, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,â Alsup wrote in his ruling. âoeThe print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.â "
Kind of fucked up that the scan can't be shared (or donated). I imagine in most cases, good copies of these books no longer exist in libraries or in the Library of Congress. With current law, all books that are covered under copyright will eventually fall into the public domain, but this is meaningless unless copies exist for people to redistribute once that limit is reached. Essentially companies are exploiting the monopoly benefit extended through copyright without allowing society to benefit from the material falling into the public domain, which is the implicit contract to using state power to enforce copyright.
Ironically, destroying physical copies in order to comply with the 1:1 rule makes the remaining copies that much more valuable.
Re: (Score:2)
If, and only if, there were more than a single copy of the book.
Re: (Score:2)
Re: (Score:2)
Can you point to any history which has been destroyed, specifically? The last copy of a historic work?
I can't believe anyone who values information to any degree would do this.
Re: (Score:3)
I can't believe anyone who values information to any degree would do this.
While I agree with you, I can also believe that a percentage of the execs at AI companies do not value information to any degree.
Digitalization can preserve history (Score:2)
Naturally the books would be destroyed, because destroying history to rewrite it is a feature and not a bug for AI companies.
Digitalization can preserve history. Copies of books that are lost as they rot and/or are thrown out by surviving children and grandchildren of the owner is far more common than an AI training copy.
The Library of Congress has a digitalization effort. They try to use non-destructive techniques but sometimes it is necessary. Maybe the AI companies can send copies to the Library of Congress once the copyright expires.
I expect the AI companies need to keep a copy their scans. As optical character recognit
Books are normally destroyed. (Score:2)
The print original was destroyed.
Books are normally destroyed, sure we might be accelerating this. But are we doing so in significant numbers? Even at one print per AI company that kind of pales compared to what will rot on people’s bookshelves and be thrown out by their children or grandchildren.
Imagine if an extra copy of what had been stored in the Library of Alexandria had been digitized? How much more would have been preserved.
Digitalization preserves. Allow the AI companies to make their images public once the copyright e
Destroying more books than the Nazis (Score:2)
Publishers prefer destruction (Score:4, Interesting)
That's why we, the people, must remove their copyright privilege after five years.
Copyright EXISTS to encourage creation not destruction.
Re: (Score:2)
And so it begins.
Publishers hire an army of temporary employees to visit garage sales and swap meets to buy up books to destroy them.
Then, they will turn their efforts against public libraries where there are billions of printed books. They'll offer to buy at first, then they will craft laws to remove public funding then buy them at pennies on the dollar.
At some point, we will have destroyed the Library of Alexandria all over again.
Remove libraries, and the whole world descends into chaos and anarchy.
Re: (Score:2)
Remove libraries, and the whole world descends into chaos and anarchy.
Nah, Encyclopedia Brittanica has an online edition. :-)
Dear ChatGPT, I feel so lazy today... (Score:3)
ChatGPT: According to my knowledge of the latest advancements in Humorism [wikipedia.org], an over-abundance of Black Bile is probably the cause. Thy shall seek to get some cupping done. I learned that from a book.
A BIG assumption that those books aren't slop too (Score:1)
Re: Model collapse is a transient phenomenon (Score:1, Troll)
Re: Model collapse is a transient phenomenon (Score:4, Informative)
You apparently have no understood what model collapse is. Here is a hint: It cannot be avoided. The first few papers about it already proved that.
Re: (Score:3)
I've just thought beyond the current state of affairs
Model collapse can be thought of as a kind of statistic "inbreeding" effect. If you can't understand that analogy, I can't help you.
I'm claiming that as AI systems 1) diversify, 2) do chain-reasoning more, including incorporating fresh searches, RAG data etc, 3) use other than LLM model building techniques (perhaps LeCun's JEPA for example), then this will all represent AI systems becoming more diverse in their input, their model bu
Re: (Score:2)
You just confirmed that you do NOT understand what model collapse is.
Re: (Score:2)
You still don't understand what model collapse is.
Re: (Score:2)
Human networks are subject to the same homogenization of popular answers, and we do get a form of model collapse in our collective "common knowledge" but we have diversity of experience and learning and forgetting which leads our individual mental models (from which we communicate) to diversify. This dive
Oh the humanity! (Score:4, Funny)
Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.
Re: (Score:2)
Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.
Maybe we'll get lucky and that'll make Sam Altman's head explode!
Looks like they are getting desperate (Score:3)
That stunt only works once. Everything to keep the hype going a bit longer, I guess.
Re: (Score:2)
One time might be enough. Even if all AI development would stop now, as in these are the models we got and no new models are created, we could still do a ton of work with the existing models. Much more than what we currently can. We are constantly learning new ways to get better results from lesser models so AI performance would continue to improve, even if the models themselves would not improve.
Like, think of a stupid student who always gets an F in a test. What if you would let him do the test with a te
Re: (Score:2)
You fell for the propaganda.
What do you do for that student? You allow them to bring a summary of the material, that they have written themselves. On paper. I have done that for about 10 years now, with excellent results and high satisfaction levels from the students.
Will this eventually kill chatgpt? (Score:2)
These general knowledge llms train by harvesting the Internet, but the Internet is rapidly being polluted by llms. What are they going to do once the pre-polluted information is too outdated to be useful?
Re: (Score:2)
There are plenty of things they could do.
1. You can hire people to generate content. This has been often done with medical LLM training.
2. You can generate content. This can be done when training LLM to generate code or math. Outside of LLM, it has also been done in medical imaging training.
3. You can make smarter LLM that doesn't need that much training data (this is what Google Deepmind aims to do, not OpenAI, but it is an alternative for them also)
A stopgap measure? (Score:3)
It seems to me that this is only delaying the inevitable. Even aside from deliberate poisoning of LLMs, over time I think they'll still be "polluted" by exposure to misinformation, and maybe even by their own hallucinations.
I'd be happy to hear that I'm wrong, and that we can end up with reliable models. Even at that though, I suspect humans will continue to prompt the models to create slop anyway, because it's profitable and/or strategically useful.
Re: (Score:2)
(scientific method, reputation analysis, interest-assessment, speech-act analysis, epistemic-chain-analysis, bayesian inference,
then apply that to AI.
There's only so much you can do in determining truth (or proximity to it) but if we can do it (some of us, sometimes) then we can almost certainly eventually develop/teach AI to do it too.
Just reverse the tachyon polarity (Score:2)
I'd be happy to hear that I'm wrong, and that we can end up with reliable models.
I believe Encyclopedia Britannica has an LLM trained only on its curated encyclopedia, which went online after they stopped printing books. If you ask the Encyclopedia Britannica site a question it tries to use its LLM. If its LLM cannot answer the question they try ChatGPT. It tells you which LLM provided the answer.
Specialized LLMs can be created for academic disciplines. Trained only on peer reviewed textbooks, papers, etc that are recognized as high quality. That way, given a physics oriented LLM, on
Re: (Score:2)
Thank you - that clarifies things for me. I'm still unclear on a couple of things though. First, do LLMs' behaviour and reliability change over time based on the questions they've processed and the answers they've given? Second, what causes the hallucinations?
What I'm thinking of just now is a Kitboga video on YouTube. In it, he totally breaks the scammers' chatbots and turns them into gibbering idiots merely by making specific requests.
Re: (Score:2)
Re: (Score:2)
look if reversing the tachyon polarity hasn't fixed it you've got bigger problems than chatgpt
That's nice and all (Score:3)
Just what the doctor ordered (Score:2)
Chatbots whose ripply muscles and chiseled features ravish the conversation like a trashy romance novel.
Re: (Score:2)
chiseled features ravish
This is beginning to sound like Ayn Rand. She was big on chiseled features and ravishing.
But Big Brother says Books have always (Score:2)
burning books (Score:3)
Funny twist on the Fahrenheit novel. AI Capitalism buys books to burn them.
Told ya (Score:2)
I've said repeatedly that before long you won't be able to trust anything that wasn't printed on paper before say, 2010 or so, maybe even 2000.
Looks that time has arrived.
I hate to say I told you but I did told you...
AI inbreeding (Score:2)
" As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors."
Sounds like AI is subject to the same issues as biological in-breeding.
Re: (Score:2)
More like mad cow disease.
Mein Kampf (Score:2)
Other slop.
what a world (Score:1)
Same deal as Pre-atomic steel (Score:1)
Geesh, How Dystopian! (Score:2)
It is bad enough that we have AI consuming the Internet for training data, which it is not paying for. Then, AI is being used more and more to replace out exactly those who made much of the training content in the first place.
Now, they are buying up all the used books, consuming them and then DESTROYING them because of copyright?! The irony here. Now we can't even buy used books to read what the author actually wrote. Nope, they'll buy them all up and then produce slop out of them.
Very sad and depressing.
Contaminated (Score:2)
AI: The new popular form of warez (Score:1)
While I believe that fair use laws only get you so far for pirating work.... there is a reason that AI companies pay licensing for news and other media. If you personally pirate work, you are more likely to be on the side of AI companies as they are now, though you should understand as the AI companies mature this gravy train will not continue indefinitely.
If you use subscriptions like Apple Music or Spotify, you already should know that these types of platform starve artists where they get an even smaller
!00 Year Old Published Book - Nonsense (Score:2)
A fine example of the nonsense some people will wilfully believe without any supporting data. Occult Chemistry [archive.org]
It was even reviewed in 2013 almost 100 years after being published. [chemistryworld.com]
Not sure if this is the type of content they expect from old books.
The Age Of AI Gray Goo Is Here (Score:2)
Old books, you say? (Score:1)
My parents built up a pretty decent collection of Horatio Alger's highly formulaic rags-to-riches "dime novels" from around 1868-1910. Perhaps AI would like to ingest many, many stories of plucky and hard-working young men who wind up making it big?
What's good for the goose ain't good for the gande (Score:2)
They're being destroyed afterwards. (Score:2)
Which is the bigger problem IMHO.
Something about copyright? (Score:2)
Wow (Score:2)
A bit like Steel and Lead from before the nuclear tests.