AI Companies Are Buying Tons of Old Books Because They're Free of AI Slop (404media.co) 114
An anonymous reader quotes a report from 404 Media: As AI companies search for more training data to improve their models, one company is offering old, printed books as an ideal source because they are guaranteed to be free of the very AI slop AI companies are producing. "The world's best AI training data is sitting on a shelf," ISBNdb, a company that produces what it claims is "the world's largest book database," and that offers high-volume book acquisition services for AI companies, says on its site. "Books represent curated, peer-reviewed, domain-specific human knowledge, structured in a way no web crawl can replicate. Dense, edited, authoritative."
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."
In one article on its site, ISBNdb explains that printed books published before 2022 are ideal for AI training data because they don't include AI generated text. As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors. The article also notes that book authors who object to their writing being scraped for training purposes can now easily poison AI models by producing writing designed to manipulate and sabotage the resulting AI models.
"Print books from the pre-LLM era are structurally guaranteed to be free of this contamination. That alone is a significant advantage [...] "Physical books published before this date [pre-2022] are structurally clean of modern poisoning tools." [...] ISBNdb advertises that it can keep the identity of AI companies secret. "Strict NDA [non-disclosure agreement] on every engagement," ISBNdb's site says. "Every project begins with a legally binding non-disclosure agreement. Your identity, strategy, and acquisition targets are never disclosed." ISBNdb notes that AI companies may not want to be caught destroying printed books during the scanning process. "The optics problem is real," ISBNdb's site says. "'AI company destroys two million books' is not a headline that generates sympathy."
Rainbows End (Score:5, Insightful)
Yeah, it's happening.
Not the best idea (Score:4, Insightful)
Re: (Score:1)
The idea is to have both. And the training is (partially) split into knowledge (or generally base training) and style training in the post-training step. So you can have a LLM describe a modern computer UI without saying delve into and "why it matters" a hundred times if you train the style on AI style free writing.
Re: (Score:3)
That doesn't make much sense either, the further back, the difference in the very way English is spoken/written.
Honestly this, again, means I get to raise a point that I never got a good reply to at the beginning of this LLM BS: what is the incentive for creating new, clean, correct, works at this point? Because "People always want to argue about shit on Reddit" isn't going to get us new information on the majority of topics with the possible exception of politics. And let's be honest, that's not a thing wh
Re:Not the best idea (Score:4)
- Wikipedia, which has an official stance of ideological filtering
In these post-truth days, impartial facts are considered left wing ideology.
Re: (Score:3)
And so easy to shatter when you look at the far left and see them calling the very idea of impartial fact "a tool of white supremacy".
It's also so easy to shatter when you look at the sky and it's bright green. Fortunately neither of those things happens.
And so easy to shatter when you look at the far left and see them calling the very idea of impartial fact "a tool of white supremacy".
No link, cool. This is probably good, you'd have probably linked to a loony website with a sort of grain of half truth am
Re: (Score:3)
look at the far left and see them calling the very idea of impartial fact "a tool of white supremacy"
It's easy to find extremists of every sort saying mad things.
This nonsense has no effect on Wikipedia's reliable sourcing requirements.
Re: Not the best idea (Score:2)
Having both only means that it will give the modern answer sometimes and the old one sometimes, with no apparent rhyme or reason. That makes it worse because it's unpredictable.
Re: (Score:1)
If I had a huge pile of money to burn, and wanted to be more of a AI goblin than a tech bro, I'd scan only golden and silver age Science Fiction novels, novellas, short stories, etc, and feed it into a LLM... to do something. Maybe try to get it licensed and implemented into work software (or would it be "platforms"? big tech love their "platforms" these days)
Isaac Asimov's "The Feeling of Power" has felt so much more realistic these days. A couple of decades ago I would have gone that the scenario was very
Re: (Score:2)
Yep. In fact a pretty bad idea. But I think they are getting desperate.
Fact Check (Score:1)
Companies are offering books for sale.
Book sellers are selling books.
The sellers seem to have no idea if the buyers are using it for AI training, just that the sales volumes have increased.
Anecdote: I see an increase in book stores. AI also see in increase in youths reading/collecting physical books at a higher rate than recent years. YMMV
$3000 per book (Score:5, Insightful)
https://techcrunch.com/2026/07... [techcrunch.com]
"The payout will deliver $3,000 per work across an estimated 500,000 works, shared among the authors and publishers who hold rights to them. While the settlement is believed to be the largest in the history of U.S. copyright law, many authors and creators still donâ(TM)t view it as a win.
Thatâ(TM)s because of how the legal question was resolved. Alsup sided with Anthropic on the core issue. He ruled that training an AI model on copyrighted text counts as fair use â" a decision widely seen as a turning point for the AI industry. But the ruling didnâ(TM)t excuse how Anthropic obtained the books in the first place. Anthropic had built its training library from two sources: books it purchased and scanned (fine), and books it downloaded from pirate sites like Library Genesis and Pirate Library Mirror. Alsup found the second method illegal on its own terms and said that piracy question could go to trial; Anthropic agreed to a settlement soon after to avoid a trial and whatever damages a jury might have awarded."
So... it's legal to scan books you own and then use them to train LLMs, but it's not legal to use scans that someone else made (I'll assume in this case, they didn't own the books in question.) Hence... the perverse incentive to buy and re-scan books that might already have been scanned... and the cheapest way of doing it is to chop the spine off.
"Internal Anthropic documents about its plan to scan millions of books, revealed in the copyright lawsuit, donâ(TM)t make clear why the company wanted to destroy the books in the process. A deposition of Tom Harvey, who Anthropic hired to lead the project and who previously helped create Google Books, shows that one company Anthropic contracted to scan the books was Datamation, which offers both âoehigh volume destructive and non-destructive book scanningâ services. In a destructive book scanning process, the spine of the book is cut so the pages can be fed into a scanning machine, which is faster and cheaper than non-destructive book scanning.
Regardless of its original intentions, the federal judge in the copyright lawsuit from authors against Anthropic, William Alsup, found that Anthropicâ(TM)s creation of digital copies of the books was legal specifically because the books were destroyed.
âoeHere, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy,â Alsup wrote in his ruling. âoeThe print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company.â "
Kind of fucked up that the scan can't be shared (or donated). I imagine in most cases, good copies of these books no longer exist in libraries or in the Library of Congress. With current law, all books that are covered under copyright will eventually fall into the public domain, but this is meaningless unless copies exist for people to redistribute once that limit is reached. Essentially companies are exploiting the monopoly benefit extended through copyright without allowing society to benefit from the material falling into the public domain, which is the implicit contract to using state power to enforce copyright.
Ironically, destroying physical copies in order to comply with the 1:1 rule makes the remaining copies that much more valuable.
Re: (Score:2)
If, and only if, there were more than a single copy of the book.
Re: (Score:3)
Re: (Score:1)
Can you point to any history which has been destroyed, specifically? The last copy of a historic work?
I can't believe anyone who values information to any degree would do this.
Re: (Score:3)
I can't believe anyone who values information to any degree would do this.
While I agree with you, I can also believe that a percentage of the execs at AI companies do not value information to any degree.
Re: (Score:2)
Fair point.
I also don't see them paying hundreds+ if not thousands of dollars for unique books, per book.
Digitalization can preserve history (Score:1)
Naturally the books would be destroyed, because destroying history to rewrite it is a feature and not a bug for AI companies.
Digitalization can preserve history. Copies of books that are lost as they rot and/or are thrown out by surviving children and grandchildren of the owner is far more common than an AI training copy.
The Library of Congress has a digitalization effort. They try to use non-destructive techniques but sometimes it is necessary. Maybe the AI companies can send copies to the Library of Congress once the copyright expires.
I expect the AI companies need to keep a copy their scans. As optical character recognit
Books are normally destroyed. (Score:2)
The print original was destroyed.
Books are normally destroyed, sure we might be accelerating this. But are we doing so in significant numbers? Even at one print per AI company that kind of pales compared to what will rot on people’s bookshelves and be thrown out by their children or grandchildren.
Imagine if an extra copy of what had been stored in the Library of Alexandria had been digitized? How much more would have been preserved.
Digitalization preserves. Allow the AI companies to make their images public once the copyright e
Destroying more books than the Nazis (Score:2)
Publishers prefer destruction (Score:3, Interesting)
That's why we, the people, must remove their copyright privilege after five years.
Copyright EXISTS to encourage creation not destruction.
Re: (Score:2)
And so it begins.
Publishers hire an army of temporary employees to visit garage sales and swap meets to buy up books to destroy them.
Then, they will turn their efforts against public libraries where there are billions of printed books. They'll offer to buy at first, then they will craft laws to remove public funding then buy them at pennies on the dollar.
At some point, we will have destroyed the Library of Alexandria all over again.
Remove libraries, and the whole world descends into chaos and anarchy.
Re: (Score:2)
Remove libraries, and the whole world descends into chaos and anarchy.
Nah, Encyclopedia Brittanica has an online edition. :-)
Re: (Score:2)
Copyright EXISTS to encourage creation not destruction.
That is so last century. Stop being naive. Copyright exists to generate as much profit as possible. The Constitution is irrelevant now. Over 2 centuries of law being created to subvert the original intentions has created the opposite of what is written. There is nothing to be done about it but let the evil people do whatever the fuck they want. The rest of humanity just doesn't care enough anymore to resist and resisting by yourself is just a pointless death. The USA is over. I nominate the name of "The Kle
Dear ChatGPT, I feel so lazy today... (Score:3)
ChatGPT: According to my knowledge of the latest advancements in Humorism [wikipedia.org], an over-abundance of Black Bile is probably the cause. Thy shall seek to get some cupping done. I learned that from a book.
Humorism (Score:2)
A BIG assumption that those books aren't slop too (Score:1)
Re: Model collapse is a transient phenomenon (Score:1, Troll)
Re: Model collapse is a transient phenomenon (Score:4, Informative)
You apparently have no understood what model collapse is. Here is a hint: It cannot be avoided. The first few papers about it already proved that.
Re: (Score:3)
I've just thought beyond the current state of affairs
Model collapse can be thought of as a kind of statistic "inbreeding" effect. If you can't understand that analogy, I can't help you.
I'm claiming that as AI systems 1) diversify, 2) do chain-reasoning more, including incorporating fresh searches, RAG data etc, 3) use other than LLM model building techniques (perhaps LeCun's JEPA for example), then this will all represent AI systems becoming more diverse in their input, their model bu
Re: (Score:2)
You just confirmed that you do NOT understand what model collapse is.
Re: (Score:2)
You still don't understand what model collapse is.
Re: (Score:2)
Human networks are subject to the same homogenization of popular answers, and we do get a form of model collapse in our collective "common knowledge" but we have diversity of experience and learning and forgetting which leads our individual mental models (from which we communicate) to diversify. This dive
Re: (Score:2)
No, you do not "think" that. You "believe" that because you failed to find out what you are talking about. Model collapse cannot be avoided. Period. This is mathematically proven and no amount of doubting on your part will change anything. The only thing that can be done is create training data using actual intelligence and that is only available in humans. And since slop cannot be reliably identified, that is a very high effort activity once they have stolen and contaminated everything online.
The more I read your anti-AI posts (Score:2)
Re: (Score:2)
I don't think you know what model collapse is either. I don't mean that pejoratively, the whole concept is a muddled mess from the name on.
"Model collapse" isn't anything unique to AI. Any model will drift if it's repeated trained on its own output. Models are all imperfect fits and that imperfection compounds if they're not anchored to actual data. Put a different way, the data processing inequality says you can't create information through computation. It's not "collapse," it's drift.
Once upon a time I wa
Oh the humanity! (Score:4, Funny)
Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.
Re: (Score:2)
Someone's going to feed it a hardcopy of "AI For Dummies" and THEN we'll be sorry. The singularity will be upon us.
Maybe we'll get lucky and that'll make Sam Altman's head explode!
Looks like they are getting desperate (Score:3)
That stunt only works once. Everything to keep the hype going a bit longer, I guess.
Re: (Score:2)
One time might be enough. Even if all AI development would stop now, as in these are the models we got and no new models are created, we could still do a ton of work with the existing models. Much more than what we currently can. We are constantly learning new ways to get better results from lesser models so AI performance would continue to improve, even if the models themselves would not improve.
Like, think of a stupid student who always gets an F in a test. What if you would let him do the test with a te
Re: (Score:2)
You fell for the propaganda.
What do you do for that student? You allow them to bring a summary of the material, that they have written themselves. On paper. I have done that for about 10 years now, with excellent results and high satisfaction levels from the students.
Will this eventually kill chatgpt? (Score:2)
These general knowledge llms train by harvesting the Internet, but the Internet is rapidly being polluted by llms. What are they going to do once the pre-polluted information is too outdated to be useful?
Re: (Score:2)
There are plenty of things they could do.
1. You can hire people to generate content. This has been often done with medical LLM training.
2. You can generate content. This can be done when training LLM to generate code or math. Outside of LLM, it has also been done in medical imaging training.
3. You can make smarter LLM that doesn't need that much training data (this is what Google Deepmind aims to do, not OpenAI, but it is an alternative for them also)
A stopgap measure? (Score:3)
It seems to me that this is only delaying the inevitable. Even aside from deliberate poisoning of LLMs, over time I think they'll still be "polluted" by exposure to misinformation, and maybe even by their own hallucinations.
I'd be happy to hear that I'm wrong, and that we can end up with reliable models. Even at that though, I suspect humans will continue to prompt the models to create slop anyway, because it's profitable and/or strategically useful.
Re: (Score:2)
(scientific method, reputation analysis, interest-assessment, speech-act analysis, epistemic-chain-analysis, bayesian inference,
then apply that to AI.
There's only so much you can do in determining truth (or proximity to it) but if we can do it (some of us, sometimes) then we can almost certainly eventually develop/teach AI to do it too.
Just reverse the tachyon polarity (Score:2)
I'd be happy to hear that I'm wrong, and that we can end up with reliable models.
I believe Encyclopedia Britannica has an LLM trained only on its curated encyclopedia, which went online after they stopped printing books. If you ask the Encyclopedia Britannica site a question it tries to use its LLM. If its LLM cannot answer the question they try ChatGPT. It tells you which LLM provided the answer.
Specialized LLMs can be created for academic disciplines. Trained only on peer reviewed textbooks, papers, etc that are recognized as high quality. That way, given a physics oriented LLM, on
Re: (Score:2)
Thank you - that clarifies things for me. I'm still unclear on a couple of things though. First, do LLMs' behaviour and reliability change over time based on the questions they've processed and the answers they've given? Second, what causes the hallucinations?
What I'm thinking of just now is a Kitboga video on YouTube. In it, he totally breaks the scammers' chatbots and turns them into gibbering idiots merely by making specific requests.
Re: (Score:2)
Re: (Score:2)
Thanks for this - it's good to have a comment that I actually need to study.
How is it that LLMs seem to have initiative and do un-prompted things such as that described here?: https://it.slashdot.org/story/... [slashdot.org] Do they misinterpret some of their training data as an instruction, or is there some other mechanism at work? I wondered if the story was made up to make LLMs seem more powerful than they are, but that seems unlikely and also strikes me as risky.
Re: (Score:2)
Re: (Score:2)
Re: (Score:2)
This joke was old in the 1980s when I first heard it. But strangely, I think it helped me understand things better, to appreciate the complexity and ambiguity, and to have more realistic expectations.
A university was working on two language tra
Re: (Score:2)
Re: (Score:2)
look if reversing the tachyon polarity hasn't fixed it you've got bigger problems than chatgpt
Re: (Score:2)
I'd be happy to hear that I'm wrong, and that we can end up with reliable models. Even at that though, I suspect humans will continue to prompt the models to create slop anyway, because it's profitable and/or strategically useful.
It is possible, but it is not easy, to create a reliable model. Which means, of course, that it is impossible because almost nobody does anything hard. They want everything for free or less. The scams are endless and the integrity is nonexistent.
That's nice and all (Score:3)
Just what the doctor ordered (Score:2)
Chatbots whose ripply muscles and chiseled features ravish the conversation like a trashy romance novel.
Re: (Score:2)
chiseled features ravish
This is beginning to sound like Ayn Rand. She was big on chiseled features and ravishing.
Re: Just what the doctor ordered (Score:1)
If training data weights itself by sheer volume, one might expect the thousand page paperweights she wrote to make the average chatbot sound just a little more objectivist than the average bear.
But Big Brother says Books have always (Score:2)
burning books (Score:3)
Funny twist on the Fahrenheit novel. AI Capitalism buys books to burn them.
Re: (Score:2)
Re: burning books (Score:2)
The only thing destructive scanning is better at, is speed/price.
Told ya (Score:2)
I've said repeatedly that before long you won't be able to trust anything that wasn't printed on paper before say, 2010 or so, maybe even 2000.
Looks that time has arrived.
I hate to say I told you but I did told you...
AI inbreeding (Score:2)
" As the article correctly notes, much of the data that AI companies can scrape from the internet today is likely to include AI generated text, which could result in "model collapse," a process by which AI models that are trained on AI generated data results in worse models that are more prone to errors."
Sounds like AI is subject to the same issues as biological in-breeding.
Re: (Score:2)
More like mad cow disease.
Mein Kampf (Score:2)
Other slop.
what a world (Score:1)
Same deal as Pre-atomic steel (Score:1)
Geesh, How Dystopian! (Score:2)
It is bad enough that we have AI consuming the Internet for training data, which it is not paying for. Then, AI is being used more and more to replace out exactly those who made much of the training content in the first place.
Now, they are buying up all the used books, consuming them and then DESTROYING them because of copyright?! The irony here. Now we can't even buy used books to read what the author actually wrote. Nope, they'll buy them all up and then produce slop out of them.
Very sad and depressing.
Re: (Score:2)
"You dicks! You can't use that without paying for it!"
"Okay, we'll buy books and scan them in"
"You dicks! You can't buy books!"
Geez...
Oh, and the books have to be disassembled in order to be scanned. The binding has to be removed so you can stack the pages in an ADR.
Re: (Score:2)
I think by and large AI will be a wonderful advancement for the species. The problem with AI isn't so much the technology but rather the politics. If we trusted in our government and our business leaders to do the right moral thing, we would be less worried about AI taking our jobs.
We don't have that kind of society. Big business has clearly showed it's not concerned with anything but the bottom dollar. They want to "appear" to be for whatever is trending, regardless if they don't believe it at all. The phr
Re: (Score:2)
And to directly address your post, if they are paying for access, they should be able to access the digital version of it.
I do get that maybe some of these books may not be digitalized but is it really such a high number or is this just cheaper then dealer with the publishers? I imagine it is cheaper and simpler to buy the books and then use them up. It's just unfortunate since the book has to be destroyed for them to scan it. If a person bought it, their is a high chance the book gets passed on to another
Contaminated (Score:2)
AI: The new popular form of warez (Score:1)
While I believe that fair use laws only get you so far for pirating work.... there is a reason that AI companies pay licensing for news and other media. If you personally pirate work, you are more likely to be on the side of AI companies as they are now, though you should understand as the AI companies mature this gravy train will not continue indefinitely.
If you use subscriptions like Apple Music or Spotify, you already should know that these types of platform starve artists where they get an even smaller
!00 Year Old Published Book - Nonsense (Score:2)
A fine example of the nonsense some people will wilfully believe without any supporting data. Occult Chemistry [archive.org]
It was even reviewed in 2013 almost 100 years after being published. [chemistryworld.com]
Not sure if this is the type of content they expect from old books.
The Age Of AI Gray Goo Is Here (Score:2)
Old books, you say? (Score:1)
My parents built up a pretty decent collection of Horatio Alger's highly formulaic rags-to-riches "dime novels" from around 1868-1910. Perhaps AI would like to ingest many, many stories of plucky and hard-working young men who wind up making it big?
What's good for the goose ain't good for the gande (Score:2)
They're being destroyed afterwards. (Score:2)
Which is the bigger problem IMHO.
Re: (Score:2)
Re: They're being destroyed afterwards. (Score:2)
This about predatory and moral-free private enterprises doing destructive scanning for distilling the statistics of the original material into a garbage generator. The book will not be preserved and people wo
Something about copyright? (Score:2)
Re: (Score:2)
Wow (Score:2)
A bit like Steel and Lead from before the nuclear tests.
Don't these folks read /. (Score:2)
Just two posts before this one we have https://yro.slashdot.org/story... [slashdot.org]
Optics solved with wording. (Score:2)
bias (Score:2)
Free of AI slop, full of human bias. Two kinds of poison.
Self referential (Score:2)
This seems to be saying that current AI output has problems that can be addressed by using sources verifiably unpolluted by any AI.
Which ultimately means no AI is better than any AI. For training.
There is no clean-room development environment for AI.
Low-background steel, AI style... (Score:2)
Low-background steel, AI style...
Childern (Score:2)
Probably unenforceable but (Score:2)
What of the future? (Score:2)
As people get exposed to more and more AI output, eventually even those works which are written fully by a human is going to be influenced by AI works.
Wonder what will happen once this becomes a cycle.
AI writing influences human writing which in return influences AI writing .....
Also because they're free of DRM (Score:2)
Old books are made of...paper.
Now AI will destroy our written heritage (Score:1)
Printed books? As in slaughtered trees? (Score:2)
So will they be scanning the printed books to produce pdf files, or whatever input formats the AI bots need? Sounds like loads of work that would keep people busy
Re: Pre-2019 for AI, pre-2010 for Woke (Score:1, Flamebait)
Then they made it law.
Re: (Score:2)
Re: Pre-2019 for AI, pre-2010 for Woke (Score:2)
Being aware of facts is useless and harmful? YOU are useless and harmful, and so is every other anti-intellectual like you.
Re: (Score:2)
That is not true. I did actually read the old encyclopedia in my parent's book shelf, well not the whole book, but the technology parts there as I was very interested in it. Since then I have found much better source for the information and that has been youtube. Yes, youtube is mostly unnecessary things, but as long as you can find it, it contains incredibly detailed information on niche things in a format that is much more easy to understand than it books. Books usually leave a lot of things out and same