FTC Faces Pressure to Probe AI Companies Over Mass Destruction of Books

Riya Sharma
7 Min Read

An increasing number of civil society organisations are urging the Federal Trade Commission to look into the major artificial intelligence companies regarding their purchase, scanning, and destruction of physical books that are used in the development of AI systems.

Over twelve groups have asked the FTC to look into whether this practice could breach competition laws, on the grounds that large amounts of books are being acquired by major AI companies, their contents are being digitized and then the physical copies are destroyed in a manner which could harm competition and access to important source material.

The campaign is introducing a further dimension to the growing controversy concerning how AI companies acquire the vast quantities of data required to train their models.

What is the reason that companies which deal with AI are purchasing books?

To train ever more sophisticated AI models, developers need vast quantities of high-quality information. Although companies have in the past placed great reliance on digital materials, physical books can offer valuable data—especially in the case of older works which may include information not available in readily accessible digital databases.

It has been reported that a number of companies have bought large quantities of used and rare books, taken off the bindings, scanned the pages and then got rid of the physical books.

The process enables the information contained in a physical book to be turned into digital training material.

Yet critics argue that destruction becomes a serious issue when the original books are rare, out of print or hard to replace.

The FTC complaint goes beyond copyright

Copyright is already one of the major legal problems that the AI industry is facing.

There has been a dispute among authors, publishers, and technology companies as to whether it is legally possible to use copyrighted works to train AI models without having to ask for permission or pay compensation.

The new request made to the FTC is concerned with a different issue: competition.

The organisations claim that companies possessing huge financial resources might be able to take over large parts of the book market, thus restricting competitors’ access to significant works and keeping valuable training material within their own systems.

It would shift the issue from whether or not an individual book had been legally copied to whether the general business practice could unfairly enhance the market power of the dominant AI companies.

Anthropic’s book-scanning project draws attention

Anthropic has come under special criticism for its attempts to obtain physical books for use in training AI.

Court documents and news reports have referred to an initiative called “Project Panama”, this involving the buying of books, the removal of their bindings and then scanning them for use in AI development; it has been reported that the company eventually got rid of the physical copies after having digitised them.

The practice has come under criticism since some of the books in question are rare or hard to replace.

The controversy has also gone beyond Anthropic.

A probe which involved the use of a tracking device put into a rare book apparently led to the identification of the bulk shipment as having been sent to an Amazon warehouse in Las Vegas; employees of the company stated that the books were having their spines cut off and then being scanned.

Why rare books matter

It might seem fairly insignificant to destroy an ordinary mass-market book if thousands of copies are still available.

Rare books are different.

Certain older editions might have only a few surviving copies. It is difficult or very expensive to replace a physical copy once it has been destroyed.

This has caused concern among booksellers and those who advocate for preservation, as they are afraid that the AI industry’s need for training data might accidentally take important physical materials out of circulation.

It has been reported that some book dealers have started to question very large orders from customers who show an interest in buying thousands of different titles.

AI’s growing appetite for data

The controversy shows that there is a more general issue affecting the AI industry.

Although the most advanced models need extremely large datasets, the amount of readily available high-quality internet data is not endless. In order to develop more capable AI systems, companies are looking for further sources of information.

What books provide is something very valuable, namely a great deal of information that has been created by humans and which is well organised, covering all sorts of areas such as science and history, as well as literature and technical subjects.

That is why AI developers find it appealing — but it also brings up questions as to who should control and gain profit from that information.

What could happen next?

The FTC has not yet announced its intention to initiate an investigation in response to the groups’ request. The organisations are requesting that the agency look into whether the acquisition and destruction of books could amount to an unfair method of competition.

Should the FTC choose to launch an investigation, the case might become a significant test of how U.S. competition law is applied to the swiftly expanding AI industry.

The way in which companies obtain training data in the future could also be affected by the outcome.

At this stage, the debate shows a growing tension lying at the heart of the AI boom, namely that companies need huge quantities of information in order to develop better models, but the struggle for data is now leading more and more questions to be raised about copyright, competition and the preservation of human knowledge.

Since AI companies are still looking for new sources of training data, the way the FTC responds will be a factor in deciding how far they can proceed in obtaining — and in the end destroying — the physical sources of that data.

Total Views: 5,642
Share This Article
Leave a Comment

Leave a Reply

Your email address will not be published. Required fields are marked *