TechToDown All articles
Investigative Tech

Stolen Intelligence: The Unregulated Data Heist Powering a $2.8 Trillion AI Industry

TechToDown
Stolen Intelligence: The Unregulated Data Heist Powering a $2.8 Trillion AI Industry

Photo: Captain Andrew M. Freeman, Air Force Institute of Technology / Air Force Research Laboratories (AFRL), Sensors, ATR, Target Recognition Branch., Public domain, via Wikimedia Commons

When a novelist spends three years crafting a book, she expects readers to pay for the privilege. What she does not expect is for a Silicon Valley corporation to quietly ingest every word of it, use it to train a commercial artificial intelligence product, and then generate revenue from that product without paying her a cent — or even asking her permission. Yet that is precisely the situation millions of American creators, publishers, and ordinary citizens now find themselves in.

The artificial intelligence industry, valued at approximately $2.8 trillion globally and growing at a pace that makes Wall Street analysts dizzy, was built on a foundation of data. Enormous quantities of it. And a significant portion of that data was taken without consent, without compensation, and without legal consequence. This is not a fringe allegation made by disgruntled artists. It is a documented pattern of behavior that has triggered dozens of federal lawsuits, drawn the attention of congressional committees, and exposed a regulatory framework so outdated it might as well have been written before the internet existed.

The Scraping Economy

To understand how we arrived here, one must first understand how AI language models are trained. Systems like OpenAI's GPT-4, Google's Gemini, and Meta's LLaMA require what researchers call "training data" — vast corpora of text, images, audio, and code from which the model learns linguistic patterns, factual associations, and generative capabilities. The more diverse and voluminous the data, the more capable the resulting model.

For years, the industry's preferred method of acquiring this data was web scraping: deploying automated bots to systematically harvest publicly accessible content from across the internet. Entire digital libraries were consumed. The Common Crawl dataset, a frequently cited training source, contains petabytes of web content scraped from hundreds of billions of web pages. Books3, a dataset used by multiple major AI labs, contained the full text of nearly 200,000 copyrighted books sourced without authorization from piracy repositories.

The companies involved did not advertise these practices. In many cases, they buried references to training data sources in dense technical papers written for academic audiences. The public-facing narrative emphasized capability — what the AI could do — while the question of where it learned to do it remained conspicuously underexplored.

A Litigation Landscape Taking Shape

The creative community has not remained silent. Since 2022, a cascade of lawsuits has been filed against the industry's leading names, collectively representing one of the most significant intellectual property battles in American legal history.

The Authors Guild, representing thousands of professional writers, filed suit against OpenAI in 2023, alleging that the company's models were trained on copyrighted works without permission or compensation. Among the plaintiffs: John Grisham, Jodi Picoult, and George R.R. Martin — names recognizable to virtually every American reader. The New York Times filed its own landmark suit against both OpenAI and Microsoft in December 2023, arguing that the companies had used millions of the publication's articles to build commercial products that now directly compete with the journalism those articles represent.

Visual artists have mounted parallel challenges. A class-action lawsuit targeting Stability AI, Midjourney, and DeviantArt alleges that AI image generators were trained on billions of images scraped from the web without artists' knowledge or consent. Getty Images filed a separate suit against Stability AI in both the United States and the United Kingdom, citing the unauthorized use of more than 12 million photographs from its licensed archive.

These cases have yet to produce definitive rulings. The legal questions at their core — whether training an AI on copyrighted material constitutes infringement, and whether the output of such a model constitutes a derivative work — are genuinely novel. Courts are navigating intellectual property law that was written for an analog era and has been imperfectly updated for the digital one. The AI industry, for its part, has leaned heavily on the doctrine of fair use, arguing that training constitutes a "transformative" application of copyrighted material. Legal scholars are sharply divided on whether that argument will ultimately prevail.

The Regulatory Vacuum

If the courts represent one arena of accountability, federal regulators represent another — and their record on AI data practices has been, at best, underwhelming.

The Federal Trade Commission, the agency most commonly cited as the appropriate watchdog for AI accountability, has issued reports, opened inquiries, and delivered strongly worded statements. What it has not done is impose meaningful penalties on any major AI company for data scraping practices. The agency's 2023 order requiring OpenAI to implement data security improvements and offer opt-out mechanisms for users was significant in symbolic terms. In practical terms, it left the fundamental business model — training on scraped data — entirely intact.

Part of the problem is structural. The FTC's authority over data practices derives primarily from Section 5 of the FTC Act, which prohibits "unfair or deceptive" business practices. Applying that framework to AI training data requires the agency to demonstrate that consumers were harmed in specific, demonstrable ways — a high bar when the harm is diffuse, delayed, and technically complex. The agency also operates with a budget and technical staff that are modest relative to the resources of the companies it is attempting to regulate.

Congress has held hearings. Senators have asked pointed questions. Legislation has been proposed — including the NO FAKES Act, targeting unauthorized use of individuals' likenesses in AI-generated content, and various data privacy bills that would impose consent requirements on data collection. None of these measures has reached the President's desk. The legislative calendar, lobbying dynamics, and genuine bipartisan disagreement about the appropriate scope of AI regulation have collectively ensured that the status quo persists.

The Personal Data Dimension

Beyond copyright, there is a second category of data that warrants scrutiny: personal information. AI systems have been trained on social media posts, forum discussions, medical forums, and other platforms where ordinary Americans shared information under the assumption that it would remain within a defined context. The concept of "contextual integrity" — the idea that information shared in one setting carries implicit expectations about how it will be used — has been systematically violated by training pipelines that treat the entire internet as a single, undifferentiated resource.

Meta's use of public Facebook and Instagram posts for AI training has drawn particular attention. The company updated its terms of service to clarify that user content could be used for AI development, but critics noted that the notification was buried and that no genuine opt-in mechanism was provided. European regulators, operating under the General Data Protection Regulation, moved to suspend Meta's AI training on EU user data in 2024. American users, lacking equivalent federal privacy protections, received no comparable intervention.

Paths Toward Genuine Accountability

The question of what accountability should look like is not merely academic. Several concrete policy interventions have been proposed by legal scholars, advocacy organizations, and a minority of legislators who have engaged seriously with the issue.

First, a federal data provenance requirement would mandate that AI companies disclose the sources of their training data in a standardized, publicly accessible format. Transparency alone does not resolve harm, but it is a necessary precondition for any meaningful accountability regime.

Second, a statutory licensing framework — analogous to the compulsory licensing system that governs music streaming — could establish a mechanism by which AI companies pay into a collective fund distributed to rights holders whose work was used in training. This approach would avoid the need to litigate millions of individual copyright claims while ensuring that creators receive some measure of compensation.

Third, a genuine federal privacy law with AI-specific provisions would establish consent requirements for using personal data in AI training, with enforcement authority and penalty structures commensurate with the scale of potential violations.

None of these proposals is without complexity. All would require sustained political will that has, thus far, been absent. But the alternative — allowing an industry worth trillions of dollars to continue building its foundational systems on unlicensed, uncompensated, and in some cases deeply personal data — is not a neutral choice. It is a choice to let the most powerful companies in the world write their own rules.

TechToDown will continue monitoring the litigation, legislative, and regulatory developments in this space. The creators, publishers, and private citizens whose work and lives have been incorporated into these systems without their knowledge deserve no less.

All Articles

Related Articles

Fired by an Algorithm: How Platform Giants Are Terminating Workers While Claiming No One Pulled the Trigger