Newly unsealed legal briefs in the Authors Guild's lawsuit against OpenAI and Microsoft, surfaced this week by Hacker News and published directly by the Authors Guild, reveal that internal communications show top executives at both companies were aware that using mass quantities of copyrighted books to train AI models likely constituted copyright infringement — before that training took place.

The case, Authors Guild v. OpenAI, centers on allegations that OpenAI systematically ingested large libraries of copyrighted works without authorization or compensation to train its large language models, including the GPT series that underlies ChatGPT. According to the Authors Guild's account of the unsealed materials, the briefs cite internal correspondence and testimony indicating that senior figures — not just lower-level engineers — recognized the legal exposure of the book-scraping strategy and discussed it explicitly. The Guild argues this undermines any good-faith or innocent-infringement defense OpenAI might raise, since willful infringement carries statutory damages under U.S. copyright law of up to $150,000 per work, compared to $750 to $30,000 for non-willful violations.

Microsoft's involvement is also highlighted in the unsealed filings. As a major investor in OpenAI and an integrator of its models into products including Bing and Azure OpenAI Service, Microsoft executives are alleged in the briefs to have had visibility into the training data sourcing process. The Authors Guild has not publicly specified exactly which executives are named or what the precise documentary evidence consists of beyond what appears in the brief summaries they released, and the defendants have not yet responded publicly to the unsealed materials.

The scale of the alleged infringement is central to the damages calculus the plaintiffs are pursuing. Publishers and authors have separately argued in related litigation that training datasets such as Books1 and Books2 — corpora believed to contain hundreds of thousands of copyrighted titles — were assembled largely from pirated shadow-library sources including Library Genesis. If individual works can be identified and matched to registered copyrights, the potential per-work statutory damages could produce aggregate liability running into the billions of dollars, though courts have historically been reluctant to apply maximum statutory figures across massive numbers of works.

What most general coverage of this story omits is the supply-chain dimension that sits underneath the legal question: the same shadow libraries alleged to be the source of OpenAI's training books — Library Genesis, Z-Library, and similar repositories — are also among the primary resources that preppers, homesteaders, and off-grid communities have historically relied on for digitizing and archiving technical manuals, out-of-print agricultural guides, and medical references. If litigation outcomes or legislative responses to AI training data ultimately result in aggressive enforcement actions targeting these repositories' mirrors and infrastructure, communities that have built offline knowledge archives sourced from those libraries could find that their existing collections represent the last accessible copies of particular texts — a resilience consideration entirely separate from the AI debate itself.