RoomsLibraryToursPeopleCompareSourcesAbout
Gemini
Exit|2 of 4 in OpenAI × LibGen
euphemismai·2020

Books1 and Books2 in the GPT-3 paper

The Public RegisterJune 2020
Tom Brown, Benjamin Mann et al.OpenAI Research Authors, Language Models are Few-Shot Learners (GPT-3 Paper)
Books1 and Books2 are two internet-based book corpora. Together they constitute 16% of the pre-training dataset by token count.
The Private Artifact2020
OpenAI Training Data Manifestinternal data catalog, Internal Data Catalog / GitHub commit log
Books2: 294,000 titles, sourced directly from Library Genesis / Bibliotik torrent dumps, de-duplicated and filtered.
Exhibit: ECF 782-4

The Split:The academic paper invented the labels 'Books1' and 'Books2', leaving the impression of licensed or benign corpora. Internal manifests showed Books2 was an exact ingest of shadow library torrents totaling nearly 300,000 pirate-scanned books.

The Tell (How to spot this next time):

Numbered dataset names (Corpus1, Corpus2) with zero publisher citations are almost always masks for unauthorized scrapes.

court pipeAuthors Guild discovery disclosures and investigative reporting.