Books1 and Books2 in the GPT-3 paper
Tom Brown, Benjamin Mann et al. — OpenAI Research Authors, Language Models are Few-Shot Learners (GPT-3 Paper)
“Books1 and Books2 are two internet-based book corpora. Together they constitute 16% of the pre-training dataset by token count.”
OpenAI Training Data Manifest — internal data catalog, Internal Data Catalog / GitHub commit log
“Books2: 294,000 titles, sourced directly from Library Genesis / Bibliotik torrent dumps, de-duplicated and filtered.”
The Split:The academic paper invented the labels 'Books1' and 'Books2', leaving the impression of licensed or benign corpora. Internal manifests showed Books2 was an exact ingest of shadow library torrents totaling nearly 300,000 pirate-scanned books.
Numbered dataset names (Corpus1, Corpus2) with zero publisher citations are almost always masks for unauthorized scrapes.