RoomsLibraryToursPeopleCompareSourcesAbout
Gemini
Exit|3 of 4 in OpenAI × LibGen
paper-trailai·2022 / unsealed 2023

Downloading shadow libraries over residential VPNs

The Public RegisterOctober 2023
OpenAI Policy Submissionregulatory filing, OpenAI Comments to the US Copyright Office
Developing AI systems involves normal web scraping technologies that browse public websites, similar to traditional search engine indexing.
The Private Artifact2022
OpenAI Data Engineering Teamscraping operations team, Internal Slack #scraping-ops
Anna's Archive and LibGen are blocking our AWS IP ranges. Routing through residential proxy pools and rotating user agents to finish the multi-terabyte download.
Exhibit: AuthorsGuild_ECF_812

The Split:OpenAI represented its collection methods to the Copyright Office as routine search-engine web indexing. Internally, engineers treated it as an adversarial scraping operation, rotating residential IP proxies to bypass shadow library rate limiters.

The Tell (How to spot this next time):

Search engine web crawlers respect robots.txt; operations that use residential proxy rotations know their ingestion is illicit.

court pipeDisclosed in Authors Guild evidentiary submissions.