Microsoft's own exec called AI scraping the 'largest theft of labor' ever

5 min read 1 source clear_take
├── "The unsealed executive quote is a devastating evidentiary shift that undermines the fair-use defense"
│  ├── top10.dev editorial (top10.dev) → read below

The editorial argues this is not a leaked intern message but a sitting Microsoft executive at the company underwriting OpenAI's compute, whose words will be quoted in every future copyright brief against foundation-model labs. Microsoft's public 'transformative fair use' posture is now materially harder to reconcile with its own internal record, and every parallel plaintiff (Getty, authors' guild, music publishers) just gained leverage.

│  └── @pluc (Hacker News, 359 pts) → view

By submitting the TechCrunch story with the incendiary quote as the headline, the submitter frames the internal characterization as the newsworthy core of the filings — treating the executive's language as a smoking-gun admission rather than a throwaway remark.

├── "Procedural context matters — this is a discovery dump, not a merits ruling"
│  └── top10.dev editorial (top10.dev) → read below

The editorial explicitly cautions that no court has ruled on the merits here; the case survived a motion to dismiss but trial is still quarters away. What changed this week is the public evidentiary record and the unsealing signal that the court is losing patience with blanket confidentiality claims — not the underlying legal question of whether training constitutes infringement.

└── "The fair-use fight will be won or lost on market effect, not transformation"
  └── top10.dev editorial (top10.dev) → read below

The editorial frames the defense as resting on two pillars — transformation (a legal argument) and market effect (an empirical one) — and argues plaintiffs have been building the empirical case that models substitute for the copyrighted works they trained on. The unsealed communications about dataset sourcing, Books3 handling, and post-complaint re-ingestion feed directly into that empirical narrative.

What happened

Newly unredacted filings in *The New York Times v. OpenAI and Microsoft* include an internal characterization from a senior Microsoft executive describing large-scale AI training scraping as "the largest theft of labor in human history." The line, reported by TechCrunch on September 17, comes from documents the plaintiffs pushed to unseal after months of redaction fights. It is not a leaked Slack message from an intern or a disgruntled researcher — it is a sitting executive at the company underwriting OpenAI's compute, saying the quiet part in a way that will now be quoted in every future copyright brief filed against a foundation-model lab.

The filings also broaden discovery into internal communications about dataset sourcing, the handling of Books3-style corpora, and whether copyrighted material was deleted, retained, or re-ingested after publisher complaints. The Times has been pushing for exactly this scope since late 2023; the unsealing suggests the court is losing patience with blanket confidentiality claims. Microsoft's public posture — that training on public web data is transformative fair use — has not changed. But the internal record it is now defending in court is materially harder to reconcile with that posture than it was a week ago.

The procedural context matters. This is not a ruling on the merits. It is a document dump inside an ongoing case that already survived a motion to dismiss on the core direct-infringement and DMCA claims. The trial itself is still likely quarters away. What changed this week is the public evidentiary record — and the leverage every other plaintiff (Getty, the authors' guild, the music publishers, the Canadian news consortium) now has in their own parallel suits.

Why it matters

The fair-use defense for training data has always rested on two pillars: transformation and market effect. Transformation is a legal argument. Market effect is an empirical one, and it is the one plaintiffs have been quietly winning. When a Microsoft executive characterizes the underlying activity as theft of labor, they are — inadvertently or not — conceding the market-effect prong. You cannot simultaneously argue that scraping does not substitute for the original and that it constitutes the largest labor expropriation in history. Pick one.

For the labs, the risk is no longer a bad ruling. The risk is a bad settlement number anchored by their own executives' words. The Times case has always been the bellwether because the plaintiff has the resources to litigate for a decade and the archive to prove verbatim regurgitation. A stipulated damages figure here — even one framed as a licensing deal — becomes the floor for every subsequent negotiation. Reddit's $60M/year deal with Google and OpenAI's reported nine-figure arrangements with Axel Springer, News Corp, and the FT already hinted at the market rate. Post-unsealing, those numbers look like the discount, not the ceiling.

The community reaction on Hacker News (359 points and climbing) split predictably but with an interesting middle. The maximalist copyright camp treated the quote as vindication. The training-is-fair-use camp argued the executive was speaking loosely, possibly in a rhetorical or devil's-advocate register that discovery strips of context. The more interesting third position, and the one that should worry practitioners: even if the quote is out of context, the fact that a Microsoft exec was thinking in those terms tells you the internal legal analysis was never as confident as the external messaging.

There is also a widening gap between what the frontier labs do and what everyone else can defensibly do. OpenAI, Anthropic, and Google are cutting licensing deals at scale — Anthropic's Reddit deal, OpenAI's dozen-plus publisher agreements, Google's opaque but expensive arrangements. That path is not available to a 12-person startup training a domain-specific model on scraped forum data. The precedent being set in *NYT v. OpenAI* will apply to that startup with none of the negotiating leverage.

What this means for your stack

If you are building agents that scrape at inference time — the entire browser-agent category, plus most RAG pipelines that hit third-party sites — the exposure surface just widened. The distinction between "training on scraped data" and "retrieving scraped data per query" has always been legally thin, and this filing gives plaintiffs a rhetorical bridge to argue they are the same activity at different points in the pipeline. Robots.txt is not a legal shield; it is a norm. TOS-based claims (see the LinkedIn v. hiQ saga) are back on the table for anyone whose agent ignores a site's stated crawl policy.

Concretely, three things worth doing this quarter. First, audit your training and RAG corpora for provenance: can you produce, for any document your model has seen, a chain of custody showing you had the right to ingest it? Most teams cannot. That is fine when the risk is theoretical; it is not fine when it is discoverable. Second, if you are fine-tuning on user-generated content from platforms like Reddit, Stack Overflow, or GitHub, revisit the platform's current TOS — several were quietly tightened in 2024–2025 specifically to enable licensing revenue. Third, get your DMCA takedown response process written down before you need it, not after.

For the open-source model ecosystem, the implications are darker. Common Crawl, The Pile, and Books3 all sit in a legal gray zone that just got grayer. Models trained on those corpora — a huge fraction of the open-weight ecosystem — carry inherited risk that downstream users have historically ignored. If the court eventually rules that ingestion itself is infringing (as opposed to only regurgitation), the liability walks downstream to anyone who fine-tuned, distilled, or deployed those weights commercially.

Looking ahead

The trial is probably 2027. The settlements will come sooner, and they will reprice the entire data-licensing market. Expect a two-tier world within 18 months: frontier labs with expensive publisher deals and a defensible provenance story, and everyone else negotiating in the shadow of a discovery record that now includes the phrase *largest theft of labor in human history* — spoken not by a plaintiff's lawyer, but by the defense.

Hacker News 793 pts 682 comments

Microsoft exec called AI scraping 'the largest theft of labor in human history'

→ read on Hacker News
thunkshift1 · Hacker News

This will lead to a massive settlement between the big boys and most people who put stuff out in good faith will be left out of it. And that will be the end of it. We will never hear anything about this ever again and the ‘theft’ will continue like normal.

haritha-j · Hacker News

I just don't understand people saying "but a human learning from a book isn't illegal".How do people not understand that some laws only make sense at a certain scale? One human learning from resources and being added to the labour pool is not the same as an infinitely copyable en

47282847 · Hacker News

“Information wants to be free“.It’s not “theft of labor”; the work was already done. If anything it is theft of “intellectual property” (aka “copyright infringement”), if you believe that is a thing, but not of the “labor” that went into it.My personal take: anyone producing content, everyone’s crea

sajithdilshan · Hacker News

If someone asked what is 'the largest theft of labor in human history' I would have thought slavery.

juvvel · Hacker News

I wouldn't have a problem with working off the fruits of other people's labor because most of us are essentially doing that everyday anyway, the issue is that big tech companies (want to) reap all the benefit and create profit from something that should be accessible to everyone. Everythin

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.