A bad redaction just leaked Google's data center water bill

5 min read 1 source clear_take
├── "The redaction failure is a damning indictment of basic operational security in public records workflows"
│  ├── top10.dev editorial (top10.dev) → read below

The editorial emphasizes that the control failed in the most basic way possible — visual occlusion over a selectable text layer. For a signed NDA to be enforced by a redaction workflow this fragile demonstrates that the technical competence protecting these agreements is nowhere near the stakes involved.

│  └── @sensanaty (Hacker News, 75 pts) → view

By submitting the story with the framing 'Improper redaction reveals Google Data Center water and electricity usage,' the submitter foregrounds the mechanical failure of the redaction itself as the newsworthy element, not just the leaked figures.

├── "The incident exposes a broken political economy where NDAs let hyperscalers hide civic-impact data from residents"
│  └── top10.dev editorial (top10.dev) → read below

The editorial argues that the AI data center buildout runs on information asymmetry — cities compete for jobs while operators use NDAs to shield water, power, and cooling inputs that residents need to evaluate the deal. The Lincoln leak accidentally revealed that this entire political economy depends on fragile paperwork rather than any legitimate public-interest balance.

└── "Accidental disclosure is worse for Google than voluntary transparency would have been"
  └── top10.dev editorial (top10.dev) → read below

The editorial points out that Google's sustainability narrative relies on carefully-staged disclosure, so having raw municipal draw numbers leak via a PDF accident undermines their messaging control. Disclosing on their own terms — with context, offsets, and framing — would have produced a far better outcome than reporters pulling the numbers out with Ctrl+A.

What happened

Lincoln, Nebraska is the latest city to learn that PDF redaction is harder than it looks. In response to a public-records request about Google's data center expansion in the city, officials released documents with the sensitive numbers blacked out — except the black boxes were a layer on top of selectable text. Anyone who opened the PDF and dragged a cursor across the 'redactions' could read Google's water and electricity consumption figures verbatim.

The local 1011 NOW report frames the story around the civic question — more questions than answers about what Lincoln signed up for — but the mechanics are the part that should land in every engineering org this week. The city had a non-disclosure agreement with Google covering utility usage. The redaction was the technical control enforcing that agreement. The control failed in the most basic way possible: visual occlusion of a text layer, with the underlying bytes still in the file.

The numbers themselves aren't yet the headline — reporters are still working through them, and Google hasn't publicly confirmed the figures. What's confirmed is that a signed NDA, a FOIA response, and a redaction workflow produced a document where the secret is a `Ctrl+A` away. For a company whose public sustainability narrative depends on carefully-staged disclosure, having the raw municipal draw leak via a PDF accident is a worse outcome than disclosing on their own terms would have been.

Why it matters

The hyperscaler data center buildout for AI is running on an information asymmetry. Cities compete for the jobs and tax base; the operators get NDAs that cover the inputs — water, power, cooling tower blowdown, backup diesel runtime — that would let residents evaluate whether the deal is a good one. The entire political economy of 'we can't tell you how much water the data center uses' assumes the paperwork holds. Lincoln just demonstrated that the paperwork is one intern with Acrobat away from not holding.

This isn't a one-off. The same failure mode has leaked court filings about Facebook ad targeting, prosecutor names in Manafort filings, airline passenger manifests, and NSA surveillance details — all because people treat redaction as a visual operation on something that is fundamentally a data structure. Adobe's own documentation has warned about this for over a decade. Government procurement guidance warns about this. It keeps happening because the workflow — print to PDF, drop black rectangles, email — is faster than the correct workflow, and nobody tests the output by trying to break it.

The second-order effect is more interesting than the leak itself. Once a few of these figures are public, the rest become inferable. Data center water draw scales with cooling architecture and local wet-bulb temperature in fairly predictable ways. Power draw scales with rack density and PUE. If Lincoln's numbers get into the record, you can build defensible estimates for Omaha, Columbus, The Dalles, and every other site where Google has a similar-generation build. Investigative reporters and academics have been trying to triangulate these numbers for years using satellite imagery, utility interconnect filings, and employment data. A verified reference point collapses the error bars across the whole map.

And the political math shifts. When a city council votes on a tax abatement for a facility whose water draw is 'proprietary,' the counterfactual is abstract. When it's 'proprietary, except Lincoln's leaked and it was X million gallons a day,' the counterfactual is a number residents can compare to their own water bill. Hyperscalers have been trading disclosure risk for siting speed; the trade just got more expensive.

There's a developer angle that's easy to miss. Every major cloud provider sells 'sustainability' as a feature of their platform — carbon-aware scheduling, region-level emissions dashboards, water-use intensity metrics in annual reports. Those metrics are produced by the same organizations whose on-the-ground operations are protected by NDAs with local governments. The reported numbers aren't necessarily wrong, but they're the version the provider chose to report, at the aggregation level they chose. If granular municipal data starts coming out through FOIA accidents rather than disclosure programs, the gap between the dashboard figure and the ground-truth figure becomes a research project anyone can run.

What this means for your stack

Two concrete implications, one for the people running infra and one for everyone else.

If you work at a company that signs utility NDAs with municipalities — hyperscalers, colos, large on-prem shops, crypto miners, anyone building AI training capacity — the 'confidential' bucket in your threat model needs to shrink. Assume that any number a city clerk types into a document can end up in a FOIA response, and that the technical control on that FOIA response will be executed by someone who has never heard of PDF layer flattening. The fix is boring and known: redact at the source (delete the text, don't cover it), flatten to a rasterized image before release, and run an automated check that greps the final PDF's text layer for the terms you were trying to hide. If your legal team is still shipping PDFs with vector-text redactions in 2026, that's a bug with a tracking number, not a policy debate.

If you're a developer picking a region or a provider, the usable signal here is that the public story about data center resource use is going to get more honest over the next 24 months whether the operators cooperate or not. Price your infrastructure plans against a world where water-constrained regions start pricing water into colocation contracts and where 'green region' badges have to survive independent audit. us-central1 looks different if Council Bluffs residents get to see the aquifer draw. Phoenix-based AI training looks different if 'we recycle our cooling water' becomes a verifiable claim rather than a marketing one.

For the rest of us: this is a good moment to notice that your own org's redaction workflow — in incident post-mortems you share with customers, in security disclosures, in legal responses — probably has the same bug. Open one. Try to select the black boxes. Act accordingly.

Looking ahead

The interesting question isn't whether Google's Lincoln numbers get officially confirmed; it's whether this becomes the forcing function that pushes hyperscalers toward voluntary municipal disclosure. The current equilibrium — NDAs enforced by hope — is cheaper for operators than real transparency, right up until a FOIA PDF leaks and the story becomes the cover-up rather than the number. Expect the next round of data center tax-abatement negotiations to look different: cities with any technical staff at all will start asking why utility figures need to be secret, and the operators who move to proactive disclosure first will get the better press cycle. The ones still shipping flattened-looking PDFs with live text underneath will keep providing content for stories like this one.

Hacker News 356 pts 458 comments

Improper redaction reveals Google Data Center water and electricity usage

→ read on Hacker News
yeag123 · Hacker News

As other commenters have pointed out, 13.3 million gallons = ~40.8 acre-feet and this being Nebraska, the obvious comparison is agriculture (Corn and Soybeans being the main crop).The average Nebraska farm is 989 acers[1] as of 2022 and uses roughly ~1,200 acre-feet of water (or roughly 390 million

aliasxneo · Hacker News

I lived in a smaller rural town near a Google datacenter. During my career there I had routinely come across some outlandish accusations made about our water/power usage from the locals. But for obvious reasons I could never dispute them. I remember we hired a local who'd grown up in the t

tptacek · Hacker News

So, basically, not a meaningful amount of water at all. Credit to the story authors for taking the time to frame what 13MM gallons actually works out to.

ilyagr · Hacker News

This article seems to focus on one of the less interesting in terms of water usage data centers (for most of us non-locals) because the paper is from Lincoln and that data center is in Lincoln.The article's data center used 13 mil gal of water, but as they point out another one used up more tha

soltanov · Hacker News

Redaction should be a tested security operation, not a visual formatting step. If copy-and-paste, a PDF parser, or OCR can recover the content, the document was never redacted.

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.