The editorial frames Zvezda as the ultimate legacy system — a 25-year-old module running 15 years past its design life, with a critical chokepoint (PrK tunnel) that can't be bypassed because every resupply mission depends on it. The piece argues this is structurally identical to the legacy auth middleware or billing service every senior engineer dreads, where patches fix individual cracks but new ones keep appearing because the underlying fatigue is systemic.
NASA's safety panel rated the Zvezda leaks among the top five ISS risks in its January 2024 report, and NASA's working theory points to high-cycle fatigue from docking micro-vibrations — a structural problem that won't resolve itself. The decision to have the crew shelter in the US segment and seal the Russian hatch reflects an institutional view that the leak rate (now ~1 kg/day, up 5x from 2019) has crossed a threshold warranting active mitigation.
By submitting the BBC live-coverage story to Hacker News where it drew 372 points and 241 comments, the submitter elevated this as a significant developing safety incident worth the technical community's attention rather than routine ISS maintenance news.
Roscosmos publicly disputed NASA's top-five-risk classification and has offered alternative root-cause hypotheses — micrometeoroid strikes, internal pressure differentials, manufacturing defects in 1990s welds — that frame the issue as localized and patchable rather than systemic fatigue. Their position implies the existing seal-and-patch approach to discrete crack sites is adequate, even though new cracks keep appearing.
On June 5, NASA and Roscosmos instructed the seven-person Expedition 73 crew to retreat into the US Orbital Segment and seal the hatch to the Russian segment while repair work proceeded on the Zvezda service module's PrK transfer tunnel — the same 1.6-meter vestibule that has been leaking, in some form, since September 2019. The BBC's live coverage put the current leak rate in the range of ~1 kg of atmosphere per day at the high end of recent measurements, up from the ~0.2 kg/day that triggered the original investigation.
Zvezda is the oldest pressurized module still in continuous use on ISS. It launched in July 2000, was based on a Mir-2 design frozen in the late 1980s, and was supposed to fly for fifteen years. It is currently in year twenty-five. The PrK tunnel connects Zvezda to the Progress cargo docking port, which means you cannot simply weld it shut: every Progress resupply (six per year) and every Soyuz crew rotation passes through it.
NASA's safety advisory panel rated the leaks a top-five ISS risk in its January 2024 report; Roscosmos publicly disputed the severity classification, and as of this week the two agencies still do not share a root-cause hypothesis. NASA's working theory is high-cycle fatigue from micro-vibrations during docking events. Roscosmos has at various points blamed micrometeoroid strikes, internal pressure differentials, and manufacturing defects in welds laid down in the early 1990s. The cracks they have found and sealed — at least four discrete sites — have not stopped new ones appearing.
If you swap "Zvezda" for "the original billing service" and "PrK tunnel" for "the auth middleware everything routes through," this story stops being about space and starts being uncomfortably familiar. The ISS is the most expensive piece of legacy infrastructure humans have ever built, and the engineering pathologies it exhibits in 2026 are the same ones every senior engineer recognizes from any sufficiently old production system.
First, the load-bearing legacy module problem. Zvezda provides life support, propulsion, attitude control, and crew quarters for the Russian segment. You cannot deprecate it without replacing four interlocking subsystems simultaneously, and the replacement — the Russian Orbital Service Station — exists mostly in PowerPoint. The same is true of the database your company can't migrate off because three teams depend on a stored procedure nobody fully understands.
Second, vendor disagreement on root cause. NASA and Roscosmos are running the same hardware in the same vacuum and cannot agree on what's breaking it. This is what happens when the original engineers have retired, the build records are paper-only in two languages, and each side has political reasons to prefer a particular diagnosis. When two vendors share a fault domain but not a root cause, you do not have a debugging problem — you have a governance problem, and no amount of telemetry will fix it.
Third, the patch-rate-vs-failure-rate curve. Roscosmos has sealed multiple cracks since 2019. The leak rate has trended up, not down. This is the canonical signature of a system that has crossed from "isolated defects" into "the substrate itself is failing." In software, this is the point where you stop merging bug fixes and start writing a deprecation plan. NASA appears to have reached that conclusion: in June 2024 it awarded SpaceX an $843M contract to build the US Deorbit Vehicle, the tug that will push ISS into the Pacific around 2030. That contract is the engineering admission that this hardware is not getting another decade.
Fourth — and this is the part worth lingering on — the cost of keeping it flying is now a meaningful fraction of the cost of replacing it. ISS operations run roughly $3B/year on the NASA side alone. Commercial LEO destinations (Axiom, Vast's Haven-1, Starlab) are bidding to take over the science mission for substantially less. The deorbit decision is not driven by a single catastrophic failure; it is driven by the slope of the maintenance-cost curve crossing the slope of the replacement curve. Every CTO who has ever signed off on a rewrite has run the same math.
The practical lesson is not "plan your deorbit." Most production systems do not get a clean retirement; they get federated, encapsulated, or strangled. The lesson is in how NASA is managing the gap between "this thing is failing" and "the replacement is ready." The mitigations they have shipped — hatch-sealing procedures, crew sheltering protocols, atmospheric supply buffers via Cygnus and Progress — are the orbital equivalent of circuit breakers, bulkheads, and capacity headroom on a system you've decided to ride into the ground.
If you are operating a system you know is being sunsetted, three things from the ISS playbook port directly. One: instrument the failure mode aggressively, even when you have no intention of fixing it. NASA measures Zvezda pressure to four decimal places not because they're going to patch every crack, but because they need to know when to evacuate. Two: pre-stage the response. The crew did not improvise the shelter protocol; it was rehearsed. Your equivalent is a runbook for the failure you've decided to accept. Three: be honest about the residual risk with everyone who depends on the system. The astronauts knew the leak rate before they boarded. Your downstream consumers should know the SLA before you cut their access.
The failure mode to avoid is the one Roscosmos is currently demonstrating: defending the system's reputation past the point where the data supports the defense. Every "it's fine, we sealed it" comes at the cost of credibility with the partner who has to make the deorbit call.
The ISS will fly until roughly 2030, leaks and all, because the alternative — uncontrolled reentry of a 420-ton structure — is worse than the controlled risk of operating a degrading module behind a sealable hatch. That tradeoff, made explicitly and reviewed quarterly, is the actual deliverable of mature legacy-systems engineering. The next time someone on your team argues that the old service "still works fine," point them at the ISS: the question is never whether it works, it's whether the failure curve and the replacement curve have crossed yet.
> After multiple inspections and sealant applications, Nasa reported in January that pressure readings suggested a stable configuration had been reached - though there remained uncertainty about whether the leak had truly been sealed or whether air was simply escaping elsewhere.I'm clearly n
Maybe someone who knows more about the ISS than I do can answer this:Naively, I would assume that there are airlocks between the different sections of the ISS. I would also assume that they would close these airlocks while doing the kind of work they are doing to repair the leaks.So, assuming I'
Can't they just get things out of the module and paint it fresh? Maybe with some special paint, or with several layers of a paint?Obviously they can't, it looks like an obvious solution they couldn't have missed. But I wonder why it is impossible to do.
They should keep some FlexSeal up there !
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
I found this interesting: NASA RELL (Robotic External Leak Detector) [1]. "NASA’s Robotic External Leak Locator (RELL) is a robotic, remote-controlled tool that helps mission operators detect the location of an external leak and rapidly confirm a successful repair. … Two instruments working in