Argues that Let's Encrypt's outage is unremarkable — operators have bad days — but the industry's pattern of building a single point of failure on a free, best-effort service is the structural problem. The expected reaction (sympathy tweets, a post-mortem, zero change) misses that ACME is a standard with multiple conforming issuers and the fix is multi-issuer fallback, not better SLAs from LE.
Points out that cert-manager's default renewBefore of 2/3 validity means sites renewing on a thin margin bleed first during an LE incident, and anything relying on OCSP stapling for revocation freshness compounds the failure. The synthesis frames this as a configuration problem operators control, not just an upstream operator problem.
The submission surfacing the LE status page to 149 points reflects community recognition that an outage at LE — which issues roughly 60% of publicly-trusted web certs — is not a niche operator hiccup but a degradation of a meaningful slice of the public web. The visibility of cert-manager retry storms and ingress controllers falling back to self-signed certs underscores how deeply embedded LE has become.
Let's Encrypt's status page has been amber-to-red for most of June 19, with the operator reporting degraded performance on ACME `newOrder`, slow OCSP responses, and elevated error rates on the v2 API. Reports on Hacker News (149 points and climbing) describe cert-manager controllers in Kubernetes clusters retrying orders for hours, ingress controllers issuing self-signed fallbacks, and CDN edges holding onto soon-to-expire leaf certs while waiting for renewals that never land.
The blast radius is wider than the status page suggests. Let's Encrypt issues roughly 60% of publicly-trusted web certificates, which means a partial outage of its issuance and OCSP infrastructure is functionally an outage of a meaningful slice of the public web. Sites that renew on a tight margin — a common pattern, since cert-manager's default `renewBefore` is 2/3 of validity — are the first to bleed. Anything still relying on OCSP stapling for revocation freshness is the second.
The ACME protocol itself isn't broken. The failure is on a single operator's side: rate-limiters tripping, the database fronting the order pipeline slow under load, OCSP responders falling behind. None of that is novel; LE has had bad days before. What's novel is how much of the production internet now treats one free CA as load-bearing infrastructure with no backup.
The industry's reaction to today will be the same as every prior LE incident: a flurry of "thanks for the free certs, take your time" tweets, a post-mortem next week, and zero structural change. That's the wrong reaction. The lesson isn't that Let's Encrypt is unreliable; it's that you built a single point of failure on top of a free service whose SLA is, charitably, "best effort."
The ACME spec (RFC 8555) is a standard. Let's Encrypt is one of several conforming issuers — Google Trust Services launched a public ACME endpoint in 2023, ZeroSSL has run one for years, BuyPass has one, and Sectigo has an ACME path for paid certs. cert-manager has supported multiple `ClusterIssuer` resources since day one. The technical pieces to fail over from LE to GTS in a matter of minutes have existed for two years. Almost nobody has wired them up.
The reason is the usual one: cert renewal is invisible until it isn't. The CNCF survey work on Kubernetes operators consistently finds cert-manager in the top tier of installed controllers, but operator surveys on whether a second issuer is configured come back near zero. The default install path in every Helm chart and tutorial assumes one issuer. The official cert-manager docs do show how to pin specific certificates to specific issuers, but the failover pattern — "if LE returns errors on `Order`, retry against GTS" — requires custom controller logic or a hand-rolled `Certificate` annotation pattern that almost no one writes.
The OCSP angle is its own mess. Browsers have been quietly walking away from OCSP for years — Chrome stopped doing OCSP soft-fails in 2012, Firefox kept it but neutered the UI — so a slow OCSP responder mostly shows up as latency on TLS handshakes in pinned environments and in server-to-server traffic where Go's `crypto/tls` actually checks staples. The shift to CRLite and short-lived certs (Apple's 47-day proposal, LE's own roadmap toward 6-day certs) is the long-term answer, but "long-term" doesn't help the queue of unrenewed certs piling up in your `cert-manager` namespace right now.
The concrete remediation is two `ClusterIssuer` resources and a small change to how you reference them. Define `letsencrypt-prod` and `gts-prod` as separate ACME issuers, each pointing at the respective directory URL with its own ACME account key. For the `Certificate` resources you actually care about — the ones that would page someone at 3 a.m. — annotate them to use whichever is healthy, and run a small reconciler (a few lines in a CronJob is fine) that flips the annotation when an `Order` has been in `Pending` state for more than, say, 30 minutes. cert-manager will pick up the change and re-issue against the backup.
If you're not ready to run a reconciler, the cheaper move is to pre-issue duplicate certs from a second CA and store them in a Secret your ingress can fall back to — a warm spare, not a hot one, but it survives an LE outage. Cloudflare Origin CA, GTS, and ZeroSSL all let you mint a long-lived cert today; stash it, set a calendar reminder for rotation, move on.
The other thing to audit today: your monitoring. If you only alert on "cert expired," you find out about an LE outage the moment your users do. Add an alert on `cert-manager_certificate_ready_status == 0` with a duration of 1h, and an alert on any `Order` resource in `Pending` for more than 30 minutes. Both are one-line PromQL queries against the `cert-manager` metrics endpoint, and both would have fired this morning before the HN thread did.
The right read of today isn't "Let's Encrypt failed us." It's that the web's TLS layer is now dependent on a free public good operated by a small team, and we've collectively decided that's fine because it's been fine. Multi-issuer ACME is the cert version of multi-region — it's annoying to set up, it's only valuable when something is on fire, and the people who already have it are the people who didn't have a bad day today. The cert-manager team will publish their post-incident guidance, ISRG will publish theirs, and most teams will close the tab and go back to shipping. The teams that don't will spend two hours this week writing a second ClusterIssuer, and the next time this happens — and it will happen — their pagers will stay quiet.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.