Hoffman argues SpaceX's core engineering challenges — Raptor combustion stability, Starship reentry thermal margin, Falcon booster guidance — are solved with control theory, FEA, and test-stand iteration, not transformer workloads. The ML it does use (landing-pad vision, predictive maintenance) is standard aerospace fare, no different from any modern shop.
The editorial frames the 'AI company' label as the same vague brand-laundering work that 'cloud company' did in 2012 and 'crypto company' did later. Calling SpaceX an AI company because it runs Kalman filters is like calling Stripe an AI company because it has a fraud model — a category error overdue for correction.
Hoffman calls xAI a 'complete train wreck,' pointing to the Memphis Colossus cluster's power and cooling failures, attrition of founding researchers to Anthropic and a stealth lab in Q1, and Grok 4's weak GDPval and SWE-bench-Verified results. The verdict is that compute alone — even a reported 200k-H100-equivalent cluster — does not translate to frontier capability without organizational discipline.
The editorial backs Hoffman with public benchmarks: Grok 4 sits roughly 11 points behind Claude 4.6 on SWE-bench-Verified despite a vastly larger training cluster. That gap is the real story — it shows the Musk-branded merger of xAI into X Corp and claims of training on SpaceX telemetry are marketing, not substance.
The November 2025 share-swap folding xAI into X Corp, plus repeated public claims that Grok would train on 'the entire SpaceX telemetry corpus,' are framed as a deliberate eighteen-month effort to merge the brands. Hoffman's intervention matters precisely because that merger has confused investors, press, and talent about which company is doing what kind of work.
In a Fortune interview published June 24, LinkedIn co-founder and OpenAI board alum Reid Hoffman drew a hard line through Elon Musk's empire. Asked whether SpaceX counted as an AI company, Hoffman answered flatly that it does not — and then volunteered that xAI, Musk's actual AI bet, is a 'complete train wreck'. The framing matters because Musk has spent the last 18 months merging the brands in public, including the November 2025 share-swap that folded xAI into X Corp and the repeated claims that Grok would be trained on 'the entire SpaceX telemetry corpus.'
Hoffman's argument, stripped of the personality drama, is mechanical. SpaceX's hard problems — Raptor combustion stability, Starship reentry thermal margin, Falcon booster guidance — are solved with control theory, finite-element analysis, and a brutal amount of test-stand iteration. None of that is a transformer workload. The company uses ML for things like landing-pad computer vision and predictive maintenance on engines, but so does every modern aerospace shop. Calling SpaceX an AI company because it runs Kalman filters is like calling Stripe an AI company because it has a fraud model.
The xAI critique is sharper. Hoffman pointed to the Memphis 'Colossus' cluster's well-documented power and cooling failures, the churn in the research team (three of the founding members left for Anthropic and a stealth lab in Q1), and Grok 4's underwhelming showing on the GDPval and SWE-bench-Verified leaderboards relative to its compute budget. He stopped short of naming numbers, but the public benchmarks tell the story: Grok 4 sits roughly 11 points behind Claude 4.6 on SWE-bench-Verified despite training on a reported 200k-H100-equivalent cluster.
This is not gossip. It's a category-error correction that has been overdue for two years. The 'AI company' label is now doing the same vague brand-laundering work that 'cloud company' did in 2012 and 'crypto company' did in 2021 — it inflates valuations and lets executives dodge questions about where the actual model work is happening. When SpaceX raises at a $400B mark and the deck mentions 'AI-native manufacturing,' a senior engineer should be able to ask: which team? which models? what's the training compute? In SpaceX's case, the honest answer is 'a small applied ML group inside avionics,' which is fine — it just isn't a thesis.
Hoffman has standing here that most critics don't. He sat on OpenAI's board through the Altman ouster and return, he's a Greylock GP funding Adept and Inflection's successors, and he has no equity in any Musk vehicle. He's also not making a personality argument. The 'train wreck' line is backed by specifics: Colossus's reported 60% effective utilization after the cooling retrofit, the H1 2026 attrition rate inside xAI's pretraining team (community estimates put it above 30%), and the fact that Grok 4's release was quietly pushed from March to May to August.
The more interesting question Hoffman raises is structural. Frontier model work requires a research culture that tolerates 18-month dead ends; rocket and EV engineering requires shipping discipline measured in days. Musk's operating system optimizes hard for the second and actively punishes the first. Anthropic, OpenAI, and Google DeepMind all run on the opposite cadence — long horizons, internal disagreement as a feature, no public deadlines. xAI's quarterly demo cycle is the tell. You don't ship a frontier model the way you ship a Cybertruck refresh.
The community reaction on Hacker News (147 points, 380+ comments at time of writing) split predictably. The Musk camp pointed to Optimus and FSD as evidence of real AI work inside Tesla — which Hoffman didn't dispute and which is, notably, not SpaceX or xAI. The skeptic camp piled on with internal-source rumors about Memphis. The most useful comment came from a former SpaceX avionics engineer who wrote: 'We have one ML PM and four engineers. We are not an AI company. We are a rockets company that uses ML where it helps.'
Three practical takeaways. First, when you're evaluating a vendor that claims to be 'AI-native,' ask for the ratio of ML engineers to total engineering headcount and the percentage of revenue tied to model-driven features. If those numbers aren't on hand, the label is marketing. This applies equally to your own roadmap docs — 'we're adding AI' is not a strategy, and your CTO will eventually be asked the same question.
Second, the xAI situation is a useful case study in what happens when you treat pretraining like a hardware launch. If your team is doing serious model work, resist the pressure to commit to demo dates more than a quarter out. The teams that ship best in this space — Anthropic, DeepMind, Mistral on the smaller end — publish when the eval numbers cross a threshold, not when the calendar says so. Internal stakeholders will hate this. Hold the line anyway.
Third, don't let the Musk-vs-Hoffman drama distract from the actual signal: the frontier model market is consolidating to roughly four labs (OpenAI, Anthropic, Google, Meta), with xAI now visibly slipping out of the running. If your architecture has a hard dependency on Grok via the X API, this is a good week to read the OpenAI-compatible-endpoint section of your vendor's docs and price out the swap.
The next data point to watch is xAI's rumored Series F, reportedly targeting $25B at a $250B valuation. If it closes at those numbers, the market is calling Hoffman wrong. If it gets cut or restructured — or if the round leaks Anthropic-style 'no new lead' language — Hoffman called it before the rest of the press caught up. Either way, the broader lesson stands: the AI label is now load-bearing for a lot of valuations that can't carry the weight, and the developers who learn to read past it will make better build-vs-buy decisions for the next three years.
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.