Kent Beck: we didn't hire you to complete tasks

5 min read 1 source clear_take
├── "Task completion was always a proxy for judgment, and AI has exposed the proxy"
│  ├── Kent Beck (kentbeck.com newsletter) → read

Beck argues companies never actually hired engineers to complete tasks — they hired them to exercise judgment about which tasks matter, when specs are wrong, and when to push back. The task was always the artifact produced on the way to demonstrating judgment, and now that LLMs make the artifact cheap, the underlying judgment gap is visible.

│  └── @rrvsh (Hacker News, 168 pts) → view

Submitted Beck's essay to HN where it reached 168 points, signaling broad agreement among senior engineers that ticket-closing metrics have always been a weak proxy for the real deliverable. The submission's traction reflects an industry quietly acknowledging this gap.

├── "AI-augmented mediocre engineers will dominate metrics-driven performance reviews while thoughtful engineers lose"
│  └── Kent Beck (kentbeck.com newsletter) → read

Beck warns that any performance rubric counting tickets closed, lines shipped, or PRs merged is now a rubric an LLM-augmented mediocre engineer will dominate. Engineers who exercise judgment — pushing back, rewriting specs, throwing out tickets — will appear less productive on paper despite producing more actual value.

├── "Junior engineers are caught in a contradiction between what they were hired to do and what actually matters"
│  └── top10.dev editorial (top10.dev) → read below

The synthesis highlights that HN comments show junior engineers reasonably asking 'then what was the offer letter for?' — they were recruited on task throughput but are now told judgment is the real currency. This creates a structural unfairness: juniors haven't had time to develop judgment, and the AI tools that boost their throughput don't compensate for that gap.

└── "AI productivity studies inadvertently prove task completion was never the bottleneck"
  └── top10.dev editorial (top10.dev) → read below

The editorial argues that every study boasting 30-55% faster task completion implicitly admits task completion wasn't the constraint — otherwise we'd be shipping 30-55% more value, which nobody claims. The real bottleneck has always been deciding what to build and noticing when the spec is wrong, neither of which AI meaningfully accelerates.

What happened

Kent Beck — the guy who wrote the book on Extreme Programming, helped invent JUnit, and has spent four decades watching software teams measure the wrong thing — published a newsletter post titled *Hey, n00b, we didn't hire you to complete tasks*. The post hit 168 on Hacker News, which for a Beck essay is roughly the baseline; the comments are the interesting part. They're full of senior engineers nodding and junior engineers asking, reasonably, *then what was the offer letter for?*

Beck's argument, compressed: companies have spent decades hiring engineers and then handing them tasks, as if the task were the deliverable. It was always a useful fiction. The actual deliverable was judgment — figuring out which task to do, whether the task as written is the task that should be done, when to push back, when to ship the 80% version, when to throw the ticket out and rewrite the surrounding system. The task was the artifact you produced on the way to demonstrating the judgment, not the thing the company was paying for.

What changed is that the artifact got cheap. A junior with Claude Code or Cursor and a decent spec can close tickets at a rate that would have looked like fraud in 2022. The throughput is real. The judgment isn't necessarily there yet. And Beck is pointing at the uncomfortable consequence: if your performance review counts tickets closed, lines shipped, or PRs merged, you have just built a rubric that an LLM-augmented mediocre engineer will dominate, and a thoughtful one will lose.

Why it matters

The industry has been dancing around this for two years and mostly chosen not to look at it directly. Every "AI productivity" study that brags about 30-55% faster task completion is implicitly admitting that task completion was never the bottleneck — if it had been, we'd be shipping 30-55% more value, and nobody is claiming that with a straight face. The bottleneck was, and remains, deciding what to build, noticing when the spec is wrong, and catching the second-order consequences before they ship.

This matters most for juniors, and Beck is honest about why. The traditional apprenticeship model assumed a junior would spend two or three years doing tasks — implementing things a senior had already decided needed implementing — and would absorb judgment by osmosis. You'd watch the staff engineer reject a ticket, ask why, and slowly build a mental model of *which fights are worth picking*. The AI tools have just removed the rung of the ladder where that absorption happens, because the tasks that used to take a junior a week now take an afternoon, and the senior doesn't bother explaining why anymore.

Compare two performance-review frameworks under this lens. The Google-style leveling rubric ("scope," "impact," "complexity") was always a judgment rubric in disguise — it just smuggled the judgment in as "scope" and let managers eyeball it. That framework gets *more* important in the AI era, not less, because the only thing distinguishing an L4 from an L5 is now almost entirely the judgment layer. By contrast, the ticket-throughput dashboards that a lot of mid-sized engineering orgs quietly run on — Jira velocity, PR counts, story points closed — are now actively misleading. They reward the exact behavior that's been commoditized.

The HN comment thread surfaced the uncomfortable corollary: a lot of the engineers who built their careers on being *reliable task completers* — show up, take the ticket, ship the ticket, repeat — are looking at a market that no longer prices that skill. Several commenters in the 10+ YOE range admitted they were rebuilding their identity around "the person who decides what gets built" rather than "the person who builds it," and finding that the second skill is much harder to demonstrate in a 45-minute interview loop.

What this means for your stack

Three concrete moves, if you run or sit on an engineering team.

First, audit your performance review template against the Beck test. Read each criterion and ask: *could an LLM-augmented engineer max this out without doing any judgment work?* If the answer is yes for more than two criteria, the rubric is now measuring noise. "Closes tickets quickly," "writes clean code," "good test coverage" — all of these now describe table stakes, not differentiators. Replace them with criteria that specifically measure judgment: tickets *rejected* with a good reason, scope changes proposed and accepted, bugs caught at design review instead of in prod.

Second, if you're a junior or mid-level reading this and feeling defensive: the honest read is that the apprenticeship contract has changed without anyone signing a new one. The tasks are no longer the training. You have to deliberately seek out the judgment work — sit in on design reviews uninvited, ask why a ticket was scoped the way it was, write the doc that explains *why* a system is the shape it is. The Twitter timeline version of this advice is "learn to think," which is useless. The concrete version is: spend less of your AI-saved time on the next ticket and more of it on understanding the *system* the tickets are coming from.

Third, for hiring: the leetcode-and-take-home loop was always a task-completion test. It's now a particularly bad task-completion test, because the candidates with the best AI workflow will smoke it and the candidates with the best judgment will look mediocre because they spend the interview asking clarifying questions. Several engineering leaders have quietly switched to interview formats built around "here's a half-broken spec, talk me through what's wrong with it," which is a judgment test in plain sight. Worth stealing.

Looking ahead

Beck has been early on this kind of thing before — he was writing about feedback loops and small batch sizes when the industry was still arguing about whether unit tests were a fad. The prediction he's effectively making is that within two or three years, "completed tasks" will be removed from performance review templates the same way "lines of code" was quietly removed in the late 1990s — not because anyone announces it, but because everyone got embarrassed to keep measuring it. The teams that figure out how to measure judgment instead will pull ahead. The teams still running velocity dashboards will spend the next decade wondering why their best engineers keep leaving.

Hacker News 184 pts 95 comments

Hey, n00b, we didn't hire you to complete tasks

→ read on Hacker News

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.