One faction in the 265-point HN thread reads the Stanford result as confirmation that frontier LLMs are now beating elite human experts on their home turf. If Stanford Law faculty lose blind-graded head-to-heads on accuracy, clarity, and completeness, the ceiling on AI-displaceable professional work is higher than skeptics admitted.
The counter-camp in the HN discussion argues that 'answering a contracts exam question' is to lawyering what 'writing a leetcode solution' is to building distributed systems. Real legal practice is about client judgment, negotiation, and stakes — none of which a standardized task battery captures.
The editorial argues the paper is more careful than either HN camp suggests: models won decisively on structured-analysis tasks with clear rubrics, but the gap narrowed or reversed on tasks involving ethical ambiguity, client judgment, and strategic recommendation. This sharpens — rather than breaks — the three-year pattern of LLM-vs-expert results.
The editorial draws a sharp methodological contrast: the bar exam result was multiple-choice against a fixed answer key — closer to a benchmark than a reasoning test. Salinas et al. uses blind expert grading of open-ended responses with matched time budgets, making it a far more credible measurement of legal-reasoning parity.
The authors built a task set from law school exams, contract review exercises, and case-analysis prompts, graded blind by a separate expert panel. Critically, they used base models with no retrieval or agentic scaffolding against Stanford Law faculty (not adjuncts) on matched time budgets — eliminating the usual confounds that inflate AI performance claims.
Stanford Law School published *Salinas et al. (2026)* — a controlled study pitting frontier large language models against Stanford law professors on a battery of standardized legal tasks. The result, distilled from the paper: the models outperformed the professors on accuracy, clarity, and completeness across most task categories, with statistically significant margins in the structured-analysis bucket.
The methodology is the part worth reading twice. The authors built a task set drawn from law school exam corpora, contract review exercises, and case-analysis prompts — graded blind by a separate panel of legal experts who did not know which responses came from humans and which from models. Professors got the same time budget as the models' inference window. No retrieval augmentation, no agentic scaffolding — just the base model versus the base professor. The professors are not adjuncts. These are Stanford Law faculty.
The HN thread (265 points) split predictably: one camp reading this as the long-predicted collapse of high-end knowledge work, another camp pointing out that "answering a contracts exam question" is to lawyering what "writing a leetcode solution" is to building distributed systems. Both camps are partially right, and the paper is more careful than either reading suggests.
The last three years of LLM-vs-expert studies followed a pattern: models matched or beat humans on narrow, well-specified tasks; humans won on anything requiring context, judgment, or stakes. The Stanford study doesn't break that pattern — it sharpens it. The models won decisively where the rubric was crisp, and the gap narrowed or reversed on tasks involving client judgment, ethical ambiguity, or strategic recommendation.
Compare this to the *GPT-4 passes the bar exam* result from 2023. That was a multiple-choice exam graded against a fixed answer key — closer to a benchmark than a test of legal reasoning. Salinas et al. is a substantively harder test: open-ended responses, expert graders, rubric dimensions beyond raw correctness. The fact that models still win on this tougher setup is the actual news. It's not that LLMs got better at law; it's that the methodology got more honest, and the result held.
The community reaction worth quoting comes from practitioners, not academics. One HN commenter — a working litigator — put it bluntly: "Eighty percent of associate work is the structured-analysis bucket. The partner work is everything else. This paper just told my firm what to automate and what to keep billing for." That's the productized read. The bearish read, also from a lawyer in the thread: "Wait until one of these models hallucinates a citation in a brief and someone gets sanctioned. We've already seen it." Both are correct simultaneously, which is why this story is a clear_take rather than a multi-viewpoint piece. The conclusion isn't *AI replaces lawyers* or *AI is a toy*. It's that the boundary between automatable and non-automatable legal work just moved, and it moved in a way that maps cleanly onto the existing associate-vs-partner labor split.
For practitioners outside law, the meta-lesson is the one that keeps repeating across domains. Radiology in 2018. Translation in 2020. Junior-to-mid engineering tasks in 2024. Each time, the structured-analysis layer of a profession gets eaten first, the judgment layer holds, and the labor pyramid reshapes around that line. Law is later than most because the data is messier and the malpractice tail is longer — but the pattern is identical.
If you ship anything that touches contracts, terms-of-service review, compliance checks, or policy analysis: the build-vs-buy calculus changed this week. The competitive question is no longer "can an LLM do this?" — it's "what does your wrapper add that a frontier model with a decent system prompt doesn't already do?" For legal-tech startups built on the premise of fine-tuned models beating general models at clause extraction, this is a problem. The Stanford result strongly implies that the general model, with no domain tuning, is already past the relevant threshold.
Practical implications for engineering teams: review your in-house legal automation roadmap. If you're paying outside counsel for first-pass contract review at $400/hour and your security review still requires a human to read every DPA, there's a workflow there that a $20/month model now handles at higher accuracy than the professors who taught those lawyers. That's not a hype claim — it's what the paper measured. The right move is not to fire your GC; it's to push the structured-analysis work down the cost curve and free your GC for the judgment calls where they still win.
For anyone building agentic legal workflows: the bottleneck is no longer model capability. It's retrieval (current law, jurisdiction-specific precedent), citation verification (the hallucination problem is real and unsolved), and the audit trail for malpractice insurance. Those are the real engineering problems, and they're tractable. The model layer is no longer the hard part.
Expect the bar associations to respond within 12 months — first with practice-guidance memos, then with formal rules around AI-assisted work product. Expect at least one BigLaw firm to publicly announce "AI-first associate workflows" by year-end, with the others doing it quietly. And expect the next version of this study to test the harder thing the current paper deliberately scoped out: not whether models can answer a contracts question, but whether they can run a deposition, negotiate a settlement, or read a room. That's the test that actually matters — and it's the one humans will keep winning for a while yet.
<a href="https://law.stanford.edu/wp-content/uploads/2026/06/salinas_et_al.pdf" rel="nofollow">https://law.stanford.edu/wp-content/uploads/2
→ read on Hacker NewsI wonder if this could be explained in a similar way to Hollywood movies. If the movies are designed to please the largest group of people, there is a greater chance people will choose to see it than another movie. The human law professors come with their own personalities, beliefs, and opinions tha
As a software engineer I have some intuition for what the risks are of letting agents do some tasks vs others.I don't have a similar intuition calibrated for what could go wrong when asking AI to draft a legal document. Some things seem harmless, i.e. drafting a will, but I don't really kn
In general it is not surprising. Even if this particular study is bad.There are certain areas of law work that are about analyzing large amounts of texts, drawing conclusions and writing other texts based on that and nothing more. That is literally the bread of LLMs.Those types of lawyers should be
I understand why the conversation on this article looks like it does, but the study is specifically focused on the potential for LLMs to operate as tutors for law students. I enjoy the extrapolation out to whether LLMs will replace lawyers, but did not find that to be discussed in the study itself.I
Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.
I find this study quite suspect. I'd have to dive deeper but there's definitely significant alarm bells that should be going off for anyone reading.Figure 2 (page 6) screams problems. There's only 16 professors (3k comparisons each?!?!) and the professors are all over the place. That&