Gemini's Computer Use model: Playwright with a brain, billed per token

5 min read 1 source clear_take
├── "Computer Use has become a shared category primitive across labs, not just another agent demo"
│  └── top10.dev Editorial (top10.dev) → read below

The editorial argues the real signal isn't that Gemini can drive a browser, but that Google, Anthropic, and OpenAI have all converged on the same interface — screenshot in, structured action out, host-side execution — within twelve months. That convergence indicates the API is solidifying into a category rather than a stunt, even though the labs disagree on tokenizers, training, and pricing.

├── "Vision-only, coordinates-out is the right abstraction because it mirrors how humans use machines"
│  └── Google (Gemini team) (blog.google) → read

Google deliberately constrained the surface to 13 primitive actions (click_at, type_text_at, scroll_document, etc.) with no DOM or accessibility tree access. The argument is that pure vision-in/coordinates-out generalizes the way a human does — to any browser or Android UI the model has never seen — rather than depending on privileged hooks.

├── "Gemini 2.5 Computer Use is currently the state of the art on browser/UI agent benchmarks"
│  └── Google (Gemini team) (blog.google) → read

Google reports 69.0% on Online-Mind2Web, 79.9% on WebVoyager, and 69.7% on AndroidWorld — each beating published scores for Anthropic's Claude Sonnet 4.5 Computer Use and OpenAI's CUA. They also claim ~225ms median action latency, which compounds meaningfully across 40-step task chains where competitors are slower.

└── "Hacker News found the launch significant enough to surface to the front page"
  └── @swolpers (Hacker News, 194 pts) → view

Submitted the Google announcement and drove it to 194 points with 124 comments, signaling that the developer community treats this as a notable release rather than routine model marketing. The traction itself is evidence that practitioners see Computer Use as a meaningful step in the agentic-browser space.

What happened

Google released Gemini 2.5 Computer Use, a specialized variant of the 2.5 Pro family built for one job: looking at a screenshot of a browser (or Android UI), deciding what to do, and emitting a structured action. The model is available today in public preview via the Gemini API and Vertex AI, and exposed through a new `computer_use` tool in the SDK.

The surface area is deliberately narrow. The tool defines 13 actions — `click_at`, `type_text_at`, `scroll_document`, `drag_and_drop`, `key_combination`, `navigate`, `wait_5_seconds`, and a handful of others. You hand the model a screenshot plus a goal, it returns a function call, your harness executes it, you screenshot again, loop. That's the entire contract. There is no DOM access, no accessibility tree, no privileged hook into the browser — it's vision-in, coordinates-out, the same way a human operates a machine they've never seen before.

Google's reported numbers, from the blog post and a companion technical report: 69.0% on Online-Mind2Web, 79.9% on WebVoyager, and 69.7% on AndroidWorld. Each of these beats the published scores for Anthropic's Computer Use (Claude Sonnet 4.5) and OpenAI's Computer-Using Agent on the same harnesses, and the report claims ~225ms median action latency — meaningfully faster than either competitor, which matters when a task is a 40-step chain.

Why it matters

There's a temptation to file this under "another agent demo," but the framing is wrong. The interesting thing about Computer Use isn't that an LLM can drive a browser — it's that Google, Anthropic, and OpenAI have now all converged on the same primitive within twelve months, which means the API is starting to look like a category, not a stunt. The three labs disagree on tokenizers, training data, and pricing, but they agree on what the interface should be: screenshot in, structured action out, host-side execution. That convergence is the signal.

The second thing worth noticing is what Google chose to optimize. Anthropic's first Computer Use was pitched at desktop automation — a model that could open Excel and shuffle cells. OpenAI's CUA targeted general computer control through a virtual machine. Gemini 2.5 Computer Use is explicitly browser-first, with Android as the secondary target and a note that desktop OS control "is not yet optimized." That's a sharper bet: most economically valuable knowledge-worker tasks happen in a Chrome tab. Salesforce, Workday, Concur, Jira, your bank's portal, the supplier-onboarding form your procurement team forwards once a quarter. None of these have decent APIs. All of them have URLs.

The benchmarks deserve a closer read. Online-Mind2Web at 69% sounds impressive until you remember that a 69% per-task success rate compounds catastrophically across a multi-step workflow — chain four 69% tasks and you're at 23% end-to-end. This is the same failure mode that killed first-generation RPA: brittle automations that work in the demo and fail silently in production. Google's response is the model's tool-call structure plus a "safety service" that can require human confirmation on high-stakes actions (purchases, sending messages, accessing sensitive data). That's the right design — fail-loud, escalate-by-default — but it pushes the integration burden onto the harness you write.

The community reaction on Hacker News (194 points by mid-morning) split along predictable lines. The optimists pointed at the latency number and the fact that Browserbase, the YC-backed headless-browser-as-a-service company, is already shipping a Gemini-powered demo. The skeptics pointed out that none of the demos include the failure modes — what does the model do when a CAPTCHA fires, when a session times out, when an A/B test changes the DOM mid-run? Google's response, implicit in the docs: that's your problem. The model emits an action; you decide whether to trust it.

What this means for your stack

If you're currently maintaining Playwright or Puppeteer scripts to drive third-party SaaS tools that refuse to ship a usable API, this is the rung above what you have. The pragmatic migration path isn't 'replace your scripts' — it's 'wrap the brittle 20% of your scripts that break every time a vendor reskins their UI in a Gemini Computer Use loop, and keep the deterministic 80% as-is.' You get the resilience of vision-based interaction where you need it, and the speed and cost of deterministic selectors where they still work. The model returns coordinates; your existing harness can still own the network layer, the session management, the cookie jar.

For anyone building an internal agent platform — and roughly every Fortune 500 has now greenlit one — this changes the build-vs-buy math on the browser-automation layer. The previous answer was "build a Playwright wrapper, accept that 30% of your engineering time will be spent on selector maintenance." The new answer is "call the Computer Use tool, accept that 30% of your engineering time will be spent on agent-loop orchestration, retries, and confirmation policy." The work doesn't disappear; it moves up the stack into territory that's more interesting and arguably more durable.

Pricing is the variable to watch. Google hasn't published Computer Use pricing separately from 2.5 Pro yet, but a 40-step browser task at 2.5 Pro's vision token rates is not trivially cheap. For high-volume automation, the per-action cost will dominate the engineering cost within weeks, and the build-vs-buy line will move accordingly. If a Concur expense submission costs $0.40 in Gemini tokens and your sales team files 50,000 a year, you've just spent $20K to save what an offshore BPO would have done for $35K. The economics get more interesting at higher-skill tasks where the labor cost being displaced is six figures.

Looking ahead

The honest read: Computer Use models are now table stakes for frontier labs, and the next 18 months will be a benchmark war fought over per-action latency, multi-step success rates, and how aggressively the safety layer interrupts the loop. The labs that lose will be the ones whose models fail silently — emitting a confident `click_at(450, 320)` on the wrong button and proceeding as if nothing happened. The labs that win will be the ones who solve the harness, not just the model: confirmation policies, retry semantics, and a clean way to hand control back to a human without losing state. Google has the model. Whether they ship the harness — or whether Browserbase, Anthropic's Bedrock partners, and a wave of agent startups ship it for them — is the next question.

Hacker News 237 pts 160 comments

Computer use in Gemini 3.5 Flash

→ read on Hacker News
smallstepforman · Hacker News

Today I asked Gemini to extract a table from an PDF appendix and create C++ data table with its contents. After 15 or so iterations with corrections and new mistakes, it eventually gave up. I was floored when it said “I’m sorry, I cannot do this simple task, I’ve exceeded my error threshold and cann

jorjon · Hacker News

Gemini Flash 3.5 (through agy) ran `git reset --hard` when I asked it to commit my changes, apparently it thought it was better to have a clean repo before `git add`. Of course I'm not trusting my computer to it. When will we have 3.5 Pro?

satvikpendem · Hacker News

There's still no MCP support in the Gemini app, which is very useful to get various pieces of info as a user just via chatting. For example I recently wanted to get an Airbnb and wanted to filter by specific criteria including house image analysis and Gemini couldn't do it so I had to do i

mlmonkey · Hacker News

It's funny how in their own graph, https://storage.googleapis.com/gweb-uniblog-publish-prod/ima... Gemini 3.5 Flash is beat hands down by both Opus 4.8 and GPT 5.5, and yet the graph is drawn as if Gemini wins ... :-D

YuechenLi · Hacker News

So... has Google provided a Codex/Claude Code equivalent to Gemini yet? I would like to use Gemini for coding tasks, but that's kind of difficult to do as I don't even know how to get Gemini to even "clone this repo and read the code in it for static analysis", much less ope

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.