Gemini 3.8 Flash Cyber: Google's first model tuned for offense

4 min read 1 source clear_take
├── "Google's move is an honest correction — offensive security capability was already leaking through refusals, and gating it behind terms of use is better than pretending"
│  └── top10.dev editorial (top10.dev) → read below

The editorial argues that every major lab has been quietly losing the offensive-security refusal war for eighteen months, with red teamers laundering prompts as CTFs or homework. Google's decision to gate the capability behind ToS, enterprise contracts, and Vertex AI allowlisting is framed as a more honest posture than continued refusal theater.

├── "Finally — a first-party model that treats security professionals as legitimate users instead of forcing prompt-laundering workarounds"
│  └── @bratao (Hacker News, 1109 pts) → view

The submitter surfaced Google's model card to a 1,109-point thread where roughly half the response is enthusiastic. This camp welcomes a supported SKU with documented offensive-security use cases as a long-overdue acknowledgment that red teamers, detection engineers, and CVE researchers are real customers, not adversaries to be refused.

├── "Oh no — shipping a first-party offensive-security model normalizes and lowers the barrier to real attack tooling"
│  └── @Hacker News commenters (concerned faction) (Hacker News) → view

Roughly half the 627-comment thread reacted with alarm that the same lab which spent two years refusing to write shellcode is now shipping a supported, documented SKU whose stated purpose includes exploit PoCs and agentic recon. The concern is that terms-of-use gating is thin protection once the weights and capability exist inside a mainstream API.

└── "The base 3.8 Flash release is unremarkable — the Cyber SKU is the entire story"
  └── top10.dev editorial (top10.dev) → read below

The editorial dismisses the base 3.8 Flash refresh as routine speed-tier bumps to context, price, and tool-calling — nothing worth writing home about. The genuinely newsworthy shift is Google shipping a first-party, ToS-covered model explicitly fine-tuned for vulnerability analysis, exploit reasoning, and authorized offensive operations.

What happened

Google dropped two models on the same page: Gemini 3.8 Flash, the routine speed-tier refresh, and Gemini 3.8 Flash Cyber, a variant explicitly fine-tuned for security work. The model card describes Cyber as trained on additional data covering vulnerability analysis, exploit reasoning, malware triage, log analysis, and — this is the phrase that made the HN thread light up — "authorized offensive security operations."

The base 3.8 Flash gets the usual bumps: longer effective context, cheaper tokens, better tool-calling. Nothing worth writing home about on its own. It's the Cyber SKU that's the story. Google is shipping a first-party model whose stated purpose includes *doing* security work, not just talking about it — writing detection rules, reasoning about CVE chains, drafting exploit PoCs against known vulnerabilities, and driving agentic tools through recon and triage steps.

Cyber isn't a jailbreak or an uncensored fork — it's a supported, documented, terms-of-service-covered SKU from the same lab that spent two years telling everyone their models refuse to write shellcode. The Hacker News comment thread (1,109 points, top comment ratio you rarely see on a vendor blog post) is split roughly evenly between "finally" and "oh no."

Why it matters

Every major lab has been quietly losing the offensive-security refusal war for eighteen months. Red teamers have been using Claude, GPT-4, and open-weights Llama derivatives for exploit development, obfuscation, and payload generation — usually by dressing the request up as a CTF, a homework problem, or "educational." The refusals were theater. The models still helped, just with more friction and worse outputs because the prompt had to be laundered.

Google's move is to stop pretending the capability is accidental and instead gate it behind terms of use, enterprise contracts, and (per the model card) organization-level allowlisting via Vertex AI. That's a more honest posture. It's also a bet that the regulatory environment in late 2026 rewards labs that build audit trails around dual-use capability rather than labs that ship the same capability with a fig leaf.

The technical claims worth watching, from the model card: Cyber posts a 73.4% on CyberSecEval 3's exploit-development subset (vs. 41% for base 3.8 Flash and 58% for the previously-leading GPT-5.1), and a claimed 2.1× improvement on MITRE ATT&CK technique identification from raw log data. Both are Google's numbers on Google's evals, so treat them the way you'd treat any first-party benchmark — directionally interesting, not gospel. The independent replication that matters will come from the SANS crowd within a few weeks.

The more interesting benchmark to me is the one Google buried in the appendix: on a held-out set of real GitHub Security Advisories from Q2 2026, Cyber correctly identified the root-cause file and function 68% of the time given only the CVE description and the repo. Base 3.8 Flash did it 34% of the time. If that number holds up on independent evaluation, triage workflows at every mid-sized security team are about to get restructured.

The uncomfortable second-order effect: if Cyber can find the bug from the CVE text, it can also find the bug from a zero-day writeup, or from a suspicious commit, or from your own private repo if someone drops the token. The defensive and offensive uses of this capability are the same capability. Google's answer — enterprise-only, audit-logged, org-allowlisted — is the least-bad answer available, but it doesn't make the underlying dual-use problem go away.

What this means for your stack

If you run a SOC, you should be piloting Cyber against your existing SOAR playbooks by end of month. The specific thing to test: replace your current "summarize this alert" LLM call with a Cyber-backed one and measure false-positive rate and analyst time-to-triage. The reported log-analysis numbers suggest a real step change, but the failure modes on your data are what matter. Budget two weeks of A/B on live traffic before committing.

If you run a red team or an internal offensive security practice, the practical shift is that you can now write Cyber into a statement of work without lawyering the AUP. The model is contractually cleared for authorized testing. That removes a genuine friction point — most engagements right now involve either running local models with worse capability or an awkward "we won't ask what you used" conversation with the client. Cyber makes the tool auditable, which is what enterprise clients have actually been asking for.

For everyone else — application developers, platform teams, anyone shipping code that touches user data — assume the attackers are already using an equivalent capability against your stack. The window where "finding the bug requires an expert" was your implicit defense is closing. This means the boring hygiene stuff matters more, not less: dependency pinning, SBOM discipline, actually reading your Dependabot PRs, turning on branch protection, rotating tokens. The cost of a mistake goes up when the search cost for the mistake goes down.

Looking ahead

The next six months will decide whether Cyber becomes a category — Anthropic and OpenAI both have obvious follow-up moves — or whether Google walks it back after the first public incident. My bet is category. The economics favor labs that ship dual-use capability with contracts and audit logs over labs that ship the same capability with refusal-training that experienced users route around in two prompts. The interesting policy question isn't whether these models should exist; it's whether the org-allowlist model scales when every mid-sized company has a security team that wants access. Watch the Vertex AI approval queue.

Hacker News 1138 pts 641 comments

Gemini 3.8 Flash and 3.8 Flash Cyber

<a href="https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash&#x2F;" rel="nofollow">https:&#x2F;&#x2F;deepmind.google&#x2F;models&#x2F;model-cards&#x2F;gemini-3-8-flash

→ read on Hacker News
simonw · Hacker News

The speed combined with the fact that this thing is really good at HTML JavaScript is pretty exciting.Here&#x27;s what I got for 1.8 cents and 13 seconds from the prompt &quot;make me a cool thing in html&quot;:https:&#x2F;&#x2F;gisthost.github.io&#x2F;?6a77bc41a81718c6aaa10d4ab243c59fTranscript her

jampa · Hacker News

I&#x27;ve been using Gemini 3.7 for my personal trip planning app. Across multiple benchmarks, it ranks higher on everything I tried:- Real world knowledge (when a thing opens and closes, the geographic region, historical facts). It&#x27;s also the best at taking a cluster of places and working out

mattlondon · Hacker News

Currently top at https:&#x2F;&#x2F;deepswe.datacurve.ai - beating Opus 5!https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;gemini-3-8-flash shows an intelligence score of 59, the same as Opus 5 medium!Wow - for a flash model this seems to benchmark powerfully. Remains to be seen what it is li

simonw · Hacker News

Pelicans (thinking effort high, medium, low): https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer?url=ht... - high cost 8.9742 centsHere are the 3.7 pelicans for comparison: https:&#x2F;&#x2F;tools.simonwillison.net&#x2F;markdown-svg-renderer.html?u... - high cost 8.4387 cents(I thi

simonw · Hacker News

The most interesting thing about the Gemini models is still their multi-modal support: they accept audio and video input, OpenAI and Anthropic&#x27;s flagships are still image-only.Gemini Flash is also pretty cheap, so it&#x27;s a great family for performing media analysis, like extracting structure

// share this

// get daily digest

Top 10 dev stories every morning at 8am UTC. AI-curated. Retro terminal HTML email.