Gemini 3.8 Flash Cyber Limited to Vetted Defenders After 70% Vulnerability Find Rate
Google has restricted Gemini 3.8 Flash Cyber, its new vulnerability-hunting model, to vetted defenders, releasing it only through an application-gated program the company calls Fairwind.
The model arrived on September 2 alongside Gemini 3.8 Flash, a general-purpose sibling that developers and consumers can use without any approval process. The two share an architecture but not an audience.
Background Reading
The split release changes how a frontier lab ships a model trained to find exploitable software flaws. Google put the access control in place on launch day rather than adding one after an incident.
The Fairwind Trusted-Defender Gate
Fairwind limits Gemini 3.8 Flash Cyber to government authorities, critical-infrastructure operators, vetted security researchers and software maintainers. Organizations apply; Google decides case by case.
The company tied the restriction directly to how the model was tuned, writing in the announcement that its safety mitigations were deliberately loosened for security work.
"3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity, and as such, is only available to trusted defenders."
The line comes from a post credited to Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini security lead at Google DeepMind.
Google says the model was built around fixing flaws rather than exploiting them, with additional guardrails against offensive cyber operations and chemical, biological, radiological and nuclear misuse.
Gemini 3.8 Flash Cyber's Benchmark Results
Google published a set of security-specific results for the variant rather than folding it into the general Gemini leaderboard numbers.
The headline figures cover discovery and repair:
- A success rate above 70% finding real vulnerabilities across 20 programming languages
- A 47.2% pass@1 score on CWE-Bench, Google's patching evaluation
- Frontier-level performance on CyberGym, ahead of the earlier 3.5 Flash Cyber
- Measurable gains on Gray Swan tests for prompt-injection robustness
Google describes the patching result as sitting on the Pareto frontier, meaning no cheaper model matched it and no better model cost less at the time of testing.
Benchmarks in this category remain young. CWE-Bench and CyberGym are both narrower than the work a security team actually does, and a 47.2% first-attempt patch rate means the majority of proposed fixes still fail.
Chrome Patches and Cloud Bug Hunting
The more concrete claims come from Google's internal deployments. In Chrome security work, the company reports the model generated 2.6 times more correct patches than larger commercial alternatives.
On Wiz's penetration-testing benchmark, Google measured a vulnerability recall rate 7.5 to 9.7 percentage points higher than competing systems, at between 2.3 and 5.2 times lower cost.
The company also says the model surfaced a critical cloud vulnerability in under two hours, a class of finding that typically takes researchers months.
Those numbers come from Google's own testing and have not been independently reproduced. The Fairwind gate itself makes outside verification harder, since researchers cannot obtain the model without approval.
Pricing and Availability for 3.8 Flash
The unrestricted sibling is the volume product. Gemini 3.8 Flash is the third Flash release in roughly six weeks and targets agentic workloads and software engineering.
It carries a 1 million-token input window and a 64,000-token output limit, and accepts text, images, audio, video and PDFs. Developers can dial effort levels to trade quality against latency and cost.
Introductory pricing runs at $0.75 per million input tokens and $3.75 per million output tokens through December 31, after which the rates double to $1.50 and $7.50.
The model is live across several surfaces:
- Gemini API, Google AI Studio, Antigravity, Android Studio and Stitch for developers
- Gemini Enterprise for business customers
- The Gemini app for AI Pro and Ultra subscribers
- Google Search AI Mode and Google Sheets
On published evaluations, Gemini 3.8 Flash scored 54.9% on Humanity's Last Exam-Verified and beat larger frontier models on the DeepSWE v1.1 coding benchmark. Google claims it delivers "significant improvements from 3.7 Flash across software engineering, agentic tasks, and critical, multi-step reasoning in specialized domains."
Unpublished Vetting Criteria
Google has not disclosed how Fairwind screens applicants, how many organizations have been admitted, or what monitoring follows approval.
That gap matters because the same capability that patches a Chrome bug can locate one to exploit. A vetting program's value rests entirely on how it is enforced, and none of that enforcement is currently visible from outside.
The approach still contrasts with the industry's recent pattern of releasing broadly and restricting later. Rivals have spent the past month responding to security incidents and capability ratings after the fact rather than before launch.
The open question is whether a gated model can build the track record that broad deployment normally produces. Until Fairwind participants publish results of their own, Google's security claims stand largely on Google's own measurements.