Google Ships Gemini 3.8 Flash With Sharper Coding Scores and a Locked-Down Cyber Variant
Google rolled out Gemini 3.8 Flash on September 2, three weeks after its previous Flash update, pitching the mid-tier model as a cheaper alternative to frontier systems from Anthropic and OpenAI. The company also released a restricted cybersecurity variant, Gemini 3.8 Flash Cyber, only to vetted defenders.
The release lands during an intense pricing and benchmark battle across the top AI labs. Google's message with 3.8 Flash is that a middle-tier model, tuned to reason harder before answering, can match a frontier system on many long-horizon coding and agent tasks.
Pricing and Availability of Gemini 3.8 Flash
The model is generally available in the Gemini app for AI Pro and Ultra subscribers, in AI Mode and Gemini in Google Sheets, and to developers through AI Studio, the Gemini API, and Google Antigravity. It supports a 1,048,576-token context window and up to 65,536 output tokens.
Introductory pricing runs through December 31 at $0.75 per million input tokens and $3.75 per million output tokens, matching the promotional rate Google set for 3.7 Flash. From January 1, 2027, the price doubles to $1.50 input and $7.50 output.
Google offers three reasoning levels — low, medium (the default), and high — and thinking tokens bill as output. The model accepts text, image, audio, and video inputs and returns text. Knowledge cutoff extends to March 2026 for most domains, with some areas capped at January 2025.
Where the New Flash Model Improves on 3.7
Google says the biggest gains show up in agentic coding and long-horizon reasoning — tasks where a model must plan, call tools, and iterate rather than answer in a single pass. The company frames the difference as 3.8 Flash working harder by executing more internal reasoning steps before responding.
Benchmark shifts against the previous Flash release include:
- DeepSWE v1.1 software engineering: 73.7% pass rate, up from 65.3%
- OSWorld-2.0 computer-use tasks: 59.0%, up from 50.6%
- BioMysteryBench, difficult tier: 56.5%, up from 43.5%
- HLE-Verified academic reasoning: 54.9% across STEM, humanities, and professional fields
Google also reports lower susceptibility to indirect prompt injection, with Gray Swan attacks succeeding on 5.5% of tries within 15 attempts, down from 9.2% for 3.7 Flash.
How Gemini 3.8 Flash Stacks Up to Claude Opus 5
The pricing gap with Anthropic's flagship is the crux of Google's pitch. Claude Opus 5 costs roughly $25 per million output tokens — more than six times the introductory 3.8 Flash rate.
On several agent benchmarks the two are close. Gemini 3.8 Flash scored 73.7% on DeepSWE v1.1 to Opus 5's 74.0%, and 89.4% on Terminal-bench 2.1 to Opus 5's 89.1%. On Vals Finance Agent v2 and Harvey's Legal Agent benchmark, Google's model narrowly led.
Opus 5 still opens wider gaps on some tasks, including 51.8% versus 19.1% on Terminal-bench 4.0 and 75.4% versus 59.0% on OSWorld-2.0. The Wall Street Journal previously reported that internal testing preferred Opus for certain workloads, though the methodology was not published.
The Fairwind Program and Gemini 3.8 Flash Cyber
Alongside the general release, Google is launching Gemini 3.8 Flash Cyber under a new initiative called the Fairwind Program. Access is restricted to what the company calls trusted defenders, including government authorities, critical-infrastructure operators, and software maintainers.
Google says the Cyber variant is tuned around defensive tasks and generated 2.6 times more correct patches to Chrome vulnerabilities than the leading commercial models it tested. On CyberGym vulnerability discovery it hit 86.2% pass@1, and on CWE-Bench automated patching it scored 47.2%, effectively tied with Anthropic's Claude Fable 5 at 47.8%.
On Wiz's penetration-testing benchmark the Cyber variant beat comparison models by 7.5 to 9.7 points on recall at 2.3 to 5.2 times lower cost. The Fairwind gating mirrors an approach used by other labs for security-focused releases and reflects industry pressure to keep aggressive security tooling out of general access.
Further Reading
Google's framing positions 3.8 Flash as evidence that a mid-tier model with more reasoning steps can close the gap on frontier coding and agent work at a fraction of the cost. Whether independent testers reproduce the benchmark claims — and whether the Cyber variant's gating holds — will shape how quickly enterprises migrate.