Gemini 3.8 Flash Brings Cheap Reasoning to Cyber

Google debuts Gemini 3.8 Flash and a Cyber variant for defenders, pairing frontier patching scores with Flash pricing via the Fairwind Program.

Google Gemini 3.8 Flash and Flash Cyber models for coding and cybersecurity
Gemini 3.8 Flash and Flash Cyber share one core tuned for long-horizon coding and vulnerability work.· Deepmind
Contents(3)

Gemini 3.8 Flash arrived Sep 2 as Google DeepMind's third Flash in six weeks, holding 3.7 Flash pricing at $0.75 input and $3.75 output per million tokens until Dec 31.

Companies working on this

StartupHub profiles of the companies this article names, with funding and a one-liner from our database.

Gemini
$50M
A regulated cryptocurrency exchange and custodian focused on security and compliance.

Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini Security Lead at DeepMind, frame both variants as one core intelligence split between a general workhorse and a gated Cyber model for trusted defenders via the Fairwind Program, according to DeepMind.

The demo targets code security across C and C++ CyberGym tasks and an internal 20-language benchmark, plus Chrome and Google Cloud codebases, with no attacker access required but defender vetting required for Cyber.

How Gemini 3.8 Flash actually works

Google says the gains came from long-running agentic loops that recursively evaluate and refine the model, paired with rigorous training in cybersecurity to harden reasoning.

At inference the model trades efficiency for diligence. It takes extra reasoning steps and iterative tool calls, burning more tokens at higher effort levels, while lower effort or staying on 3.7 Flash preserves cost.

Think of it like a tireless junior engineer who reruns tests and re-reads docs until the patch holds, not a faster autocomplete that stops at the first plausible answer.

Why this matters, and what it does not fix

At Flash cost it nears frontier results, with 54.9 percent on HLE-Verified, wins on DeepSWE v1.1 long-horizon engineering, plus Vals Finance Agent V2 and the Harvey Legal Agent Benchmark. For security, it posts 47.2 percent pass at 1 on CWE-Bench versus 47.8 percent for a larger frontier model, over 70 percent success across 20 languages, 2.6 times more correct Chrome patches than larger commercial models, and 7.5 to 9.7 percent higher recall for Wiz at 2.3 to 5.2 times lower cost.

That cost curve matters for defenders who need speed, and Google reports a real Cloud Vulnerability Research find in under two hours, plus prompt injection robustness gains measured by Gray Swan. Yet Cyber remains gated, and the base model keeps CBRN and cyber offense safeguards per the Frontier Safety Framework.

What isn't fixed is access and overhead. Introductory pricing doubles on Jan 1, 2027 to $1.50 and $7.50, high effort inflates tokens, and Google hasn't disclosed false positive rates or independent reproductions beyond its chosen benchmarks, so builders should budget for verification and keep 3.7 Flash in reserve for efficiency-first paths.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer

Startups in this story

Profiles for the companies named above.