Google’s third Flash release in six weeks arrived on September 2, 2026. Gemini 3.8 Flash and its cybersecurity-focused sibling Gemini 3.8 Flash Cyber push the Flash family to a new intelligence ceiling while keeping the same $0.75 per million input tokens introductory price introduced with 3.7 Flash. Benchmarks are real: gains over 3.7 on most Artificial Analysis evaluations, parity with Claude Opus 5 on several reasoning suites, outright leads on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. To reach those scores, 3.8 Flash spends more reasoning tokens per task, runs longer end-to-end, and ends up with higher cost per completed task than its predecessor despite the headline price staying flat.
This guide covers the official release, the DeepMind model card, independent data from Artificial Analysis, and the practical implications for developers, enterprises, and security teams.
Quick Facts: Gemini 3.8 Flash at a Glance
- Release date: September 2, 2026, just one day after Anthropic’s Claude Open 5.1 and Claude Fable 5.1 launches took the Artificial Analysis Intelligence Index top spot.
- Variants: Gemini 3.8 Flash (general availability) and Gemini 3.8 Flash Cyber (trusted defenders only via the new Fairwind Program).
- Inputs and outputs: text, image, audio, and video input with up to a 1M-token context window; text output up to 64K tokens.
- Reasoning levels: tunable low, medium, and high effort, identical to the 3.7 Flash API surface.
- Knowledge cutoff: March 2026 for most domains; January 2025 on others, in line with the broader Gemini 3 family.
- Pricing through December 31, 2026: $0.75 per million input tokens, $3.75 per million output tokens, $0.075 per million cached input tokens. Standard rates of $1.50 in and $7.50 out take over on January 1, 2027.
- Headline performance: 59 on the Artificial Analysis Intelligence Index at high effort, ranking 16 of 195 models, with 305.1 tokens per second output speed.
- Where it leads: MMMU-Pro, GPQA Diamond, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, HLE-Verified (54.9%), and DeepSWE v1.1 according to Google and Artificial Analysis.
- Where it slips: AutomationBench-AA Score and Objectives Completed both fall below 3.7 Flash, per Artificial Analysis.
From 3.7 to 3.8: What Changed in Three Weeks
Gemini 3.7 Flash landed on August 13, 2026. Twenty days later, Google shipped 3.8 Flash with the same introductory price, the same 1M-token context window, and the same reasoning-effort dial. The team frames 3.8 Flash as its “most intelligent workhorse model,” with substantial gains over 3.7 across software engineering, agentic tasks, and critical multi-step reasoning in specialized domains. Both 3.8 Flash and 3.8 Flash Cyber share a single foundational intelligence, hardened further by training in cybersecurity.
The DeepMind model card confirms the dependency: 3.8 Flash builds on 3.7 Flash, with most sections referring readers back to the 3.7 model card for architecture, training data, hardware, and software details. The new release is a refinement, not a clean-slate retrain, consistent with Google’s recent cadence of tightly sequenced upgrades sharing infrastructure and tuning advances.
For developers, the API surface, prompt ergonomics, and tool integrations carry over cleanly. For procurement teams, capacity planning, integration testing, and pricing model assumptions made for 3.7 Flash do not need to be redone. The change is on the inside.
Benchmark Performance: Where 3.8 Flash Leads
Across the public benchmark tables published by Google, by the DeepMind model card, and by independent evaluators like Artificial Analysis, 3.8 Flash either leads outright or matches the frontier on the workloads Google positioned it for. The headline numbers, all cited as published in September 2026, are:
- HLE-Verified (multidisciplinary reasoning): 54.9% on the Google blog chart, leading the comparison set.
- MMMU-Pro visual reasoning: leads the published Google comparison.
- GPQA Diamond scientific reasoning: leads the published Google comparison.
- DeepSWE v1.1 (long-horizon software engineering): leads the Google comparison, with the announcement noting 3.8 Flash “outperforms most larger frontier models in autonomously solving complex engineering problems end to end, only at a fraction of the cost.”
- Vals Finance Agent v2 (financial analyst tasks): leads the published Google chart. The Vals Finance Agent v2 benchmark measures performance on quantitative and professional workflows.
- Harvey’s Legal Agent Benchmark (complex legal workflows): leads the published Google chart, validated against the Harvey Legal Agent Benchmark.
The Artificial Analysis Intelligence Index, which combines nine evaluations across agentic work, terminal coding, scientific reasoning, physics, knowledge, and long-context reasoning, places 3.8 Flash (high) at 59 out of a theoretical 100, ranking 16 of 195 models and well above the comparison-class median of 36. That positions 3.8 Flash in the same conversation as Claude Opus 5 and ahead of GPT-5.6 Terra and GPT-5.6 Sol on most benchmarks in Google’s own comparison set. From a 3.7 Flash baseline it also pulls ahead of the previous generation’s flash-tier peers: Muse Spark 1.2, GLM 5.3 Flash, and Qwen 3.8 2.4T A95B on the Intelligence Index, as cited in Google’s release materials.
For visual clarity, here is the leader-versus-rivals picture across the benchmarks where Google claims a lead:
The chart values are illustrative roundings derived from the comparison ranges published by Google. They are intended to convey the shape of the lead, not exact pairwise deltas; check the source charts on the Google blog and Artificial Analysis for the precise numbers in your scenario.
The non-obvious story is regression. According to Artificial Analysis, 3.8 Flash drops below 3.7 Flash on AutomationBench-AA: Score and on AutomationBench-AA: Objectives Completed. Google framed 3.8 Flash as an “intelligent workhorse” for software engineering and agentic knowledge work, so the automation regression is notable rather than damning: the model trades breadth on routine enterprise workflows for depth on harder reasoning and coding problems.
The Reasoning-Budget Tradeoff: 3.8 “Works Harder”
The key sentence in Google’s announcement: “These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence, executing extra reasoning steps, and calling tools iteratively. At times, the model might use more tokens to maximize performance, especially at higher effort levels.”
Artificial Analysis quantifies it. At high reasoning effort, 3.8 Flash generated 120 million output tokens for the full Intelligence Index, against a class median of 71 million, roughly 70% more tokens per benchmark run than the median reasoning model in the same price tier. Time-to-first-token is 13.30 seconds against a class median of 2.99 seconds.
The cost effect is striking. With the per-token price identical to 3.7 Flash, the cost per completed Intelligence Index task lands at $0.58, ranking 3.8 Flash 59th of 195 models on the cost-per-task axis. The same model is simultaneously cheap per token and moderately expensive per finished job. Developers who build budget forecasts on price-per-million will need to re-estimate against completed-task costs.
Google acknowledges this directly. For applications where compute efficiency is the primary constraint, the announcement directs developers to use lower effort levels or stay on 3.7 Flash, which remains fully supported for efficiency-first workloads. 3.8 Flash is positioned for tasks where intelligence gains dominate cost considerations, not for high-volume lightweight calls.
Across Reasoning Levels: A Smarter 3.8
Looking across reasoning effort levels changes the picture. Artificial Analysis tracks Gemini 3.7 Flash and 3.8 Flash at low, medium, and high effort on the same Intelligence Index. The full grid:
Three patterns emerge. First, 3.8 Flash moves up at every level: +1 point at low, +4 at medium, +3 at high. Second, the biggest jump is in the middle of the dial. Medium reasoning captures the largest single-version delta, suggesting Google concentrated tuning effort there. Third, and most operationally important: 3.8 Flash at medium effort scores 57, essentially matching 3.7 Flash at high effort (56). The headline intelligence for a typical production workload is reachable at lower token spend, lower latency, and lower per-task cost.
For teams currently running 3.7 Flash at high effort, the migration path is not “upgrade to 3.8 high to spend more.” It is “upgrade to 3.8 medium to spend less for the same result.” Workloads that already saturated 3.7 high, however, benefit most from 3.8 high: a +3-point Intelligence Index lift plus the coding gains is what unlocks the agentic use cases Google is targeting.
The radar view below shows the multi-dimensional shape of the upgrade, with reasoning-level quality on one axis, code depth on another, knowledge-work performance on a third, multimodal on a fourth, and cybersecurity capability on a fifth. Larger area means more capable across the workload mix:
The radar is illustrative, scaled from the public benchmark charts on Google’s release and from Artificial Analysis. The Cybersecurity axis reflects Gemini 3.8 Flash Cyber’s CyberGym Pass@1 leadership; for the Flash variant specifically, the cyber axis reflects the shared underlying intelligence rather than the Fairwind-Program-only fine-tune.
Speed, Verbosity, and the Cost-per-Task Reality
Speed remains one of 3.8 Flash’s strongest stories. Artificial Analysis measured sustained output at 305.1 tokens per second, placing it third of 195 models and well above the comparison median of around 70 tok/s, roughly the same throughput regime as 3.7 Flash despite the higher reasoning depth.
Verbosity is the flip side. 120 million output tokens for the Intelligence Index is the “red” rating on the cost-per-task dimension. The model is not slow; it is generous. It spells out more reasoning steps, calls tools more often, and explains more before finalizing. For debugging sessions or research workflows, that is an asset. For high-volume pipelines that relied on tight Flash-style responses, it represents a budget shift.
Cost per Intelligence Index task at $0.58 places 3.8 Flash at 59th of 195 models, behind peers that produce tighter responses for the same reasoning quality. If your workload is many short calls with crisp outputs, 3.7 Flash or 3.6 Flash (repriced to match 3.7) may still be the right tool. If your workload is fewer, harder calls where reasoning depth moves the needle, 3.8 Flash is the new best in class at the Flash tier.
Pricing and API Access
Pricing is identical to 3.7 Flash through the end of 2026:
| Component | Introductory Rate (through Dec 31, 2026) | Standard Rate (from Jan 1, 2027) |
|---|---|---|
| Input per 1M tokens | $0.75 | $1.50 |
| Output per 1M tokens | $3.75 | $7.50 |
| Cached input per 1M tokens | $0.075 (90% discount) | Not yet published |
| Image input | Billed separately, see Gemini API docs | Same |
The introductory window is the same promotional schedule Google opened for 3.7 Flash, so anyone planning Q1 2027 spend should model against the standard rates rather than the current headline. On a blended 7:2:1 cache-hit/input/output workload, Artificial Analysis places 3.8 Flash at $0.58 per million tokens, competitive within the Flash tier.
API access opens immediately: the Gemini API and Google AI Studio for developers, Gemini Enterprise for organizations, Google Antigravity for the agent-first workflow demos Google highlighted, Google AI Mode in Search for consumers, and the Gemini app for end users on Google AI Pro and Ultra. No waitlist for the general Flash variant. The Cyber variant is restricted to the Fairwind Program.
Gemini 3.8 Flash Cyber: A Frontier Cybersecurity Model

Gemini 3.8 Flash Cyber is the more strategically interesting release. Available only to trusted defenders through the new Fairwind Program, it provides a decisive advantage in vulnerability detection and automated patching at Flash-tier speed and cost.
The benchmark results justify the framing. On the standard industry benchmark for finding vulnerabilities, CyberGym, Gemini 3.8 Flash Cyber demonstrates frontier-level performance in autonomous vulnerability discovery, surpassing 3.5 Flash Cyber and significantly larger frontier models. On CWE-Bench, run by Collinear, the model sits on the Pareto frontier with a Pass@1 of 47.2% against a leading frontier model at 47.8%, offered at a significantly lower cost. On Google’s internal benchmark spanning 20 programming languages, the model surpasses 70% success rate, a substantial leap over previous Gemini cybersecurity models.
Real-world validation is already in motion. The Chrome Security team reports that 3.8 Flash Cyber produced 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models they tested, even though those competing models were much larger. Wiz found that 3.8 Flash Cyber achieves +7.5 to +9.7 percentage points higher recall on their internal penetration testing benchmark, at 2.3 to 5.2 times lower cost compared to other leading frontier models. Google’s Cloud Vulnerability Research team used 3.8 Flash Cyber to find a critical foundational vulnerability in less than two hours, a vulnerability that would typically take months of manual research to surface.
The threat model is explicitly defensive. Google prioritized vulnerability fixing over offensive capabilities; Cyber ships with a more permissive set of cybersecurity mitigations than the standard Flash variant, which carries the usual CBRN and cyber-offense safeguards defined by Google’s Frontier Safety Framework. Cyber is a defender tool with broader capabilities, gated behind the Fairwind Program’s trusted-defender criteria for governments, critical infrastructure operators, and software maintainers.
What 3.8 Flash Means in a Crowded Release Week
September 2, 2026 was not a quiet day for frontier models. Anthropic’s Claude Open 5.1 and Claude Fable 5.1 releases took the Artificial Analysis Intelligence Index top spot one day earlier, and OpenAI’s GPT-5.6 family has held a strong position since July. Google’s response, one day later, is a Flash-tier model that closes much of the intelligence gap to those flagships at a fraction of the per-token cost, while introducing a dedicated cybersecurity variant that no other frontier lab is shipping today.
The strategic takeaway is the pattern, not any single benchmark. Google is shipping Flash updates every few weeks, each tuned for a specific workload mix, and using the same underlying intelligence to power multiple specialized variants. The flash tier is no longer the cheap-but-shallow option; it is now the deployment-focused tier where Google moves fastest, and where most production agents should be tested before reaching for a more expensive model. For a deeper comparison of how this fits the broader 2026 model landscape, see our Gemini vs ChatGPT 2026 comparison and our coverage of competing open Flash-tier releases like Qwen 3.8 Flash.
Safety, Limitations, and the Knowledge Cutoff
The DeepMind model card publishes an unusually transparent safety comparison. Versus Gemini 3.7 Flash, 3.8 Flash shows -0.4 pp on Text-to-Text Safety (better), 0.0 pp on Image-to-Text Safety (flat), +0.2 pp on objective tone (better), +1.1 pp on unjustified refusals (slightly more refusals on borderline prompts), and +5.4 pp on Multilingual Safety (a regression, since lower is better). The card notes the regression is “slight” and manual review found losses overwhelmingly false positives or not egregious. Anyone deploying 3.8 Flash in non-English markets should validate for their own content domains before turning it loose.
The standard Flash model inherits Gemini 3 family limitations: occasional hallucinations, possible jailbreak susceptibility that Google continues to harden, occasional slowness or timeout issues, and the documented verbosity tradeoff discussed earlier. The knowledge cutoff is March 2026 for most domains, with January 2025 data still surfacing in some areas, in line with the broader Gemini 3 family. Gemini 3.8 Flash was not assessed as reaching any Tracked or Critical Capability Levels under Google’s April 2026 Frontier Safety Framework, based on the unchanged 3.7 Flash assessment.
Prompt-injection robustness has improved. Gemini 3.8 models show a significant leap in prompt-injection resistance as measured by Gray Swan, which is a meaningful gain for any deployment that exposes the model to untrusted external content, from email summarization to web-research agents.
How to Try 3.8 Flash Today
Three audiences get different on-ramps.
Developers can start building in Google AI Studio, the Gemini API, or Android Studio right now, with Google Antigravity as the agent-first workflow surface for the demos Google highlighted in the announcement. Stitch is available for generating user interfaces.
Enterprises can access 3.8 Flash through the Gemini Enterprise Agent Platform and existing Google Cloud procurement channels.
Consumers get 3.8 Flash in the Gemini app for Google AI Pro and Google AI Ultra subscribers, alongside AI Mode in Google Search and Gemini in Google Sheets.
Cybersecurity teams should apply to the Fairwind Program for trusted-defender access to 3.8 Flash Cyber. Eligibility is restricted to government authorities, critical infrastructure operators, and software maintainers with prioritized access needs.
Frequently Asked Questions
When was Gemini 3.8 Flash released?
September 2, 2026, three weeks after Gemini 3.7 Flash and one day after Anthropic’s Claude Open 5.1 and Claude Fable 5.1 launches.
How much does Gemini 3.8 Flash cost?
Through December 31, 2026, the model costs $0.75 per million input tokens and $3.75 per million output tokens, identical to 3.7 Flash’s introductory pricing. Cached input is $0.075 per million tokens. From January 1, 2027, standard rates of $1.50 per million input and $7.50 per million output take effect.
Is Gemini 3.8 Flash better than Gemini 3.7 Flash?
On most reasoning, coding, and knowledge-work benchmarks, yes. The Artificial Analysis Intelligence Index improves from 56 (3.7 high) to 59 (3.8 high), and Google reports outright leads on HLE-Verified, MMMU-Pro, GPQA Diamond, DeepSWE v1.1, Vals Finance Agent v2, and Harvey’s Legal Agent Benchmark. On AutomationBench-AA, however, 3.8 Flash regresses relative to 3.7, so for routine enterprise workflow automation 3.7 may still be the better fit.
Is Gemini 3.8 Flash better than Claude Opus 5?
On most benchmarks in Google’s comparison set, yes. 3.8 Flash leads or ties Opus 5 on MMMU-Pro, GPQA Diamond, HLE-Verified, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, and DeepSWE v1.1, while running at Flash-tier pricing. Claude Open 5.1 and Claude Fable 5.1 hold the overall Intelligence Index top spot, but Flash-tier value for the dollar now firmly belongs to Google.
Is Gemini 3.8 Flash better than GPT-5.6?
On most benchmarks in Google’s comparison set, yes. 3.8 Flash beats GPT-5.6 Terra and GPT-5.6 Sol on HLE-Verified, Vals Finance Agent v2, Harvey’s Legal Agent Benchmark, and DeepSWE v1.1. GPT-5.6 Sol retains the highest peak intelligence on the Intelligence Index at max reasoning, but Flash-tier value for the dollar across most workloads is Google’s.
Why does 3.8 Flash use so many more tokens than 3.7 Flash?
Google describes 3.8 Flash as a model that “works harder”: on complex tasks it executes extra reasoning steps, calls tools iteratively, and produces more explanatory output. Artificial Analysis measured 120 million output tokens for the full Intelligence Index versus a class median of 71 million. Verbosity is the cost of the benchmark gains.
Should I switch from 3.7 Flash at high effort to 3.8 Flash at medium effort?
For many production workloads, yes. 3.8 Flash at medium effort scores 57 on the Intelligence Index, essentially matching 3.7 Flash at high (56), while using fewer reasoning tokens and finishing faster per task. That is the most efficient migration path for the largest group of existing Flash users.
What is Gemini 3.8 Flash Cyber?
A cybersecurity-specialized variant of 3.8 Flash, available only to trusted defenders through Google’s Fairwind Program. It leads CyberGym, sits on the Pareto frontier of CWE-Bench at 47.2% Pass@1 versus 47.8% for the leading frontier model, and exceeds 70% success rate on Google’s internal 20-language vulnerability benchmark. The Chrome Security team reports 2.6x more correct Chrome vulnerability patches than larger competing models.
What is the context window and knowledge cutoff?
Gemini 3.8 Flash has a 1 million-token input context window with up to 64K output tokens. The knowledge cutoff is March 2026 for most domains, with January 2025 data still surfacing in others.
Is Gemini 3.8 Flash multimodal?
Yes. 3.8 Flash supports text, image, audio, and video input, and produces text output, the same multimodal surface as 3.7 Flash.
Conclusion
Gemini 3.8 Flash is the third Flash release in six weeks, and the one that finally gives Google’s efficiency tier a credible claim to the “intelligence per dollar” crown for production agentic workloads. The benchmark table is real: leadership on HLE-Verified, MMMU-Pro, GPQA Diamond, DeepSWE v1.1, Vals Finance Agent v2, and Harvey’s Legal Agent Benchmark, plus parity with Claude Opus 5 and outright wins over GPT-5.6 on most Google comparison points. The cost story is more nuanced: the same $0.75 per million input tokens as 3.7 Flash, but $0.58 per completed Intelligence Index task because the model thinks out loud more often. For many workloads, the right migration is 3.7 high to 3.8 medium, capturing nearly equal intelligence at lower per-task spend.
For security teams, Cyber is the headline: frontier CyberGym performance, Pareto-frontier CWE-Bench scores, 2.6x Chrome vulnerability patching gain, and a critical vulnerability discovered in under two hours.
Google’s broader pattern is the real story. Cheap, fast, capable Flash models updated every few weeks, each tuned for a specific workload mix, all sharing the same underlying intelligence. For most production deployments, 3.8 Flash deserves a careful evaluation: run your hardest workloads at high reasoning, your typical workloads at medium reasoning to capture the 3.7-high-equivalent efficiency gain, and your lightweight calls on 3.7 Flash or 3.6 Flash where verbosity does not pay for itself.
