the short version
- On a comprehensive internal benchmark spanning 20 programming languages, Gemini 3.8 Flash Cyber exceeds 70% success rate in autonomous vulnerability discovery.
- CWE-Bench patching performance reaches 47.2% pass@1, within margin of a leading frontier model at 47.8% while costing 2.3-5.2x less.
- Chrome Security measured 2.6x more correct patches from 3.8 Flash Cyber than the best commercial models; Wiz reported 7.5-9.7% higher recall at 2.3-5.2x lower cost.
Gemini 3.8 Flash Cyber delivers quantifiable improvements in autonomous vulnerability discovery and patch generation compared to 3.5 Flash Cyber. On a comprehensive internal benchmark covering 20 programming languages, the model exceeds 70% success rate in finding vulnerabilities across complex codebases. On CWE-Bench, the industry standard for patching, it achieves 47.2% pass@1-within 0.6 percentage points of a leading frontier model at 47.8%, while costing 2.3–5.2x less. Real-world deployment data from Google and partner organizations confirm these gains translate to measurable security impact.
Vulnerability discovery across programming languages
The standard industry benchmark for vulnerability detection, CyberGym, tests detection capability in C and C++ codebases. On CyberGym, Gemini 3.8 Flash Cyber surpasses both 3.5 Flash Cyber and significantly larger frontier models. However, real-world defensive work is not limited to C and C++. To capture this, Google evaluated 3.8 Flash Cyber against an internal benchmark requiring the model to discover vulnerabilities across complex codebases spanning 20 programming languages. The model achieves a success rate exceeding 70%.
This multi-language benchmark matters to teams deploying security tooling in production because codebases span multiple languages. A model that detects vulnerabilities only in C or C++ cannot serve research teams, cloud infrastructure operators, or security-focused enterprises working with TypeScript, Python, Rust, Go, or Java. The 70% threshold across 20 languages suggests the model can handle the polyglot environments where actual vulnerability research happens.
Patching performance on CWE-Bench
Patching-the ability to generate correct fixes-is more difficult than detection and more valuable to defenders. CWE-Bench, run by Collinear, is a challenging external benchmark for vulnerability patching. Gemini 3.8 Flash Cyber achieves 47.2% pass@1 on this benchmark, placing it on the Pareto frontier. A leading frontier model scores 47.8% pass@1, a difference of 0.6 percentage points, but Gemini 3.8 Flash Cyber is available at 2.3–5.2x lower cost per token.
For teams running patch generation in production, this cost difference compounds. A security team iterating on patch suggestions across hundreds of findings per day will see material savings. The model is priced at $0.75 per million input tokens and $3.75 per million output tokens at introductory rates through December 31, 2026. Starting January 1, 2027, pricing increases to $1.50 per million input tokens and $7.50 per million output tokens.
Prompt injection robustness
Gemini 3.8 models have made a significant leap in prompt injection robustness as measured by Gray Swan, protecting Gemini model users from prompt-injection-related malicious attacks. The source material does not specify the Gray Swan benchmark score for 3.8 Flash Cyber or comparison figures against 3.5 Flash Cyber, only that an improvement exists and was measured.
Performance in production deployments
Google and partner organizations have deployed Gemini 3.8 Flash Cyber against real codebases and published specific results. The Chrome Security team found that 3.8 Flash Cyber produces 2.6 times more correct patches to vulnerabilities in Chrome than the best commercial models that are much larger. Wiz, a cloud security vendor, reported that Gemini 3.8 Flash Cyber achieves 7.5–9.7% higher recall on its internal penetration testing benchmark while costing 2.3–5.2x less than other leading frontier models. Google's Cloud Vulnerability Research team used the 3.8 Flash Cyber model to find a critical foundational vulnerability in less than 2 hours, where similar research typically takes months.
These three data points come from different contexts-browser security, cloud infrastructure, and foundational research-and they align on cost and output quality. However, each measurement is scoped: Chrome's comparison names no specific commercial models, Wiz's benchmark is internal, and the foundational vulnerability discovery is a single incident. Practitioners should treat these as evidence of capability rather than guarantees of performance on their own codebase.
Access through the Fairwind Program
Gemini 3.8 Flash Cyber is not available via the standard Gemini API. It is available only to trusted defenders through Google's new Fairwind Program, which provides prioritized access to government authorities, critical infrastructure operators, and software maintainers. Teams interested in deploying the model must apply for access. The source material does not specify approval criteria, timelines, or whether access is restricted by geography or organization type.
Safety mitigations and cyber offense constraints
Gemini 3.8 Flash Cyber ships with a more permissive set of mitigations for cybersecurity than standard Gemini 3.8 Flash, which ships with safeguards against misuse in Chemical, Biological, Radiological, and Nuclear (CBRN) and cyber offense domains. This permissiveness reflects the model's focus on defensive use cases. The restricted access through the Fairwind Program is designed to enable beneficial defensive security work while limiting exposure to actors who might misuse the enhanced cyber capabilities.
What remains unknown
The announcement does not report end-to-end latency for vulnerability detection or patch generation tasks, making it difficult to estimate how quickly a security team could iterate or how many concurrent scans a given cost budget could support. The multi-language benchmark result of exceeding 70% across 20 programming languages lacks detail on which languages are covered, whether performance is uniform across them, or how that result compares to 3.5 Flash Cyber. The Gray Swan prompt injection robustness improvement is mentioned but not quantified. Teams should request benchmark data tailored to their own codebases and languages before committing to production deployment.
Questions this raises
How does Gemini 3.8 Flash Cyber perform on vulnerability detection?
On an internal benchmark covering 20 programming languages, Gemini 3.8 Flash Cyber exceeds 70% success rate in finding vulnerabilities across complex codebases. This multi-language capability matters because real-world codebases span TypeScript, Python, Rust, Go, Java and other languages beyond C and C++.
How does Gemini 3.8 Flash Cyber compare on CWE-Bench patching?
It achieves 47.2% pass@1 on CWE-Bench, placing it on the Pareto frontier. A leading frontier model scores 47.8% pass@1, only 0.6 percentage points higher, but Gemini 3.8 Flash Cyber costs 2.3-5.2x less per token, compounding savings for teams iterating on patch generation at scale.
How do I access Gemini 3.8 Flash Cyber?
Gemini 3.8 Flash Cyber is not available via the standard Gemini API. It is available only through Google's Fairwind Program, which provides prioritized access to government authorities, critical infrastructure operators, and software maintainers. Teams must apply for access.
These daily notes are drafted by a model I run and operate myself - the same kind of pipeline this site is about - from sources published in the previous 24 hours, and every one lists what it read. The longer essays, the talks and the preprint are mine, written by hand.
