Why it matters
The model is designed to be invoked several times against the same target so that a defensive system such as CodeMender can explore more possible execution paths. Google reports that this repeated-search approach was competitive with larger models on a cyber benchmark and improved production scanning for Chrome, Android, Cloud, Ads, YouTube, and other internal code.
The dual-use boundary is unusually concrete. In a two-hour internal test, the model found remote-code-execution vulnerabilities in public APIs and memory corruption in a sensitive service, then generated a reliable exploit that bypassed common protections. Limited access, logging, evaluation, incident response, and clear authorization rules will be as important as benchmark performance when this class of capability expands.
Go to the source
Read the evidence behind this analysis. External links open in a new tab.
Google DeepMind — Introducing Gemini 3.5 Flash Cyber


