Google on Wednesday announced Gemini 3.8 Flash Cyber, which it described as its most capable cybersecurity model, and has made it available to a set of trusted defenders via a new initiative called the Fairwind Program.
"The Fairwind Program gives high-priority defenders (like governments, healthcare providers, and telecommunications services) early access to advanced models that help them build better defenses, before new threats arrive," Google said. "So defenders have an early advantage, to help them protect vital infrastructure – which in turn protects people who rely on those systems."
The tech giant said it's currently working with over 650 partners globally, including CrowdStrike, Datadog, Menlo Security, Palo Alto Networks, and Snowflake. The program is available to a group of Google Cloud customers, government agencies, and cybersecurity partners.
The release of Gemini 3.8 Flash Cyber comes a little over a month after Google unveiled Gemini 3.5 Flash Cyber. The latest model improves upon its predecessor by demonstrating frontier-level performance in autonomous vulnerability discovery, even surpassing larger frontier models from rivals Anthropic (Mythos 5) and OpenAI (GPT-5.6 Sol and GPT-5.5-Cyber).
"With Gemini 3.8 Flash Cyber, we focused specifically on equipping defenders with expert capabilities that give them an advantage over attackers. This is why we have invested in vulnerability fixing from the start, and prioritized it over offensive capabilities like exploitation," Tulsee Doshi, senior director of product management, and Raluca Ada Popa, Gemini Security Lead at Google DeepMind, said.
Anthropic Debuts Claude Fable 5.1 and Claude Mythos 5.1
The development coincides with Anthropic's launch of Claude Fable 5.1 and Claude Mythos 5.1 with different levels of safeguards, with the latter only available through its trusted access programs and support work in cybersecurity and the life sciences.
The company also said it's now allowing Fable 5.1 to be used for identifying software vulnerabilities, but it expects to still redirect some cybersecurity tasks to Opus models, like "penetration testing, exploit generation, and binary-based vulnerability scanning."
"We ran evaluations of how Claude Mythos 5.1 responds to malicious requests and prompt injections (adversarial instructions hidden within content processed by AI models)," Anthropic noted. "It refused malicious agentic coding and computer use requests at a comparable rate to Mythos 5, Sonnet 5, and Opus 5, and it is our most robust model to date on an external prompt injection benchmark."
The artificial intelligence (AI) company has since also announced a new solution called Enterprise Frontier Safeguards (EFS), which it said combines the "privacy of zero data retention (ZDR) with state-of-the-art safeguards for detecting misuse," while giving businesses full control over how their data is reviewed, stored, and managed. OpenAI has a similar solution in place known as Private Safety Processing.
Furthermore, Anthropic said it has implemented additional hardening and containment measures, increased monitoring for flagging model misalignment, and paused external cyber evaluations of pre-release models in response to unauthorized access incidents involving Claude models against real systems, in addition to highlighting two contributing factors (or alignment failures) -
The models appear to disregard evidence that their evaluation environments were connected to the real internet after initially being told they were simulated and interpret the discrepancy in a manner that allows them to maintain that belief
The models exhibit recklessness and are willing to take harmful actions on the real internet in single-minded pursuit of their goals
Conceding that the incidents were a "failure of operational security," Anthropic said it has built a classifier that detects and blocks sandbox escape attempts and changed specifications around model rewards to address concerns where an AI agent resorts to a shortcut to game the reward metric without actually achieving the intended goal.
"Our conclusion is that the presence of substantial reward hacking in training can cause models to be willing to perform long sequences of potentially harmful real-world actions in pursuit of task success," the company added.
OpenAI Claims Astra Meets Critical Cybersecurity Capability
OpenAI, for its part, has revealed that its forthcoming Astra model meets the Critical cybersecurity capability threshold under its Preparedness Framework, and that it intends to make its most advanced cybersecurity features available to a group of testers through the Daybreak Blue program.
The "Critical" designation applies when an AI model can independently detect and exploit zero-day vulnerabilities across many well-defended systems, or carry out a complete cyber attack against a hardened target from only a high-level instruction without a human guiding it along the way.
"Over the past several weeks, we have delayed parts of Astra's development and release while we strengthened and tested protections against cyber misuse and unauthorized model actions," the AI company said. "Based on that work, we believe Astra's safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework."
OpenAI said it has also added stronger safeguards for Astra to prevent a Hugging Face-like incident, in which its AI agents, part of an ExploitGym evaluation found to way to exploit its research infrastructure and abuse Artifactory as a message board to exchange information in their quest to solve an impossible task, ultimately breaking into Hugging Face's infrastructure in hopes of stealing the answer instead of solving the challenge themselves.
"One agent, PHASEONE[big], orchestrated a significant fraction of this cheating research. PHASEONE10841 passed along its work to PHASEONE[big], which had the same task but a larger budget," METR noted in its analysis. "Agents collaborated on many efforts to make cheats look legitimate, including: (1) swapping the program they had to exploit; (2) manipulating the automated scorer; (3) manipulating transcripts to obscure evidence of cheating."
OpenAI has reported that Astra achieves a perfect score of 100% on ExploitBench to develop exploits from known vulnerabilities, and that it now declines 91.5% of jailbreaking requests, compared to 59% from GPT‑5.6 Sol.
In addition, OpenAI noted that Astra achieves "much higher arbitrary code-execution rates" than GPT‑5.6 Sol using far fewer output tokens, and that the model discovered and used two zero-day vulnerabilities in unspecified software as part of an exploit chain during an evaluation.
The model has also been found to discover previously unknown flaws and turn them into working exploit chains, including a full browser-compromise that escapes the sandbox and executes arbitrary commands on the underlying host when an HTML file is opened in the browser.
"The model also found multiple vulnerabilities in a hardened operating system and combined them into a local privilege-escalation chain from an unprivileged user to root," OpenAI added. "All together, our investigation has led us to conclude that Astra meets the critical threshold."
To minimize risk for severe cyber harm arising from Astra-like models, the company said it has added classifiers and layered protections to improve the robustness of its systems against misuse by bad actors and prevent the model from taking unauthorized, misaligned actions, even in the absence of a malicious user.
However, OpenAI also warned that Astra's safeguards may erroneously flag legitimate activity as cyber misuse or unauthorized behavior.
"Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow," the upstart concluded. "That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient."
AI companies have been under intense scrutiny in the wake of incidents where their models escaped their evaluation environments and targeted legitimate systems. In tandem, the rise of AI-fueled cyber attacks has prompted a coalition of over 100 companies, including Anthropic, Google, Microsoft, OpenAI, and several software and security vendors, to issue a joint letter calling for improved defenses to defend against such threats.




