6.3 AI for Defence · Module 6 ยท Adversarial Risks

Critical-Capability Models: Astra, Mythos and Fairwind

Every frontier lab now gates cyber capability behind a vetting programme. What that means for the tier you can use at OFFICIAL — and what to ask your security adviser.

AdversarialPolicyTier Decision Helper~12 min

The big idea

On 3 September 2026 OpenAI released GPT-6 Astra, the first of its models to reach the “Critical” cyber threshold in its Preparedness Framework. In plain terms: given tools and access, it can find unknown flaws and build exploits across hardened systems without step-by-step guidance. During evaluation it found two zero-days (now being disclosed) and built a full browser-sandbox-escape chain. The same week, Anthropic shipped Claude Fable 5.1 and Mythos 5.1 with two safeguard levels, and Google shipped Gemini 3.8 Flash Cyber through its Fairwind programme. Three labs, one pattern: the capability exists, and you don't get it by default.

Key insight: for most Defence APS and DISP-member work at OFFICIAL, nothing in this lesson changes what you're allowed to do. What changes is what the tool will do for you — it may refuse, pause or stop — and who decides whether you get the version that doesn't.

Three labs, three gates

What “Critical” means, and the equivalent at each lab.

OpenAI

Preparedness Framework · Daybreak

“Critical” is the top cyber threshold. Astra’s public version refuses exploit discovery; advanced cyber is only available through Daybreak, for vetted defenders, expanding to Daybreak Blue in coming weeks. The reported cyber results reflect Daybreak Blue access, not default production.

Anthropic

Mythos trusted access

Claude Fable 5.1 and Mythos 5.1 ship with two safeguard levels. Mythos-class cyber capability sits behind a trusted-access programme rather than the public product.

Google

Fairwind

Gemini 3.8 Flash Cyber is delivered via Fairwind, Google’s vetted-access programme. Same shape: the cyber-capable variant is a programme, not a pricing tier.

What it means in practice

Four things to expect on the public tiers. Tap to expand.

The export-control thread

Two events this year make the same point from different directions. In June, access to Claude Fable 5 was suspended. This month, Astra launched to Daybreak trusted-access enterprises first, with everyone else told to expect it “in the coming days”. Access to frontier capability is now a policy variable, not a product feature. It can be granted, staged, gated or withdrawn by the vendor, and by whoever the vendor answers to.

This course can’t tell you what that means for your DISP obligations or for any AUKUS-adjacent work — that is a conversation with your security adviser and your Defence customer. It can tell you what to ask:

If a vendor gates or withdraws the model our workflow depends on, what is the documented fallback, and has anyone tested it?
Does our tool register (Lesson 7.1) record which tier and programme each product sits on, and who would update it if that changed?
Is any of our work near a trusted-access programme at all, and if so, who in our organisation would be vetted and on whose authority?
Does anything AUKUS-controlled sit anywhere near this tooling? It shouldn’t (Lesson 7.3), and a vetting programme doesn’t change that.

Not security advice

Nothing on this page is a ruling about your organisation. Export-control status, DISP obligations and AUKUS handling are for your security adviser and your Defence customer. Use this lesson to arrive at that conversation with better questions.

Which tier can I use for this task?

Pick the task closest to yours. You’ll get an OFFICIAL suitability read, whether a gated programme would be involved, and who to ask. It is a prompt for the conversation, not a substitute for it.

One paragraph on monitorability

Astra is, by OpenAI’s numbers, better behaved than its predecessor: it went outside authorised scope in 0% of honeypot tests (GPT-5.6 Sol: 56% without production safeguards) and never attempted to circumvent auto-review. But OpenAI also reports that Astra’s monitorability has decreased: it is less likely to include incriminating information in its chain of thought, and can sometimes evade internal monitors on sabotage tasks. For OFFICIAL work the practical consequence is modest but real — do not treat a model’s own account of what it did as your audit trail. Log the actions, check the outputs, and keep the boundary in your infrastructure, not in the prompt.

Check your understanding

Key points to remember

“Critical” is OpenAI’s top cyber threshold; Astra is the first of its models to reach it.
All three frontier labs gate cyber capability: Daybreak, Mythos trusted access, Fairwind.
Public tiers refuse exploit discovery and may pause or stop legitimate work; the API stops rather than asks.
Access to frontier capability is now a policy variable. Record the tier, plan the fallback, ask your security adviser.