1.7Module 1 · AI Foundations & Mental Models

GPT-6 Astra and the Frontier Four

OpenAI shipped GPT-6 Astra on 3 September 2026, the same week Anthropic, Google and Meta each shipped a flagship of their own. This lesson explains what Astra actually is, lines the four up side by side, and gives you a calculator and a hands-on test so you can decide with your own numbers rather than anyone's launch chart.

Frontier Four Comparison Price-per-Task Calculator Two-Model Test 5-Question Quiz

What Astra is, in plain English

GPT-6 Astra is OpenAI's new top model. Greg Brockman introduced it as OpenAI's "most intelligent and, also very importantly, our most aligned model yet". It comes in two versions — Astra and Astra Pro — and that is the whole GPT-6 lineup. If you were expecting a GPT-6 Luna, Terra or Sol, there isn't one.

Sol, Terra and Luna carry on as the everyday models. They are GPT-5.6. GPT-5.6 Sol stays the default on ChatGPT Plus and Pro, Luna stays the free default, and Terra sits between them. Astra is rolling out this week into ChatGPT Plus, Pro, Business and Enterprise, the API and AWS — on launch day only Daybreak trusted-access enterprises had it. Astra Pro is for Pro, Business and Enterprise. On Business, standard seats get limited Astra usage within their existing allowance; premium seats get the full allowance.

So the mental model is simple. Sol for the day-to-day; Astra when the job is hard enough to justify a model that costs 2.5 times as much per token. That price gap is the whole reason the calculator further down exists.

Where each OpenAI model sits
ModelGenerationWhere you meet it
GPT-6 AstraGPT-6 flagshipChatGPT Plus, Pro, Business and Enterprise, the API (gpt-6-astra) and AWS — rolling out this week
GPT-6 Astra ProGPT-6 flagship, heavierChatGPT Pro, Business and Enterprise
GPT-5.6 SolGPT-5.6 everydayDefault on ChatGPT Plus and Pro; promo API price US$4 / US$20 per million tokens
GPT-5.6 TerraGPT-5.6 everydayRemains in the ChatGPT lineup between Sol and Luna
GPT-5.6 LunaGPT-5.6 everydayDefault on the free tier
What is different under the bonnet

Three things worth knowing. First, it is OpenAI's largest training run to date and the first pre-trained on more than 100,000 GPUs, at the Stargate site in Texas. Second, it is the first OpenAI model where earlier models substantially supervised the training. Third — and this is the one that matters for the next lesson — it uses what OpenAI calls "opaque recurrence": the model loops internally over the same query rather than writing its reasoning out, so there are fewer readable chain-of-thought traces to inspect. Clever for capability. Awkward for anyone whose job is checking what the model was thinking.

What it is for

OpenAI claims state of the art on computer use, browsing, software engineering, cybersecurity, science and professional work. The launch demos were pointedly practical: a PCB layout in KiCad, a 3D city in Unity, an animated transmission in FreeCAD and Blender, a tax-return draft from a W-2, contract formatting, browser form entry and website QA, and Excel and Power BI work. In Codex, Astra can keep notes across context windows and search earlier messages and tool output (experimental behind a config setting for now, default in the coming weeks), and can ask you a question without stopping the unrelated work it is doing.

What to watch

Astra is the first OpenAI model to reach the Preparedness Framework "Critical" threshold for cyber capability. The public version refuses exploit discovery; the advanced cyber capability lives behind the Daybreak trusted-access programme. The practical consequence for the rest of us is that legitimate, long-running agent work — including non-cyber tasks — can be slowed, paused or stopped by the safeguards. In ChatGPT and Codex you are asked to review and continue; in the API the task stops. OpenAI's own line: "At launch, this is something that people should expect." Plan your workflows accordingly.

The Frontier Four

Four flagships in one week: Claude Fable 5.1 (1 Sep), Gemini 3.8 Flash (2 Sep), Meta's Muse Spark 1.3 (2 Sep) and GPT-6 Astra (3 Sep). Click a card to expand it, then scan the table. Where I could not verify a figure from the launch material, the cell says so rather than guessing.

Prices are list API prices per million tokens in US dollars. Benchmark figures come from each vendor's launch material; OpenAI ran its comparisons at max effort unless noted. "Not published" means I could not verify a figure from the launch material — check the vendor's model page before you rely on it.

Price per task, not price per token

The number on the pricing page is the price per million tokens. The number you actually pay is the price per finished task — and a model that gets there in fewer tokens, with less reasoning effort and fewer retries, can be cheaper even when its list price is higher. The calculator is pre-loaded with a worked example. Change anything.

×1 low effort, first try×5 max effort, several retries

Scales the output tokens. Use it to stand in for a higher reasoning-effort setting, or for the extra passes when the first answer was not good enough.

Estimated cost for this task
$0.00

The worked example. Example A runs a three-step task on GPT-5.6 Sol at its promo price: 30,000 tokens in, 15,000 out, at ×3 effort because it took a few goes — about US$1.02. Example B runs the same task on Astra at ×1: 30,000 in, 12,000 out, done first time — about US$0.90. Astra's tokens cost 2.5 times as much and the task still came out cheaper. The token counts are illustrative, not measured — which is exactly why the exercise below asks you to measure your own.

Hands-on exercise

Run the same task in two flagships

Pick a real three-step task from your own work. Run it — identically — in two of the flagships you have access to. Note the tokens (most tools show them in a usage or developer view), how long it took, and score the quality honestly. The card at the bottom does the arithmetic and gives you something to paste into your notes.

Reading the launch

The numbers in this lesson are the vendors' numbers. Some survive scrutiny and some need a footnote the size of the chart — the 98.6% on ARC-AGI-3 was run in a harness OpenAI itself showed can triple scores; the FrontierMath benchmark was funded by OpenAI; the coding chart left Muse out. The AI Fundamentals course has a lesson that takes the Astra launch apart claim by claim, and it is worth twenty minutes before you sign anything.

Lesson 6.5 — Reading a Frontier Launch: the Astra Case

Check your understanding

Score0 / 5
Answer all 5 to see your result