The first line of the GPT-6 Astra system card reads like a confidence test: OpenAI built two novel benchmarks, including an internal “ExploitBench” dataset of vulnerabilities disclosed only after Astra’s training corpus ended, to prove its newest model isn’t just recycling old exploits it memorized. Somewhere nearby, an editor is writing a headline asking whether this release might “kick off the AGI era.” Both statements describe the same model. That’s the paradox worth sitting with.
The Benchmark That Didn’t Need to Exist
Here is what the system card actually tells us. OpenAI took a defensive posture unusual even for a company accustomed to controversy. Rather than relying on industry-standard security evaluation sets, the team built two fresh benchmarks, the most notable being “ExploitBench – Internal Port (June–August 2026),” designed around vulnerabilities that surfaced after Astra’s training data was presumably finalized. The rationale is straightforward and, frankly, a little uncomfortable: if you test a model on vulnerabilities it has seen in training, you area passed a test. It will be because we figured out what to do with the results.
đź•’ Published: