OpenAI released GPT-6 Astra on September 3rd in what could be a quantum leap in AI capability, but it is not AGI, and they themselves did not claim the status. Astra is a multimodal, tool-using digital agent that achieves extraordinary benchmark scores: 98% on FrontierMath Tier 4, 99.9% on ARC-AGI-3, 100% on ExploitBench, and a "Critical" cybersecurity classification. The model completes OSWorld tasks in 40 minutes versus 75 minutes for its predecessor, operates across coding, scientific workflows, CAD design, and professional software, and has been integrated into real-world products from Legora (financial review) to Playco (game prototyping). However, OpenAI's own charter defines AGI as "highly autonomous systems that outperform humans at most economically valuable work," and Astra has not demonstrated that breadth, reliability, or long-horizon autonomy. The public record shows a frontier system with emerging AGI capabilities, not a proven AGI.
Compared to previous models and competitors, Astra establishes a new tier of agentic performance. Against Claude (cited at 52.6% on Terminal-Bench Science), Astra scores 64.6%; against GPT-5.6 Sol (83.3% on BenchCAD), Astra hit 95.9%; and its ARC-AGI-3 performance allegedly exceeds human action-efficiency baselines on 96% of levels. The system's tool ecosystem enables it to complete multistep workflows that earlier models could not. Yet these comparisons are harness-dependent and vendor-reported, while independent replication is yet to be seen. Crucially, the Polymarket betting market currently places the probability of AGI being achieved by year-end 2026 at approximately 18-22%, a growing figure that still put market skepticism at the forefront that any current system constitutes true AGI.
Where does this leave us? Astra is best described as a "near-AGI precursor" being exceptionally close on digital task slices like cybersecurity, coding, and tool-mediated reasoning, but materially short on the dimensions that define strict AGI. OpenAI's own system card reveals unresolved monitorability issues, residual prompt-injection risks, and adversarial failure modes. Still, excitement does not stop there, with OpenAI announcing 10,000 of its AI agents running on a post-GPT-6 model solved the 90-year-old Navier–Stokes Millennium Prize Problem. It is genuinely unbelievable that an AI system constructed and formally verified a 165-page proof of a finite-time fluid breakdown in a mere 88 hours. The focus is now on future models built on top of the previous, giving confidence that AGI might come from OpenAI after all.
Sources: OpenAI, Reuters, ArcPrize
Photos: Unsplash
Written by: Ariff Azraei Bin Mohammed Kamal