OpenAI's GPT-6 Astra has landed with impressive benchmark scores that demonstrate significant advances in AI reasoning and task execution. On the ARC-AGI-3 benchmark, which measures an AI's ability to adapt to entirely new situations, Astra scored 99.9%—far outpacing Claude Opus 5's 30.2% and GPT-5.6 Sol's 7.8%.
The model also excels at complex professional-level work. On the Agents' Last Exam, which evaluates the capacity to complete sophisticated computer tasks similar to human professionals, Astra achieved 59.3%, compared to Claude Fable 5's 48.7% and GPT-5.6 Sol's 53.6%. For programming specifically, Terminal-Bench 4.0 tested the ability to solve complex problems using the terminal, including software engineering work—Astra scored 57.9% versus Claude Fable 5.1's 55.8% and GPT-5.6 Sol's 37.3%.
GPT-6 Astra will roll out over the coming days to users on Plus, Pro, Business, and Enterprise plans, as well as via the API. Pricing is set at $10 per million input tokens and $50 per million output tokens.