OpenAI Debuts GPT-6 Astra Amid Mixed Benchmark Results and Higher Costs

In independent IA Index measurements, GPT-6 Astra scored about 60–61 points and sits roughly 14th among ~202 models, with a cost per Intelligence Index task around $0.96 (median model ~36), highlighting that Astra is not the clear top performer on this index despite OpenAI's generational framing.
ARC-AGI-3 results show a dramatic gain when using a Provider Adapter harness, with Astra reaching about 99.9% versus 62.7% under the standard harness, illustrating how preserved reasoning across turns can materially boost agentic performance (though exact numbers vary by report and setup).
The ARC-AGI-3 leaderboard as of recent reports places Claude Opus 5 at roughly 30.2% top performance, with OpenAI Sol at 7.78% officially (38.3% with certain API settings); there are no confirmed Astra ARC-AGI-3 scores yet, and a viral claim of 98.6% remains unverified.
Astra appears to leverage memory-enabled agentic behavior, turning unfamiliar environments into compact symbolic world models and using a domain-specific language to track state and plan actions, contributing to its observed gains in ARC-AGI-3 under favorable harnesses.
Astra's deployment characteristics include a 1M-token context window, support for text and image input with text output, and initial rollout through Daybreak Access (with broader access via AWS Bedrock and Microsoft Azure), factors that influence cost and integration for users.
OpenAI has unveiled GPT-6 Astra, its newest flagship model, claiming it is itechpost 'the most intelligent and aligned system the company has ever built.' However, independent benchmarking reveals a more complex picture: Astra scores roughly on par with OpenAI's prior Sol model on the Artificial Analysis Intelligence Index, ranking around 14th among 202 models at a cost of $0.96 per task — well above the median of $0.36.
The model shows meaningful gains in specialized agentic scenarios, such as ARC-AGI-3 testing with advanced reasoning tools, where it can reach 99.9% accuracy compared to 62.7% under standard conditions. Yet Astra still trails competitors like Claude Opus and other leaders, and rfi reports the rollout has begun only to select customers via limited access.
The Artificial Analysis Intelligence Index, an independent measure, places Astra at approximately 60–61 points. This ranks it roughly 14th among roughly 202 models — far from a clear top performer despite OpenAI's generational framing. For comparison, the median model score sits around 36 points, meaning Astra does outperform the average but faces stiff competition from established leaders.
Cost efficiency tells another story. Astra costs $0.96 per Intelligence Index task, significantly higher than the typical model at $0.36. This price premium raises questions about whether the incremental performance gains justify the higher expense for users and enterprises.
When tested on ARC-AGI-3 benchmarks using a Provider Adapter harness — a tool that preserves reasoning across multiple turns — Astra reaches approximately 99.9% accuracy. This represents a dramatic jump from 62.7% under standard testing conditions, showing that preserved reasoning across turns materially boosts performance.
The leaderboard context is important. Claude Opus 5 currently leads ARC-AGI-3 at roughly 30.2% under official testing, while OpenAI Sol scores 7.78% officially. No confirmed Astra scores exist on the public leaderboard yet, and a viral claim of 98.6% remains unverified by independent sources.
Astra appears to leverage memory-enabled agentic behavior differently from prior models. gadgetpilipinas reports the system converts unfamiliar environments into compact symbolic world models and uses a domain-specific language to track state and plan actions. This architecture explains its dramatic gains in agentic evaluation scenarios.
The model supports a 1-million-token context window, allowing it to process large amounts of text and images as input while producing text output. yahoo notes Astra excels at computer-use and browsing tasks, suggesting these capabilities were central to the design.
rfi confirms OpenAI has begun rolling out GPT-6 Astra to select customers, with the company claiming it has built in safeguards. Initial deployment is through Daybreak Access, a limited program designed for early users and testing partners before broader availability.
Broader access will follow via AWS Bedrock and Microsoft Azure, expanding availability beyond OpenAI's direct channels. The higher cost per task and tiered rollout strategy suggest OpenAI is positioning Astra as a premium offering for complex, mission-critical workloads rather than a general-purpose replacement for existing models.
Publishers
22
Articles
19
Reach
41