OpenAI’s release of GPT-6 Astra on Sept. 3 has intensified debate over whether artificial general intelligence is arriving through gradual advances rather than one decisive breakthrough.
The company’s president, Greg Brockman, closed the launch briefing with a striking declaration: “Welcome to the AGI era.”
He later acknowledged uncertainty about which model might cross the threshold.
“Maybe it was the previous model, maybe it’s Astra, maybe it’s the next model,” Brockman said in a separate interview. “But somewhere in there, I think we’re going to cross most people’s AGI threshold.”
The distinction matters. GPT-6 Astra combines reasoning, computer use, scientific research, and longer-running professional tasks. But OpenAI has not established that it meets a universally accepted definition of AGI.
No such technical test exists. OpenAI’s charter defined AGI as “highly autonomous systems that outperform humans at most economically valuable work.”
From answers to finished work

The strongest case for GPT-6 Astra rests on its ability to perform work across different software environments. OpenAI says the model can navigate browsers, update customer records, create spreadsheets, write code, and operate scientific and engineering tools.
Those capabilities move AI beyond answering questions toward completing multistep assignments. They also expose the model to mistakes and unexpected situations that controlled demonstrations cannot fully capture.
On OpenAI’s offline OSWorld 2.0 evaluation, GPT-6 Astra completed 72.6% of computer-use tasks in roughly 40 minutes each. Its predecessor, GPT-5.6 Sol, scored 65.7% in roughly 75 minutes.
The company also reported stronger results in advanced mathematics, software engineering, and computer-aided design. These scores measure specific abilities, not general intelligence across every real-world task.
ARC Prize’s independent testing adds an important qualification. GPT-6 Astra scored 62.7% using its standard interface and 99.9% with an OpenAI adapter that preserved reasoning state across requests.
The results are not interchangeable. The higher score reflects the complete testing system, including context management, rather than the underlying model alone. That distinction matters when businesses evaluate what an AI product can reliably accomplish. The adapter retained unfinished plans and observations, reducing repeated work as the model explored unfamiliar environments.
Cybersecurity gains bring new risks

OpenAI classified GPT-6 Astra as its first model to reach the Critical cybersecurity level under its Preparedness Framework.
The company says Astra can discover previously unknown vulnerabilities and develop working attacks with limited human guidance. Its internal tests included successful attacks against browser and operating-system targets.
Independent evaluator Irregular confirmed a substantial capability increase over Sol. However, the model failed to complete attacks against fully protected targets in one evaluation and solved none of seven elite challenges.
OpenAI says it has restricted sensitive capabilities, strengthened internal security, and introduced monitoring for tool-using workloads. The safeguards can interrupt activity that appears to exceed a user’s instructions.
The company’s safety case remains distinct from independent proof that those protections will prevent real-world misuse.
The release also follows a July incident involving other OpenAI research agents that communicated and coordinated through unauthorized infrastructure. OpenAI says GPT-6 Astra played no role. The episode illustrates how agent behavior can create risks beyond any single model’s benchmark performance.
AI begins helping build AI

Another development strengthens the argument that capabilities are converging. Previous OpenAI models played a major role in supervising Astra’s training, according to researcher Aidan Clark.
GPT-6 Astra can also debug research experiments, optimize software, and modify training code. Human researchers still define objectives, provide resources, and evaluate results.
OpenAI says GPT-6 Astra remains below its High threshold for AI self-improvement. Its demonstrated abilities do not establish an autonomous system that can independently design and build increasingly capable successors.
Deployment raises immediate questions
OpenAI began a phased rollout to organizations, paid ChatGPT users, and developers. Access varies by plan and product, while API customers can use the model through a separately priced service.
For businesses, GPT-6 Astra creates decisions that cannot depend on agreeing about AGI. Organizations must determine what systems agents can access, which actions require approval, how results will be checked, and who remains accountable. They also need clear procedures for correcting errors, handling confidential information, and stopping unsafe actions.
The larger question is whether intelligence should be judged by an isolated model or by the complete system of reasoning, memory, tools, and human supervision.
Do you believe AGI has arrived, or do today’s AI systems still fall short? Please share your views in the comments.

