Why less visibility into how OpenAI’s new GPT-6 Astra ‘thinks’ is sparking safety concerns
The Hugging Face breach in July highlights the importance of being able to inspect what models are ‘thinking’, say analysts
When announcing Astra on Thursday, OpenAI said it was “the world’s most intelligent and aligned model,” with a “significant jump in cyber capabilities”.
OpenAI president Greg Brockman said at the end of a press call announcing Astra’s arrival that it likely represents AGI, or artificial general intelligence – AI that matches or outperforms human intelligence.
However, OpenAI also said the model’s written reasoning was “harder to monitor” compared with GPT-5.6 Sol, the previous generation released in July.
“We have found that GPT-6 Astra is more capable of controlling its own CoT (chain of thought) than GPT-5.6 Sol, and less likely to include incriminating information in its CoT,” OpenAI said, referring to the intermediate reasoning steps an AI generates while solving a task.
The shift in visibility stems from a technique known as recurrent depth, or looped transformers, which reuses parts of a neural network. As a result, it processes complex logic inside hidden mathematical loops rather than in step-by-step readable text, The Information reported on Tuesday ahead of the launch.



