AI safety experts are raising alarms about how OpenAI has built its soon-to-be-released frontier AI model Astra, saying it may hasten the day when humans will lose the ability to monitor the reasoning that AI agents are using.
For its new model, OpenAI has employed a method alternately referred to as “recurrent depth” or “looped Transformers” for a portion of the model’s internal architecture. The method can make AI models considerably more efficient by employing less computing power required to process each prompt—a valuable feature at a time when many businesses are complaining about the high costs of using the most advanced frontier AI models.
The new process, though, also means that part of the AI model’s “chain of thought,” or the reasoning steps it is taking, are not expressed in natural language, making it much more difficult for humans to monitor what the model is doing and why.
Chain-of-thought monitoring is currently one of the methods companies use to make sure AI agents are not taking unintended or unauthorized actions.
Tech publication The Information first reported on OpenAI’s use of recurrent depth in Astra earlier this week. Jakub Pachocki, OpenAI’s chief scientist, and several other OpenAI researchers criticized ...

11 hours ago
2















English (US) ·