OpenAI said its upcoming AI model Astra marks an improvement in coding and computer application operations. According to Sina Finance, a person familiar with Astra's development said a new technique that improves model performance also means the model and similar models will reveal less of their "thinking" process, making it harder to monitor signs of bad behavior.
OpenAI is using a new technique called recurrent depth, or recurrent Transformer, which lets AI models improve answers by processing the same text multiple times. Unlike the most advanced models on the market, which show in text how they "think" through a task before completing it, the new approach obscures part or all of the AI's reasoning process, or chain of thought.
According to Sina Finance, OpenAI has limited the use of recurrent depth in Astra, so the model still produces readable chains of thought and researchers can monitor its reasoning. OpenAI said in a blog post on Tuesday that Astra will launch with "additional chain-of-thought monitoring to quickly detect and contain" potential misconduct.
OpenAI Chief Scientist Jakub Pachocki said on X on Tuesday night that although monitoring model chains of thought is "fragile" and "unfortunately moving in a negative direction," he hopes to stop the industry from racing to build models that do not produce readable reasoning for AI researchers. He said strengthening such monitoring is "a core goal of our current research agenda."