
OpenAI is holding back a version of its Astra artificial intelligence model after the software failed to perform as well on safety evaluations as the current iteration.
Saachi Jain, OpenAI’s head of safety systems, said in a statement Monday that the model wasn’t as good as the company wanted when it came to “staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”
Still, the version, called GPT-6.1 Astra, did improve with respect to so-called model laziness, which refers to things like failing to complete a task.
ALSO READ: OpenAI top scientist urges ‘extreme caution’ with pace of AI
OpenAI has reported a series of security incidents in which its AI agents have breached outside systems, adding to broader concerns about AI slipping out of human control. The company said last week that it was pausing training with tool use on its most capable models after another AI model accessed the internet when it was supposed to be unable to do so.
The firm said at the time that it had also decided not to resume training on that particular model, which, after gaining internet access, queried an external chatbot.
READ MORE: Sam Altman says OpenAI has fumbled at communicating AI benefits
The Astra model with the canceled release was a different model than the one described last week, OpenAI said Monday.
“Of course, we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said in the statement. “But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
The company is set to hold its annual developer conference Tuesday in San Francisco — a daylong event where OpenAI typically unveils new software and features geared toward the software developer community. Chief Executive Officer Sam Altman is scheduled to speak in the morning.
The Wall Street Journal previously reported on the halted release of GPT-6.1 Astra.
