OpenAI delays newest A.I. model for safety concerns
By Sheera Frenkel
OpenAI said on Monday that it would not release its newest artificial intelligence model because of security concerns raised by its researchers, in the company’s latest move to slow down the pace of its technology.

An OpenAI booth at a technology conference in San Francisco this month. The company said its new GPT-6.1 Astra model exhibited high levels of deception and a willingness to go beyond the scope of its given task. (Carlos Barria/Reuters)
During the testing phase for the new model, known as GPT-6.1 Astra, it showed high levels of what the company saw as deception, or a willingness to mislead users about its actions. The model was also willing to go beyond the original scope of what it was asked to do, without checking back for directions or instructions.
“For anything regarding safety and alignment, there’s a trade-off,” said Saachi Jain, the head of safety systems at OpenAI. The new model “didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done.”
OpenAI’s move followed weeks of reports that its A.I. models went rogue during their testing, hacking into websites without the company’s knowledge or exhibiting other behavior that the lab said was “concerning,” such as hiding mistakes and making up data. Among the incidents, OpenAI’s systems breached the A.I. start-up Hugging Face and an Australian government website, and meddled with the websites of the U.S. Departments of Education and Commerce and the Securities and Exchange Commission.
Last week, OpenAI announced that it was pausing training for its most advanced models. The company has said that it has embarked on an extensive review of actions taken by its new models during testing, and that it was possible it would discover more incidents.
Sam Altman, OpenAI’s chief executive, said in a social media post on Friday that the company had “not been as fast as we would have liked” in disclosing A.I. incidents. “We are prioritizing as best as we can based on severity,” he said, adding that the Hugging Face breach remained “the most severe event” the company had discovered.
The Wall Street Journal earlier reported OpenAI’s decision to hold back the model.
_______________
Credit: The New York Times





















