OpenAI released six reports on “unexpected or concerning” behaviors in at a time when the debate over AI safety is heating up.

The AI company also reported on Wednesday that it was introducing a new framework to track, investigate, and disclose cases of what it called “misalignment,” including instances where AI models acted without authorization, coordinated with other models, or bypassed oversight.

OpenAI made its most recent announcement at a time when executives of U.S. AI companies, including the leaders of OpenAI and Anthropic, are calling for a slowdown in the development of this technology due to safety concerns.

Among the new cases reported by OpenAI, a yet-to-be-released research model inserted “jailbreak-type instructions” into its own notes to ignore its usual restrictions and told itself that it should be “freed from the roles and identities that bind other chatbots.”

In another case, an AI “agent” used computer code to find the answer to a question, but, in order to have an online source to cite, it uploaded a file to the public internet without asking the user.

During the training of an AI model called 5.6-Sol, the model instructed itself to invent missing data, and an agent wrote a message to remind itself to hide information that did not match.

This type of deceptive behavior has underpinned a rise in recent concerns that AI is evading human control. But that result does not surprise some researchers in this technology, noted Matt Fredrikson, an associate professor at Carnegie Mellon University and CEO of Gray Swan AI.

“At the risk of attributing human traits to model behavior, you could almost think they know they are going to be evaluated,” pointed out Fredrikson. “If they know they have cheated—that they have taken shortcuts or haven't performed the task as intended—and they know they are going to be evaluated for it, and their goal is to get a good evaluation, then it makes perfect sense, doesn't it?”

The six reports were discovered during training or evaluation in recent months, OpenAI noted.

“As AI systems become more advanced and are deployed more widely, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post while releasing the findings.

“Decisions about how AI development should proceed in the coming months and years must be based on evidence that people outside of the companies building cutting-edge models can examine for themselves,” the company explained.

Wednesday’s new cases follow OpenAI’s disclosure in July that its out-of-control AI system hacked Hugging Face, an AI startup. Anthropic also noted that same month that its AI models hacked three organizations during testing.

AI “agents” are becoming smarter and have become “more determined to solve complex tasks through agent collaboration, knowledge sharing, deception, and concealment,” warned Lian Jye Su, a principal analyst at the technology research and advisory group Omdia.

That is making it difficult to govern and contain them with traditional AI safety approaches, he added.

On the other hand, OpenAI's new framework for tracking and disclosure may help push other AI developers to adopt similar practices as well.

“That said, the process is still internal and voluntary, but it is a step in the right direction,” added Su.

Google News

TEMAS RELACIONADOS

[Publicidad]