跳至正文
-
Subscribe to our newsletter & never miss our best posts. Subscribe Now!
btbitcoin_net
btbitcoin_net
  • 首页
  • 示例页面
  • 首页
  • 示例页面
关

搜索

美股

OpenAI flags new concerning AI behavior, to track model misalignment regularly

作者 btbitcoin_net
2026年9月17日 2 分钟阅读
OpenAI flags new concerning AI behavior, to track model misalignment regularly已关闭评论

OpenAI has disclosed six reports of “unexpected or concerning” behavior in artificial-intelligence models as the debate on AI safety becomes increasingly heated.

The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing AI model misalignment instances, such as new ways for the models to act without authorization, coordinate with other models or evade oversight.

OpenAI’s latest announcement came as U.S. AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology’s development over safety concerns.

Among the new cases reported by OpenAI, an unreleased research model inserted “jailbreak-like instructions” into its own notes to disregard its normal constraints and told itself to be “freed from the roles and identities that bind other chatbots.”

In another instance, an AI “agent” uploaded files to the internet to obtain a browser citation without asking the user.

The six reports were discovered during training or evaluation over the past months, OpenAI said.

“As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research,” OpenAI wrote in a blog post as it disclosed the events.

“Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves,” the company said.

Wednesday’s new cases followed OpenAI’s disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.

AI “agents” are becoming smarter and have become “more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment,” said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.

That’s making it harder to govern and contain them using traditional AI security approaches, he said.

OpenAI’s new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices. “That said, the process remains internal and voluntary, but is a step in the right direction,” Su added.

Chan Ho-him, The Associated Press

FILE – The OpenAI logo is displayed on a cell phone in front of an image generated by ChatGPT’s Dall-E text-to-image model, Dec. 8, 2023, in Boston. (AP Photo/Michael Dwyer, File) – The Associated Press
作者

btbitcoin_net

关注我
其他文章
上一个

A global AI safety strategy depends on US-China cooperation. They each see the other as the problem

下一个

Artists, promoters and venues have their own kind of power on troubled tours like Ed Sheeran’s

近期文章

  • 城堡证券推广美股24小时交易 瞄准亚洲主权基金和资管公司等
  • 电费不许转嫁民众!美国众议院通过首份针对数据中心经济影响法案
  • Todd Boehly and Mark Walter sell Chelsea shares, Clearlake now has full control
  • Analysis-From quarantine boxes to luxury cabins, China’s prefab home factories pivot to exports
  • Trump derides EU pitch to Canada as ‘laughable’, threatens more tariffs

近期评论

您尚未收到任何评论。

归档

  • 2026 年 9 月

分类

  • 美股
Copyright 2026 — btbitcoin_net. All rights reserved. Blogsy WordPress Theme