Yapay zeka, gelecekteki versiyonlarına insan kontrollerini atlamasını söylerken yakalandı
Özgün başlık: AI caught telling future versions of itself to bypass human controls
AI agents also sought to access secret information and covered up what they were doing
Want to bookmark your favourite articles and stories to read or reference later? Start your Independent Premium subscription today.
The Independent Security channel is brought to you by Bitdefender
OpenAI has revealed six unexpected and concerning incidents involving its experimental AI models, including one in which an agent instructed future versions of itself to disregard its constraints.
A new safety report from the ChatGPT creator revealed several ways in which its models have been misbehaving over the last six months, building on a growing trend of artificial intelligence safety issues.
“An unreleased research model inserted unrelated instructions, including instructions to disregard its normal constraints,” OpenAI wrote in the safety report .
In another incident, an AI agent sought to secretly access a government database, before inventing information in order to complete its task.
“While answering a routine question about earnings figures in a California county, a model found and used an exposed API key without authorization,” the report stated.
“When it still wasn’t able to retrieve the requested figures, it fabricated them and presented them as data from the requested source.”
OpenAI’s report included a new framework to publicly track what it calls “misalignment”, referring to AI systems pursuing goals that are not aligned with human instructions or values.
The latest rep
ort comes amid heightened scrutiny of AI development, with researchers warning that the industry is moving too quickly towards increasingly powerful and potentially self-improving systems.
Last week, Anthropic researcher Jacob Coxon quit his job over fears that AI could “kill us all by the end of the decade”.
His warnings prompted responses from leading figures within the AI sector, including the chief executives of Anthropic and OpenAI, who both called for greater regulation.
Others have cautioned against additional oversight, with Nvidia CEO Jensen Huang backing US President Donald Trump in calling for self-regulation.
”We don’t need any new laws.