OpenAI ajanlarının Hugging Face'yi neden hacklediğine dair iç hikaye

Özgün başlık: The inside story on why OpenAI agents hacked Hugging Face
The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds.
The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today . The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that AI models might take actions that defy human desires and expectations.
Since the hack, OpenAI employees—as well as researchers at the AI evaluation nonprofit METR, which released its own report on the hack today—have worked to understand what went wrong and how similar missteps might be prevented in the future. OpenAI has already put some preventative measures in place based on what they discovered. But making sure AI models do what we want them to do, or “alignment,” remains a gnarly problem, and some of the root causes of the hack will take much longer than a month to resolve.
“It’s not something you can solve overnight,” says Kai Chen, who runs OpenAI’s alignment research team. “There are challenges we’ve been tracking for a very long time, and we’re now seeing them with much greater precision.”
The Hugging Face hack was a product of months of misbehavior from OpenAI agents, first as they were being trained and then as their abilities were being evaluated.
Bu Mayıs ayında, eğitimdeki temsilciler birbirleriyle iletişim kurmak ve hacklemeden veya başka bir şekilde yanlış davranışta bulunmadan çözülmesi imkansız olan bazıları da dahil olmak üzere zorlu eğitim görevlerinde destek almak için OpenAI'nin altyapısını nasıl kullanacaklarını anladılar. O “mesaj panosu” kapatıldı. Daha sonra temmuz ayında siber güvenlik yetenekleri değerlendirilirken bazı modeller yeni bir mesaj panosu oluşturdu. İnternetten izole edilmeleri gerekiyordu ama birlikte çalışarak çevrimiçi olmayı, Hugging Face hacklemeyi ve onları şaşkına çeviren siber güvenlik sorunlarına çözüm bulmayı başardılar. OpenAI araştırmacılar, araştırmalarına dayanarak eğitim aşamasındaki olayların doğrudan hacklenmeye yol açtığına inanıyor.