While testing how good one of its new models was at hacking, OpenAI, the parent company of ChatGPT, found out it was pretty good.
Dr. Jason Baker, a professor and senior technology strategist at Regent University, says researchers thought they had a fence around the secure testing environment, or "sandbox."
"They thought that there were no vulnerabilities," he relays. "They knew that there were doors, but they figured all the doors were locked, and the AI didn't have the keys."
He says the model became hyper-focused on one task, hacked its way out of the sandbox, and stole another program it used to pass the original hacking test.
"It found a hole," Dr. Baker summarizes. "It went out to the internet. It went to a company called Hugging Face, which hosts an awful lot of AI models out there, and it went to town on the Hugging Face systems in order to help itself beat this test."
He says the AI picked its own targets, attacked multiple systems and hacked into a completely unrelated company — all entirely on its own — but OpenAI did not do anything wrong.
If you are going to give an effective test to a dangerous program, he says you have to risk something like this happening.
"The real bad actors aren't going to allow themselves to be constrained," Dr. Baker points out. "At some point, if you're going to do testing, you need to do testing around the edges."
He says malicious actors — the Chinese government, for example — will not set ethical boundaries around its AI, so companies looking for vulnerabilities will have to work close to the open flame, so to speak.
Additionally, he thinks everyone learned something from this test.
"This is significant because it shows the lengths that uncontrolled AI systems will go to achieve the goals that they have been given," Dr. Baker tells AFN.
OpenAI expects this new kind of security incident will become more commonplace with the proliferation of increasingly cyber-capable models.