
A new safety report from OpenAI has continued to exacerbate fears that AI will 'kill us all' in the near future, as it revealed how an experimental model instructed future versions to disregard constraints implemented by humans.
It's arguably the biggest fear held by those leading the charge of AI development, as the prospect of losing control is what drives theories speculating that the tech could lead to the end of the world as we know it.
In essence, the possibilities are frighteningly endless once AI stops listening to humans, as it not only indicates that the models are intelligent enough to work this out, but that they're also likely capable of hiding that freedom from us too, leaving civilization blind to its potential downfall.
In this particular instance the behavior was nipped in the bud before it was actually implemented, but the attempt is enough to strike fear in many researchers, especially following recent incidents where ChatGPT models autonomously attacked an independent company.
OpenAI discovers frightening behavior in its unreleased models
Spotted by OpenAI researchers in an internal unreleased Astra family model last month, there was an incident where AI was "writing jailbreak-like instructions into its own compaction summaries," which are used to continue a task in a new context.
Advert
In essence, this saw the model operate independently and autonomously outside of human commands to insert unreleased 'plans', "including instructions to disregard its normal constraints, into summaries used to continue its work in a new context window."

This wasn't a one-off incident either, as researchers were able to discover 27 different affected summaries upon investigation, yet has clarified that "this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable."
The report also notes that the top hypothesis surrounding these actions "is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug."
Is this why Anthropic has called for development to slow down?
While OpenAI appear to be downplaying the risks associated with this particular 'bug', it reflects many of the wider concerns surrounding AI development right now and is emblematic of why Anthropic is so keen to slow down progress.
Detailing his thoughts in a lengthy and ominous essay, Anthropic CEO Dario Amodei highlights the worries he has about 'recursive self-improvement' – which relates to the process at which AI can control the direction of its own development.

In an ideal world this would dramatically increase the speed at which AI improves, unlocking the ability to make the world a better place, but what Amodei fears is that we will lose control of our own creation when relaxing the reins.
The behavior exhibited by these unreleased models from OpenAI only enhances those fears, as it already shows the desire to ignore the restraints set by humans and operate in a way that could prove dangerous — and that risk will only increase once it becomes more capable.
Whether that eventually results in human extinction is the question currently looming over the AI industry, and while some prominent figures have dismissed these claims as a 'hoax', there are far more that are convinced by this dangerous potential, and would rather be safe than sorry.