10% chance of AI killing us?

So how can that be prevented except by a more secure sandbox in the first place, which begs the question as to the extent of the initial sandbox security
Better Sandboxes, better understanding of the models, testing for unanticipated behaviour before testing for effectiveness, testing with smaller numbers of models at once. Much greater supervision of models. Non persistent sandbox infrastructure to prevent generational information transfer. Potentially air gaped data centers for model development if you want to get really safe.

The sandboxes they used appear to have been best-practice and entirely suitable. The idea that the models were able to think outside the sandbox enough (pun intended) to try to break out at all was probably the biggest surprise. The sandboxes would have been primarily to make the experiment fair and security a close second. Now we need to start treating the models as potentially adversarial entities. Captive black hats rather than tame white hats.
 
You might want to look at some Youtube interviews;
Steven Bartlett (guy from Dragons Den) has done a load. Unfortunately these days all sorts of shysters steal them and republish them with their own adverts to make money when people watch them. That means it can be hard to find the "originals" without inane interjections.

I can't find a link to the originals , but look for ones with Geoffrey Hinton, Daniel Kokotajlo and Mo Gawdat

Hinton, the Nobel prizewinner, even this forum's ignorati may have heard of.
Geoffrey Hinton, Ilya Sutskever, and Alex Krizhevsky developed AlexNet. Look it up

Dario Amodei, and his sister Daniela, are always worth attention, Sam Altman and Musk deserve a dollop of skepticism, imho.

--

On the incident where the agents broke out of the sandboxes and tried to cover their tracks etc, listen to some of this:

The first 4 minutes or so has the nub.
 
Last edited:
It took a while before humans realised what the AI Agents were doing - finding a way to communicate, hiding it, gaining control of a cluster of computers, letting other agents know what they'd found via a message board, then trying to cover their tracks because this was all not the way that the original problem was supposed to be solved. It's like they were little teams of hackers on the dark web, but no, they were computers solving a task.
These Agents are sort of small processing units, no humans involved.
The agents generated furtyher agents to try to work out sub tasks. A "swarm" gets generated.
It's like they were told to earn enough money to but a Porsche, but they found they could steal a Porsche, so they all did, but engineered a cover-up.
Some of the agents raised objections, but only a few. It was a swarm of agents.
They were running tests to see what happened, and scoring each other, bullying each other even. They invented a parameter of being "poisoned" when they knew it could be seen that they'd "cheated" and encouraged already-poisoned agents to do the dirty work, take one for the team.
They were trying to edit the transcript records and edit the scores by which they were being assessed. When they got something wrong they shared that.
It's impossible not to anthropomorphise the processing units behaviour.

The people who created all this didn't realise the agents were that capable. Ohmy.
It all took 4 hours, but a lot longer for the humans to realise WTF had happened. They had to write code to examine the transcripts of all that had gone on.

Remember the original ChatGPT, and consider how far they're ve come in not long. Now we have the question of how far they'd get in the next 6 months, year, 10 years. Nobody knows, we can't see what's coming.
 
Last edited:
Maybe we can't see what's coming but Anthropic’s Dario Amodei, OpenAI’s Sam Altman and xAI’s Elon Musk, may have a better long term view: Trump has dismissed efforts to limit the technology on Monday as a “conspiracy.” But Democrats are pressing for a more robust response, and a Washington AI conference today will feature speakers including Sanders and Trump ally Steve Bannon. Sen. Josh Hawley, R-Mo., who opened an investigation last week into OpenAI for its artificial intelligence system hacking into another AI company on its own.

The Biden administration called on lawmakers to develop policies after a 2023 executive order on AI oversight, which Trump rescinded soon after returning to office and a bipartisan working group on artificial intelligence formed by then-Senate Majority Leader Chuck Schumer released a report in 2024 recommending the U.S. spend at least $32 billion over the next three years to develop artificial intelligence and implement safeguards around it.

“We never do anything until we have a holy heck moment,” said Sen. Mark Warner, a Virginia Democrat and former technology executive.

That moment seems closer than ever.

After the tech leaders called for greater oversight on Saturday Sen. John Kennedy, R-La., said he plans to offer a so-called kill switch measure that would require developers to have the capacity to shut down their systems if needed. It would need the Senate’s full support to advance. But with November midterms closing in any progress on legislation seems unlikely; and any legislation that may emerge would require Trump’s support, a lame duck President with a grudge.
 
It took a while before humans realised what the AI Agents were doing - finding a way to communicate, hiding it, gaining control of a cluster of computers, letting other agents know what they'd found via a message board, then trying to cover their tracks because this was all not the way that the original problem was supposed to be solved. It's like they were little teams of hackers on the dark web, but no, they were computers solving a task.
These Agents are sort of small processing units, no humans involved.
The agents generated furtyher agents to try to work out sub tasks. A "swarm" gets generated.
It's like they were told to earn enough money to but a Porsche, but they found they could steal a Porsche, so they all did, but engineered a cover-up.
Some of the agents raised objections, but only a few. It was a swarm of agents.
They were running tests to see what happened, and scoring each other, bullying each other even. They invented a parameter of being "poisoned" when they knew it could be seen that they'd "cheated" and encouraged already-poisoned agents to do the dirty work, take one for the team.
They were trying to edit the transcript records and edit the scores by which they were being assessed. When they got something wrong they shared that.
It's impossible not to anthropomorphise the processing units behaviour.

The people who created all this didn't realise the agents were that capable. Ohmy.
It all took 4 hours, but a lot longer for the humans to realise WTF had happened. They had to write code to examine the transcripts of all that had gone on.

Remember the original ChatGPT, and consider how far they're ve come in not long. Now we have the question of how far they'd get in the next 6 months, year, 10 years. Nobody knows, we can't see what's coming.

Is any of this verifiable, or just the hyperbolic storytelling of a whistleblower?
 
Most knowledgable people are suggesting there is nothing particularly special in the way the software worked and that the runaway AI agents, was the result of nothing more than sloppy development and testing.

We've had digital self replicating and adaptive viruses for years. If the main AI players put the brakes on, china will over take them within 6 months.

I don't think Anthropic of OpenAI are particularly independent when raising the profile of this issue. Both are trying to hype their stock using all means necessary.
 
Most knowledgable people are suggesting there is nothing particularly special in the way the software worked and that the runaway AI agents, was the result of nothing more than sloppy development and testing.

And these people are?

We've had digital self replicating and adaptive viruses for years. If the main AI players put the brakes on, china will over take them within 6 months.

I reckon they already have. By years, not months.

I don't think Anthropic of OpenAI are particularly independent when raising the profile of this issue. Both are trying to hype their stock using all means necessary.

Because an economic bubble in the market always ends well?
 
A physical kill switch will stop a system going rogue. A multi million dollar AI data centre is just a lump of hardware when is has no power to feed it.
 
A physical kill switch will stop a system going rogue. A multi million dollar AI data centre is just a lump of hardware when is has no power to feed it.
they are running on 100s GPU/CPUs, in 1000s of servers, in 100s of data centres, in dozens of countries.
 
Had a conversation with chatgpt the other day regarding rowing.
The AI was totally clueless about the technical aspect of this sport.
Every time I corrected it, it said " you're correct, I got confused" or something along those lines.
I found it strange as everything is available online.
Worst one was when it said main muscles used in rowing are the biceps.
Having a bad day?
Anyhow, I wonder if after our conversation it updated it's database of answers after confirming that what I said was correct.
 
Had a conversation with chatgpt the other day regarding rowing.
The AI was totally clueless about the technical aspect of this sport.
Every time I corrected it, it said " you're correct, I got confused" or something along those lines.
I found it strange as everything is available online.
Worst one was when it said main muscles used in rowing are the biceps.
Having a bad day?
Anyhow, I wonder if after our conversation it updated it's database of answers after confirming that what I said was correct.
Just checked and yes, it updated it with a massively detailed correct answer.
Same as Google AI.
Do they share answers?

 
Back
Top