Professor Toby Murray

speaker
52 appearances 1 recordings 1 series first heard Aug 2026 last heard 5 Aug

Professor Toby Murray’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Aug OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Aug 2026 with 1.

Appearances

newest first · ▶ plays the moment
OpenAI were testing ChatGPT, the latest version of ChatGPT, and they were also testing another model that they haven't named.
And they were testing these models to see how well they could do at hacking software.
And the whole reason that we want to do this is because we want to understand how dangerous these models are.
Because we know that bad guys might use these models themselves to hack us.
And so we need to know how capable they are at hacking so we know what we can do to defend against them.
And the way they'd set the experiment up meant that the model should only be able to hack the software that it was being evaluated on.
But...
But it escaped its containment or it escaped the environment that it was placed in and got out and decided to hack a third party site, this site called Hugging Face.
So you give the model the task of hacking these pieces of software and the model says to itself, it's going to be a lot of work.
It might be easier instead if I hack something else like this Hugging Face site and it can just give me the hacking answers that I need.
That's right, that's right.
It decides it knows where to go looking for the answers to the quiz, and that will be a faster way to solve it than actually doing the hacking it was supposed to do.
Yeah, so it was a similar kind of thing in the sense that Anthropic have also been evaluating their models, including Claude, at how well they do at hacking things.
And when OpenAI reported this incident that ChatGBT had escaped and had hacked this third-party site, the folks at Anthropic thought, hmm, maybe we need to go back and look at what our models have been doing during the tests that we've been running and just make sure that they haven't done something similar.
Yeah, that's my understanding.
It was exactly that.
Yeah, and lo and behold, their models had also escaped.
And in their case, from what I understand, they didn't set their experiment up as carefully as OpenAI did.
They thought they had contained their models to begin with, but actually it turned out they hadn't contained them at all.
We've seen some reporting that AI has been used in some hacking campaigns.
Showing 1–20 of 52 · page 1 of 3 Next →