Kevin Roose

speaker
381 appearances 2 recordings 1 series first heard Oct 2025 last heard 6 Aug

Kevin Roose’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Aug OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Aug 2026 with 1.

Appearances

newest first · ▶ plays the moment
Like, this is what the X-Files were set up for. matched
There was an autonomous AI agent, and it's hacking into it. matched
It's like, oh, God, where's Mulder? matched
Get me Mulder! matched
They put up a blog post and they explain that they had been using GPT 5.6 Sol, which is a model that you can now use, plus an unnamed more capable pre-release model. matched
And they were testing an internal benchmark called Exploit Gym, that's GYM. matched
And it is basically a test of how good a model is at discovering new vulnerabilities and creating exploits so that, for example, it could hack into somebody else's system. matched
But then it broke out of its cell without them noticing it? matched
That's right. matched
So if nothing else, PJ, it did pass the test. matched
And from there, it is able to start hacking other computers on OpenAI's network until finally, and crucially, it finds one that has internet access. matched
Yeah, because basically the model, and here I'm going to use anthropomorphizing language that's going to drive listeners insane, so I do apologize. matched
But, you know, the model, I'm speaking metaphorically here, essentially thinks to itself, hey, I need to solve this problem. matched
Where might I find the answer to this problem? matched
I bet Hugging Face, the company that stores all of the data sets, including for all these various benchmarks that I'm being tested on, I bet I could find the information there. matched
And so that is why it goes to Hugging Face, and it is then able to mail itself in through the front door of the company. matched
Just a few days before the disclosure about Hugging Face, OpenAI published another blog post where they revealed that an internal model spent about an hour finding a vulnerability in a sandbox so that it could post its results to GitHub. matched
I'm not sure why it wanted to post, but it did. matched
The important thing there is it had been explicitly instructed to only post to Slack, but it just sort of ignored that instruction. matched
And then in April, Anthropic's Mythos model had found some sort of multi-step hack that let it get out of its sandbox, get onto the internet, and actually it emailed a researcher. matched
Showing 41–60 of 381 · page 3 of 20 ← Previous Next →