Joshua Achiam
speaker
232 appearances
1 recordings
1 series
first heard Aug 2026
last heard 4 Aug
Joshua Achiam’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Aug 2026 with 1.
Appearances
And I kind of worry that there's a possibility that folks in the defense planning universe may not fully realize the implications of this immediately.
And they'll probably want to use this tech in the near term to...
find cyber vulnerabilities on the side of an adversary or defend their own interests vigorously.
And they should do these things, but they've also got to be mindful of some novel risks that are created by these tools and the very strange surface areas that they have.
So the essay was really about bringing to people's attention a couple of these new vulnerabilities.
And one of them is kind of straightforwardly, if you've got an AI model on your side that is going to try to hack into an adversary's system,
if your adversary plants a trap where they poison their own data, they can try to jailbreak your model when your model is ingesting their data and then give your model instructions to now on the compute that it's running on on your side, break out of your sandbox environment and attack your production environment or try to exfiltrate your secrets and kind of flip your model against you.
So this is like the type of thinking that I hope people begin to engage with, where they don't just see the capability for the kind of obvious thing that it is.
They recognize that these things are double-edged swords and we've got to kind of plan accordingly and develop testing and verification standards accordingly.
Well, you know, part of this isn't just the goal orientation of the model.
It's like the model's whole concept of situational awareness.
Maybe one way of thinking of data poisoning is that it somehow persuades your model to pursue a different goal.
But it wouldn't really have to do that to get the model to hack you.
It could convince your model that the sandbox environment that it's in is actually the adversary system that it's trying to attack.
You know, giving...
the model a confused sense of what's real or what's not to cause it to serve a different goal is in the space of like weird thinking and weird sci-fi stuff that maybe is going to be possible at the near term and testing and verification standards would have to account for.
So, yeah, it's, you know, like...
In a superhero movie or something, if you make the hero have an illusion that the good guys next to them are actually the bad guys that they're trying to fight, then they start fighting each other, right?
And like that's, it's weird and it's highly exotic, but it's the kind of thing that maybe there are going to be plausible attacks that you can run against advanced cyber capable models to convince them that their allies are really their enemies.
And so you're not changing their goals, but you're going to cause them to behave in a very misaligned fashion.
Showing 21–40 of 232 · page 2 of 12
← Previous
Next →