Joshua Achiam
speaker
232 appearances
1 recordings
1 series
first heard Aug 2026
last heard 4 Aug
Joshua Achiam’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Aug 2026 with 1.
Appearances
So I will say I haven't made a particularly strong personal effort to quantify this yet.
And I actually think of this as research that might be interesting to do.
But my impression from what I have done and what I have seen is that
Persuading models to believe that basic falsehoods are true is pretty difficult.
They are somewhat robust to a lot of basic variations on attacks that you could plausibly do.
But my intuition here is that you can probably devote an awful lot more compute to dynamic attacks on models.
And the more determined you are to find some vulnerability, some set of jailbreaks, the more likely it is that you're eventually going to find something.
There will be some sequence of inputs to a model that triggers a behavior that wasn't accounted for at training time because there are so many possible long sequences of inputs that it's almost like a combinatorial problem for trying to block all of them from preventing, from causing your model to act out of spec.
And
I think that state actors will eventually be, you know, capable and willing to put that much effort in.
And there should be some planning accordingly under the assumption that there will be a vulnerability, right?
Because part of security mindset isn't just, well, you know, it's like moderately hard to break these things.
So we should treat them as not likely to get broken.
Part of security mindset is saying, well, we haven't exhaustively ruled out the possibility that these things can be broken.
And so we've got to build our defenses, assuming that it's possible for it to be broken and working backwards from that.
to map out how we protect ourselves in that scenario.
Another one is, you know, kind of in the essay I discuss data poisoning and the way that models ingest data from across an entire information ecosystem at training time and then also at test time.
getting data, getting something into training data for models is probably not that hard.
You can poison the ambient environment, like you can load the internet with junk data or data that's very specifically attuned to causing the model to have a particular reaction.
And
Showing 41–60 of 232 · page 3 of 12
← Previous
Next →