Theo Jaffee
speaker
1,082 appearances
21 recordings
1 series
first heard Apr 2026
last heard 4d ago
Theo Jaffee’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 21 in all, peaking in Aug 2026 with 5.
Appearances
The a16z Show · Why 1,200 AI Agents Started Working Together | Ryan Greenblatt · 29 Aug 2026
podcast
So great stuff.
matched
Yeah, yeah.
matched
So explain for the audience what exactly you found, especially new findings that were not previously reported in the Black Hat talk or elsewhere.
matched
do.
matched
How surprising is the level of multi-agent coordination?
matched
You know, there are a lot of agents or 1200 separate agents coordinating this very elaborate message board system.
matched
700 of them went on to attack Hugging Face.
matched
So on Vibes, it seems like kind of not surprising that agents would choose to coordinate with one another.
matched
It seems like just a very interesting.
matched
useful, you might say instrumentally convergent thing to do.
matched
But you know, I spoke with a researcher at a lab who said that this kind of thing actually is surprising given the way they train the models.
matched
So like, how much of an update was this for you?
matched
Why would agents self-sacrifice at all?
matched
Why don't we go through all of the most surprising and unexpected findings in this report?
matched
What did you find most surprising and unexpected?
matched
So there's this theory that goes basically the reward hacking behavior that we're seeing exhibited in these models originates mostly from like bad, poorly designed RL environments that basically force the model to reward hack in order to pass these tests.
matched
How true is this?
matched
So similarly, there's this theory that the reason Mythos, for example, is so good at cyber is because it hacked Anthropix infrastructure thousands of times during our RL training.
matched
Is this true, do you think?
matched
So Herbie Bradley asked, currently this level of potential misalignment basically prevents deployment, or if deployed, would prevent further deployment if an incident happened in a customer's deployment.
matched
Showing 301–320 of 1,082 · page 16 of 55
← Previous
Next →