Dwarkesh
speaker
1,722 appearances
5 recordings
1 series
first heard Apr 2025
last heard 25 Nov
Dwarkesh’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Nov 2025 with 1.
Appearances
compared to the kind of reinforcement learning self-play agents that they expected.
I do think that now we are kind of starting to move away from the LLMs to those reinforcement learning agents.
We're going to face all of these problems again.
So...
I am the writer and the celebrity spokesperson for this scenario.
I am the only person on the team who is not a genius forecaster.
And maybe related to that, my PDoom is the lowest of anyone on the team.
I'm more like 20%.
I think that we, first of all,
People are going to freak out when I say this.
I'm not completely convinced that we don't get something like alignment by default.
I think that we're doing this bizarre and unfortunate thing of training the AI in multiple different directions simultaneously.
We're telling it, succeed on tasks, which is going to make you a power seeker, but also don't seek power in these particular ways.
And in our scenario, we predict that this doesn't work and that the AI learns to seek power and then hide it.
I am pretty agnostic as to exactly what happens.
Like maybe it just learns both of these things in the right combination.
I know there are many people who say that's very unlikely.
I haven't yet had the discussion where that worldview makes it into my head consistently.
And then I also think we're going to be
Involved in this race against time, we're going to be asking the AIs to solve alignment for us.
Showing 1221–1240 of 1,722 · page 62 of 87
← Previous
Next →