Dwarkesh
speaker
1,722 appearances
5 recordings
1 series
first heard Apr 2025
last heard 25 Nov
Dwarkesh’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Nov 2025 with 1.
Appearances
Right.
There's actually a really interesting foretaste of this.
At some point...
Somebody asked Grok, like, who is the worst spreader of misinformation?
And it responded... I think it just refused to respond to Elon Musk.
Somebody kind of jailbroke it into telling it it's prompt, and it was, like, don't say anything bad about Elon.
And then there was enough of an outcry that the head of XAI said, actually, that's not consonant with our values.
This was a mistake.
We're going to take it out.
So we kind of want more things like that to happen where people are looking at, like...
Here it was the prompt, but I think very soon it's going to be the spec where it's kind of more of an agent and it's understanding the spec in a deeper level.
And just thinking about that and being, and if it says like, by the way, try to manipulate the government into doing this or that, then we know that something bad has happened and
If it doesn't see that, then we can maybe trust it.
This is actually part of our misalignment story, is that if the AI is sufficiently misaligned, then yes, we can tell it it has to follow the spec.
But just as people with different views of the Constitution have managed to get it into a shape that probably the founders would not have recognized, so the AI will be able to say, well, the spec refers to the general welfare here.
I think we agree.
I think that's kind of why all of our policy prescriptions are things like more transparency, get more people involved, try to have lots of people working on this.
I think our epistemic prediction is that it's hard to maintain classical liberalism as you go into these really difficult arms races in times of crisis.
But I think that our policy prescription is let's try as hard as we can to make it happen.
Yeah, so I agree that the AIs are currently getting more reliable.
Showing 1301–1320 of 1,722 · page 66 of 87
← Previous
Next →