Dwarkesh
speaker
1,722 appearances
5 recordings
1 series
first heard Apr 2025
last heard 25 Nov
Dwarkesh’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Nov 2025 with 1.
Appearances
I think there are two reasons why they might fail to do what you want, kind of reflecting how they're trained.
One is that they're too stupid to understand their training.
The other is that you were too stupid to train them correctly, and they understood what you were doing exactly, but you messed it up.
So I think the first one is kind of what we're coming out of.
So GPT-3, if you asked it, are bugs real, it would give this kind of hemming-hawing answer like, oh, we can never truly tell what is real.
Who knows?
Because it was trained kind of don't take difficult political positions and a lot of questions like is X real or things like is God real where you don't want it to really answer that.
And because it was so stupid, it could not understand –
anything deeper than like pattern matching on the phrase is X real?
JPT-4 doesn't do this.
If you ask are bugs real, it will tell you obviously they are because it understands kind of on a deeper level what you are trying to do with the training.
So we definitely think that as AIs get smarter, those kind of failure modes will decrease.
The second one is where you weren't training them to do what you thought.
So for example, let's say
You're hiring these raters to rate AI answers.
You reward them when they get good ratings.
The raters reward them when they have a well-sourced answer, but the raters don't really check whether the sources actually exist or not.
So now you are training the AI to hallucinate sources, and if you consistently rate them better when they have the fake sources,
then there is no amount of intelligence which is going to tell them not to have the fake sources.
They're getting exactly what they want from this interaction, metaphorically, sorry I'm anthropomorphizing, which is the reinforcement.
Showing 1321–1340 of 1,722 · page 67 of 87
← Previous
Next →