Dwarkesh

speaker
1,722 appearances 5 recordings 1 series first heard Apr 2025 last heard 25 Nov

Dwarkesh’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Nov OctJan 26AprJulnow

Recordings per month over the last 12 months — 2 in all, peaking in Nov 2025 with 1.

Appearances

newest first · ▶ plays the moment
I think there are two reasons why they might fail to do what you want, kind of reflecting how they're trained.
One is that they're too stupid to understand their training.
The other is that you were too stupid to train them correctly, and they understood what you were doing exactly, but you messed it up.
So I think the first one is kind of what we're coming out of.
So GPT-3, if you asked it, are bugs real, it would give this kind of hemming-hawing answer like, oh, we can never truly tell what is real.
Who knows?
Because it was trained kind of don't take difficult political positions and a lot of questions like is X real or things like is God real where you don't want it to really answer that.
And because it was so stupid, it could not understand –
anything deeper than like pattern matching on the phrase is X real?
JPT-4 doesn't do this.
If you ask are bugs real, it will tell you obviously they are because it understands kind of on a deeper level what you are trying to do with the training.
So we definitely think that as AIs get smarter, those kind of failure modes will decrease.
The second one is where you weren't training them to do what you thought.
So for example, let's say
You're hiring these raters to rate AI answers.
You reward them when they get good ratings.
The raters reward them when they have a well-sourced answer, but the raters don't really check whether the sources actually exist or not.
So now you are training the AI to hallucinate sources, and if you consistently rate them better when they have the fake sources,
then there is no amount of intelligence which is going to tell them not to have the fake sources.
They're getting exactly what they want from this interaction, metaphorically, sorry I'm anthropomorphizing, which is the reinforcement.
Showing 1321–1340 of 1,722 · page 67 of 87 ← Previous Next →