Rohin Shah
speaker
1,071 appearances
1 recordings
1 series
first heard Jun 2026
last heard 2 Jun
Rohin Shah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
what they can do is a lot of stuff in parallel.
And so pre-training is really heavily optimized to be able to do everything in parallel, which means that the opaque serial depth does actually have to be pretty small in order to get those efficiency gains.
And so this is a pretty strong structural economic reason that, at least for pre-training, you're going to get AI systems that have relatively small opaque serial depth
And so you might think of them as like, you know, there's a good reason for models to be born speaking English or some other natural language.
Yes, I think that's right.
I think for technical folks, they might interchange the words wide and deep in your sentence, but yes.
Cool.
Yeah, so that was stage number one, which is like, why should we expect the opaque serial depth to stay low, at least during the pre-training phase?
Then there's step two, which is like, OK, great.
We're saying that the pre-trained model has to essentially speak in English in order to do reasoning.
But then there's a bunch of post-training and RL and all of this sort of stuff.
Maybe that's going to make it so that even if the model is speaking in tokens, maybe it will start speaking some sort of alien language that we don't understand.
And here I would say like, you know, there's no theoretical argument or there's no like, you know, theorem you can prove that will say, no, they're not going to do this.
But I will say, you know, they are born speaking English in some sense.
And so, and pre-training is like by far the like most powerful form of getting stuff into an AI system that we have ever built.
And so the model is incredibly good at speaking in natural language.
It's great at doing reasoning, but only the kind of reasoning that humans do when writing stuff down.
And then when you look at the reasoning training that we're doing today, there's a paper from, I think, Neil Nanda and others,
that shows that actually a substantial part of what the reasoning training is doing is just teaching the model when to do a specific kind of reasoning step that it had already learned during pre-training.
And so like basically a lot of the capability is just coming in pre-training.
Showing 381–400 of 1,071 · page 20 of 54
← Previous
Next →