Rohin Shah
speaker
1,071 appearances
1 recordings
1 series
first heard Jun 2026
last heard 2 Jun
Rohin Shah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
We know that the pre-training for good reason is going to like stay the way that it is and be done using human-like reasoning and speaking in English.
And then RL is just like really inefficient relative to pre-training.
And for it to like build an entirely new epistemic language that is something that like
you know, we wouldn't be able to understand even with a decent chunk of effort is just so far beyond what RL is doing currently that it would be pretty surprising to see that in the near future.
Now, do I expect to see it ever?
Yes.
But I think six months ago, I think I split a room in a conference by saying, agree or disagree, chain of thought monitoring will continue for two years.
I think I said two years.
It might have even been one year.
I forget.
Whereas I'm like, I don't know, my median is probably like four years, five years.
I don't know.
It's not totally clear.
mostly because of this argument about how hard it would be to go beyond the current situation.
I agree that sort of thing might work in the relatively near future.
I'm not sure it diminishes the value of chain of thought monitoring all that much.
If you were in fact just saying, we're going to keep a probability distribution of tokens and feed that back in,
you can still inspect that probability distribution of tokens and interpret them again as normal English reasoning.
It's harder because now you have this like larger set of things that the model could be thinking about at any given time.
But I would still expect that it's like not too bad and you would be able to do some monitoring of it.
Showing 401–420 of 1,071 · page 21 of 54
← Previous
Next →