Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
So in the limit, if you can keep improving your overseer, if the overseer becomes sufficiently smart...
then you can recover the performance of what you would get with just straight RL backpropagating through time, but without the reward hacks.
In practice, we're probably not going to get that far, but I think you could actually push this quite far, especially just by, I think, the most simple thing of get the AI system to explain why its plan is good is, I think, a very simple baseline that I think can make this quite effective.
I think plausibly.
It only matters once you start doing reinforcement learning over multi-step trajectories over a reasonably long period of time.
That's something that I think has only really started this year and to varying amounts at different companies.
It's...
not totally clear how much reward hacking is a big problem.
But yeah, to the extent that the sort of multi-step reward hacking does become a big problem, I think it's quite plausible that this should be used now.
And in fact, I think at current capability levels, my guess would be that the non-myopic approval part, rather than being a competitiveness hit as it would be in something like AlphaZero,
I would guess that it would actually improve capabilities overall, although at the cost of you need a lot more human input to provide those rewards.
Not just because of that.
I mean, that is definitely one reason.
I think also reinforcement learning has a credit assignment problem, where you do a ton of different actions, and then at the end, you get a reward.
And it's the RL algorithm's job to figure out, based on that reward, which of these actions are actually most relevant to that reward.
A difficult problem.
Yeah, it's a difficult problem.
And in some sense, the thing that the RL algorithm does is like, eh, we're just going to make everything more likely to happen if the reward was positive, or it was unusually good, and everything less likely to happen if it was unusually bad.
Whereas with something like MONA, you can be much more granular.
You can say like, this particular part where you wrote these tests, those were some really great tests.
Showing 621–640 of 1,071 · page 32 of 54 ← Previous Next →