Rohin Shah
speaker
1,071 appearances
1 recordings
1 series
first heard Jun 2026
last heard 2 Jun
Rohin Shah’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
I think mostly our contribution was showing that this actually works with existing LLMs and giving examples of that happening.
Because as far as I know, there weren't actual experiments demonstrating this before.
Yep, that's right.
I mean, to some extent, the answer is you don't get around it.
Like, ultimately, both the incredible, amazing creative insights that are awesome and the incredible creative insights that are reward hacks
look basically the same to you as the observer.
And if you want to stop the reward hacks, yeah, you also give up some amount of the competitiveness.
That being said, I do think that we call it myopic optimization with non-myopic approval.
So the non-myopic approval part is basically the part where you say, when you're grading this particular step, ask some sort of intelligent overseer, a human, possibly an LLM,
to judge whether this particular step, how good will it be for getting future reward?
And the overseer should then take into account everything that it knows and judging how good this is going to be for the future.
And what this guarantees is whatever incentives from the future affect the AI system in this particular step have to be things that the overseer understands.
Well, I would suggest that they don't look at whether the company made a bunch of money.
And instead, they look at the action that the AI system takes and they predict to themselves, will this lead to a bunch of money?
Yes or no.
And based on that, provide a reward.
Now, the AI system on step one can do lots of things to help the overseer with this.
It can give an explanation of how its plan is going to lead to tons of money in the future.
And then the overseer just has to verify it.
you can use things like debate or other AI assistance to make the overseer better at predicting what's going to happen in the future so that they can give better rewards.
Showing 601–620 of 1,071 · page 31 of 54
← Previous
Next →