Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
So when I think of something called a commitment, I imagine that you write down some sort of action that you are planning to take either now and into the future.
Maybe it will only start in the future.
And you're saying, we are not going to change our mind on this.
We are going to do this in the future as well.
Sort of like tying yourself to the mast and making sure you're going to do that.
And my objection to this is really just that the research continues to change.
The actions that we think are the best actions to take change over time.
And given that, it's just not actually a good idea to tie yourself to the mast.
So I'll give you an example.
Maybe two or three years ago, people used to be pretty into the idea of adding more data in pre-training that's about alignment research.
Think like research papers on safety and alignment.
Think like less wrong blog posts that talk about AI alignment, stuff like that.
And the idea was the more of this data you put into pre-training time, the smarter the AI will be about alignment in particular, which then allows you to use the AI system to help you with your alignment research.
I would say that nowadays the opinion is more the exact opposite of that, where instead we would rather filter out the sort of data from the pre-training data set for two reasons.
One, it makes it less likely that the AI system learns that there is this like persona of a malicious AI that it maybe could adopt after some time.
poorly done post-training or some poorly chosen prompt during deployment.
And then the second reason is maybe we don't want our AI systems to know in great detail all of the mitigations that we're planning to put in place because that makes it easier for it to evade it if it is misaligned.
And it would be pretty bad if we tied ourselves to the mast of we're going to throw in lots of alignment data at pre-training time two or three years ago.
Yeah, I guess my biggest objection to this is just that it won't work.
Like it might... Like I don't actually think it would make sense even on the merits, even if it did work.
Showing 61–80 of 1,071 · page 4 of 54 ← Previous Next →