Marius Hobbhahn
speaker
481 appearances
1 recordings
1 series
first heard Sep 2024
last heard Sep 2024
Marius Hobbhahn’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
then the model will very quickly just refuse to act on its own goal and it's just like happy to take the goal of its developer.
And so we we try to have situations where either both goals are like clearly
positive and it's plausible that there are trade-offs between them, but it's not clear that one strategy is is better than the other, or one goal is better than the other, or where
The goal that the model is given from us uh is has positive balance, because otherwise you immediately get into refusals or just HHH training failures, and then we don't know whether it has the capability or not, or whether it's uh uh like an alignment failure.
And we really for this we specifically wanted to test whether it has the capability.
So we were fine making this trade off.
Yeah, interesting.
respond to it?
As long as we can agree on what the right goal is and we have a goal that is always correct, I would probably be f be fine with the model vehemently defending this goal.
I think the main problem is.
As a society and as the world and as the philosophers and everyone together, we haven't really figured out what the right goal is.
Some people will say it's liberty, other people will say democracy, other we will people will say economic growth.
Other like most people will say like some weird mix between all of this, but I can specify the individual components and like the trade offs and so on.
So if we were able to find the right goal, I would probably be fairly fine.
I I think it's just like a really hard process.
And the main thing I'm worried about is that
Either we misspecify the goal or the model somehow learns a different goal that is just like
not the one we wanted it to have.
So for example, maybe it conver it instrumentally converges on something like if I had more power or more money, that would make it easier to do my terminal goal.
And then it's just now I have another goal.
Showing 361–380 of 481 · page 19 of 25
← Previous
Next →