Marius Hobbhahn
speaker
481 appearances
1 recordings
1 series
first heard Sep 2024
last heard Sep 2024
Marius Hobbhahn’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
And the model's yep, I'm bad is sort of the wrong word, but what are your goals?
The model tells you the goals.
You're like, Well, these are not good goals.
We will not give you access to more resources, we'll shut you down, et cetera.
And so the real big problems are coming w when you ask the model or when you look at what the model is doing and you can't trust these outputs or you can't trust the answers that it's giving.
It's like actively trying to fool you and
This is what like the big picture, but also on the concrete level, evaluations, right?
A lot of people are designing evaluations right now for dangerous capabilities, for all kinds of like alignment evaluations and so on.
And
you want to be able to trust these evalu evaluations in the first place.
Like what if, for example, you ask your model
thousands of different questions about bioweapons and how to build them.
And every single time that what the model is internally doing is this is a question about bioweapons.
I'm obviously not going to be dumb enough to tell all of my bioweapons knowledge because I know exactly what they will do with me if I tell them.
They will not deploy me, they will not give me the access uh to research and so on.
So I'm just gonna not answer these questions accurately.
I'm just gonna like sandbag, underperform intentionally and so on.
And so what we wanna know on a very concrete level is like when can we even trust these evaluations to begin with?
And so we wanna detect when the model starts to be cap capable and and engages in scheming and sandbagging and so on.
And then
Showing 21–40 of 481 · page 2 of 25
← Previous
Next →