Marius Hobbhahn

speaker
481 appearances 1 recordings 1 series first heard Sep 2024 last heard Sep 2024

Marius Hobbhahn’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
problem in this particular way, but it won't hack you.
It will not try to accumulate power.
It will not try to do all of these other things.
Whereas if you train it on outcomes and you train it on hard tasks, on long, like really long run run tasks, I think you have all of these like
instrumentally convergent goals where in order to achieve this long-term goal, right, you wanna have more power, you wanna have more money, you wanna have less constraints on you so that you can do more things and learn more things along the way and so on.
I feel really this is where the danger is coming from.
And it's the flattery aspects of the I I think they're like, it's clearly suboptimal.
We should do something about it, but it's not the main thing that I'm worried about, if that makes sense.
had done before?
I have absolutely no insights in what exactly they're doing, what the setup exactly looks like, and so on.
I think reading between the lines of what people say publicly, what the model card says and so on, like just to give you some intuitions, w what I imagined this roughly to look like is they take, let's say, GPD 40 is my guess as like a base model, and then they do nor normal agentic training, and then they also start to
Do outcome based R L training.
On reasoning, and maybe they also reward the reasoning in in in some more incremental fashion where you have another model judging this or something.
And sort of I think the main hints at this are Noam Brown and Greg Bruckman in like various different tweets very explicitly saying that they have reinforcement learning on outcome and longer trajectories.
And then also Noah Brown just historically being a person that for the last, I don't know, many years, maybe more than five probably, just being just saying something like.
Like, hey guys, inference compute is really undervalued.
You can get a lot of alpha out of this, and we can definitely make this.
We can definitely make use of more inference compute to get models to do the kind of thing.
So my guess is that it's some mix of
inference compute trickery plus then actually training models to use this inference effectively, but this is all based on the public tweets and so on.
Showing 81–100 of 481 · page 5 of 25 ← Previous Next →