Marius Hobbhahn
speaker
481 appearances
1 recordings
1 series
first heard Sep 2024
last heard Sep 2024
Marius Hobbhahn’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
The third component here was or the third question is you asked was where should we expect this to come from in the first place?
And
I think almost all of the worry for me at least comes from reinforcement learning, outcome based training, where the default story that I might have in mind for what the AI companies want to do with their AIs is they want them to be more and more capable, but also more and more autonomous.
So right now an AI can maybe solve tasks that would take a human, a few minutes, maybe thirty minutes in some cases, but this is already like f some more on the extreme end of things.
And then like extremely rarely something would take a human maybe an hour.
And
My guess is that the A A GI companies will want to just continue this and at some point
They will be trained for or they will be able to do set tasks for eight hours or like a full week just on their own.
You'd give them write this.
This mini research sprint, do this mini research sprint, they will just go ahead and do that.
And so what the way in which I expect us to come there is through outcome-based training, where we tell the model, hey, write that research report, and then we reward it based on whether the final outcome, the result that it ended up with was good or bad.
And my guess is that from this, the model will just become very outcome-oriented, very goal-driven, goal-directed.
And that somewhere in this sort of science loop
It just becomes more and more consequentialist.
It cares more and more about this output.
and cares less and less about the kind of guardrails and the surroundings and so on that we and the constraints give it, right?
Like maybe the way the fastest way to get to this output is just by cutting a bunch of corners here and there on the ethical uh side of things or um like explicitly removing some of the guardrails because they slow it down.
Or just in other ways like constrain uh ignoring the constraints that humans give it.
Uh one one very interesting example here is the Sakana AI paper on the automated scientist.
They had a section where they were like, we also gave the model access to its own code when it did the science and it just changed, it started another copy of itself somewhere because it wanted to run a longer experiment and it knew that it would otherwise not be able to finish the experiment or something along those lines.
Showing 41–60 of 481 · page 3 of 25
← Previous
Next →