Garrison Lovely

speaker
631 appearances 3 recordings 1 series first heard Jul 2026 last heard 15 Sep

Garrison Lovely’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Sep OctJan 26AprJulnow

Recordings per month over the last 12 months — 3 in all, peaking in Sep 2026 with 2.

Appearances

newest first · ▶ plays the moment
Um, and so this is like, yeah, just kind of an unbelievable story across the board.
Um, yeah, happy to go into more detail on any of it.
Yeah, I mean, I think that if it were given like a biological, you know, understanding evaluation and then it's like, well, I'm going to hack and find the bio answer key, then like that's more evidence, right?
That it's just generalizing to like, I am going to do something you did not intend that's unrelated to the activity that you're testing me for.
So that's worse.
But I think it's also bad that OpenAI wanted the model to do this evaluation the right way, which would not involve cheating, stealing the answer key.
And they definitely didn't want their model to do crimes and to hack another multi-billion dollar company.
And so the fundamental point is that these companies, the most sophisticated in the world in many respects, are not able to reliably steer or control their own models.
And those models are now autonomous and capable enough to take actions in the real world that could have really big consequences.
And those consequences are just going to grow with time as the models get more capable and autonomous.
uh and the problem just gets harder right because like they'll just they're superhuman hacking in many respects now but they could be like super superhuman in the next generation and you'll need the very best models to keep up and like hugging face was struggling to fend off this attack because they couldn't use the models that open ai and anthropic like the very best models on the market because those models are like trained to not help with this type of thing because they can't distinguish reliably whether like you're actually defending against an attack
or you're using that as a cover to do your own attack.
And so Hugging Face had to use open weight models, which are less capable.
And then opening eyes, like we let them into our trusted access partner program as a result of this.
And it's like, it feels like a mob boss.
Yeah, it's like, oh, like nice code base you got there.
Like shame if something were to happen to it.
Yeah, yeah.
So these open-weight models, I mean, when they're published online, they have guardrails built into them.
The thing is, it's trivial to remove them if you have any technical capacity, and it costs not very much time or money to do so.
Showing 521–540 of 631 · page 27 of 32 ← Previous Next →