Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
And they won't release it. Why? Because chains of thought are scary. Right. And they are legitimately scary. Right. If you look at R1, it flips back and forth between Chinese and English. Sometimes it's gibberish. And then the right answer comes out. Right. And like for you and I, it's like great.
doing this it's amazing I mean you talked about that sort of like chain of thought for that philosophical thing which is not something they trained it to be philosophically good it's just sort of an artifact of the chain of thought training it did but like that's super important in that like Can I inspect your mind and what you're thinking right now? No.
And so I don't know if you're lying to my face. And chain of thought models are that way, right? Like this is a true quote unquote risk between, you know, a chat application where, hey, I asked the model to say, you know, bad words or whatever, or how to make anthrax. And it tells me that's unsafe. Sure. But that's something I can get out relatively easily.
What if I tell the AI to do a task and then it does the task automatically? all of a sudden randomly in a way that I don't want it, right? And now that has like much more task versus like response is very different, right? So the bar for safety is much higher. At least this is Anthropic's case, right? Like for deep seek, they're like ship, right? Yeah.
And they killed that dog, right? And all these things, right? So it's like.
And there's an interesting aspect of just because it's open-weighted or open-sourced doesn't mean it can't be subverted, right? There have been many open-source software bugs that have been like... For example, there was a Linux bug that was found after 10 years, which was clearly a backdoor because somebody was like, why is this taking half a second to load? This is the recent one.
Why is this taking half a second to load? And it was like, oh crap, there's a backdoor here. That's why. And it's like, this is very much possible with AI models. Today, the alignment of these models is very clear. I'm not going to say bad words. I'm not going to teach you how to make Anthrax. I'm not going to talk about Tiananmen Square.
I'm not going to, you know, things like, I'm going to say Taiwan is part of, you know, is, is just an Eastern province, right? Like, you know, all these things are like, depending on who you are, what you align, what, you know, whether, you know, and even like XAI is aligned a certain way, right? You know, there, they might be, it's not aligned in the like woke sense.
It's not aligned in like pro China sense, but there is certain things that are imbued within the model. Now, when you release this publicly in an instruct model, that's open weights, um, this can then proliferate, right? But as these systems get more and more capable, what you can embed deep down in the model is not as clear, right?
And so that is like one of the big fears is like if an American model or a Chinese model is the top model, right, you're going to embed things that are unclear. And it can be unintentional too, right? Like British English is dead because American LLMs won, right? And the internet is American and therefore like color is spelled the way Americans spell it, right?
This is just- This is just the factual nature of the LLS now.
It is. Taking it as something silly, right? Like something as silly as the spelling, like which British and English, you know, Brits and Americans will like laugh about probably, right? I don't think we care that much. But like, you know, some people will, but like this can, this can boil down into like very, very important topics. Like, Hey, you know, subverting people, right?
You know, chatbots, right? Character AI has shown that they can, like, you know, talk to kids or adults. And, like, it will, like, people feel a certain way, right? And that's unintentional alignment. But, like, what happens when there's intentional alignment deep down on the open source standard? It's a backdoor today for, like, Linux, right?
right, that we discover, or some encryption system, right? China uses different encryption than NIST defines, the US NIST, because there's clearly, at least they think there's backdoors in it, right? What happens when the models are backdoors not just to computer systems, but to our minds?
Because once it's open weights, it doesn't like phone home. It's more about, like, if it recognizes a certain system, it could... Now, it could be a backdoor in the sense of, like, hey, if you're building a software, you know, something in software, all of a sudden it's a software agent. Oh, program this backdoor that only we know about.
Or it could be, like, subvert the mind to think that, like, XYZ opinion is the correct one.
There's this very good quote from Sam Altman who, you know, he can be a hype beast sometime, but one of the things he said, and I think I agree, is that superhuman persuasion will happen before superhuman intelligence. Yeah. And if that's the case, then these things before we get this AGI-ASI stuff, we can embed superhuman persuasion towards our ideal or whatever the ideal of the model maker is.
And again, today, I truly don't believe DeepSeek has done this. But it is a sign of what could happen.
Yeah, recommendation systems hack the dopamine induced reward circuit, but the brain is a lot more complicated. And what other sort of circuits, quote unquote, feedback loops in your brain can you hack slash subvert in ways like recommendation systems are purely just trying to do? you know, increased time and ads and et cetera.
But there's so many more goals that can be achieved through these complicated models.
Showing 1501–1520 of 1,814 · page 76 of 91
← Previous
Next →