Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
But this task and form of training only works when it's verifiable. And from here, the thought is, okay, we can continue to scale this current training method by increasing the number of verifiable tasks. In math and coding, coding probably has a lot more to go. Math has a lot less to go in terms of what are verifiable things.
Can I create a solver that then I generate trajectories toward or reasoning traces towards and then prune the ones that don't work and keep the ones that do work? Well, those are going to be solved pretty quickly, but even if you've solved math, you have not... actually created intelligence, right?
And so this is where I think the like, aha moment of computer user robotics will come in because now you have a sandbox or a playground that is infinitely verifiable, right? Did you, you know, messing around on the internet, there are so many actions that you can do that are verifiable. It'll start off with like, log into a website, create an account, click a button here, blah, blah, blah.
But it'll then get to the point where it's, hey, go do a task on Tasker or whatever these other, all these various task websites. hey, go get hundreds of likes, right? And it's going to fail. It's going to spawn hundreds of accounts. It's going to fail on most of them. But this one got to a thousand. Great. Now you've reached the verifiable thing.
And you just keep iterating this loop over and over. And that's when... And same with robotics, right? That's where, you know, where you have an infinite playground of tasks like, hey, did I put the ball in the bucket? All the way to like, oh, did I like build a car, right? Like, you know, there's a whole...
trajectory to speed run or you know what models can do but at some point i truly think that like you know we'll spawn models and initially all the training will be in sandboxes but then at some point you know the language model pre-training is going to be dwarfed by what is this reinforcement learning you know you'll pre-train a multimodal model that can see that can read that can write you know blah blah blah whatever vision audio etc but then you'll have it play in a sandbox and
infinitely and figure out figure out math figure out code figure out navigating the web figure out operating a robot arm right and then it'll learn so much and the aha moment i think will be when this is available to then create something that's not good right like oh cool part of it was like figuring out how to use the web now all of a sudden it's figured out really well how to just get hundreds of thousands of followers that are real and real engagement on twitter because all of a sudden this is one of the things that are verifiable
And it's verifiable, right?
I think one of the things that people are ignoring is Google's Gemini flash thinking is both cheaper than R1 and better. And they released it in the beginning of December. And nobody's talking about it. No one cares.
I get the same question from earlier. Uh, the one about the, the human nature.
Oh, and it latched onto human, and then it went into organisms, and oh, wow. Yeah.
I think when you, you know, to Nathan's point, when you look at like the reasoning models, to me, even when I used R1 versus O1, there was like that sort of rough edges around the corner feeling, right? And flash thinking, you know, earlier, I didn't use this version, but the one from December, and it definitely had that rough edges around the corner feeling, right?
Where it's just not fleshed out in as many ways, right? Sure, they added math and coding capabilities via these verifiers in RL, but it feels like they lost something in certain areas. And O1 is worse performing than chat in many areas as well, to be clear. Not by a lot. Not by a lot though, right?
And it's like R1 definitely felt to me like it was worse than V3 in certain areas, like doing this RL expressed and learned a lot, but then it weakened in other areas. And so I think that's one of the big differences between these models is
and the and and and what oh one offers and then open ai has oh one pro and what they did with oh three which is like also very unique is that they stacked search on top of chain of thought right um and so chain of thought is one thing where it's able it's one chain it backtracks goes back and forth but how they served solved the arc agi challenge was not just the chain of thought
It was also sampling many times, i.e. running them in parallel and then selecting.
And then what?
Another form of search is just asking five different people and then taking the majority answers. Yes.
right there's a variety of like you know it could be complicated it could be simple we don't know what it is just that they are they are not just issuing one chain of thought in sequence they're launching many in parallel and in the arc hgi they launched a thousand in parallel for their uh the one that like really shocked everyone that beat the benchmark was they they would launch a thousand in parallel and then they would get the right answer like 80 of the time or 70 of the time 90 maybe even
Whereas if they just launched one, it was like 30%.
Showing 1541–1560 of 1,814 · page 78 of 91
← Previous
Next →