Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
So there's just like a lot of every different type of training and serving has these considerations you need to scale.
Like we talked about pre-training, we talked about RL and then inference time scaling is like, how do you serve a model that's thinking for an hour to a hundred million users?
I'm like,
I don't really know about that, but I know that's a hard problem.
And in order to give people this intelligence, there's all the systems problems and we need more compute and you need more stable compute to do it.
Some Reddit data is very coveted and excellent for training.
Yeah, I'm trying to learn so much about AI.
I was learning about pre-training parallelism.
I'm like, I lost something, and I don't know what it was.
A few things that could be helpful for people.
A lot of people think of synthetic data as being bad for training the models.
You mentioned the DeepSea got an OCR, which is optical character recognition paper.
A lot of labs did.
AI2 had one.
had multiple.
And the reason that each of these labs has these is because there's vast amounts of PDFs and other digital documents on the web that are in formats that aren't encoded with text easily.
So you use these almost CR, these were deep seek OCR.
And we called ours almost CR to extract what can be trillions of tokens of candidate data for pre-training and,
And pre-training data set sizes on the order of trillions is measured in trillions of tokens.
Smaller models from researchers can be something like 5 to 10 trillion.
Showing 301–320 of 1,814 · page 16 of 91
← Previous
Next →