Nathan Lambert
speaker
1,814 appearances
3 recordings
2 series
first heard Feb 2025
last heard 1 Feb
Nathan Lambert’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 2 in all, peaking in Feb 2026 with 1.
Appearances
The other thing to be complete is that some people are trying to train on only licensed data, where Common Crawl is a scrape of the whole internet.
So if I host multiple websites, I
happy to have them train language models, but I'm not explicitly licensing what governs it.
And therefore, the Common Crawl is largely unlicensed, which means that your consent really hasn't been provided for how to use the data.
There's another idea where you can train language models only on data that has been
licensed explicitly so that the kind of governing contract is provided and i'm not sure if apparatus is the copyright thing or the license thing i know that the reason that they did it was for an eu compliance thing where they wanted to make sure that their model um fit one of those checks and also on that note also for example there's also the distinction between um the licensing so some people like you said they um
I think on the data thing, this is one of the things where like this happened in 2025 and we totally forget it, is Anthropic lost in court and was owed $1.5 billion to authors.
Anthropic, I think, bought thousands of books and scanned them and was cleared legally for that because they bought the books and that is kind of going through the system.
And then the other side, they also torrented some books.
And I think this torrenting
was the path where the court said that they were then culpable to pay this billions of dollars to authors, which is just like such a mind-boggling lawsuit that kind of just came and went.
Like, that is so much money from the VC ecosystem.
largest problems in infrastructure and systems, but from an AI point of view, it's kind of inevitable.
Yes, and I think that a lot of open source contributors are legitimately burning out.
If you have a popular open source repo, somebody's like, oh, I want to do open source AI.
It's good for my career.
And they just vibe code something and they throw it into the... You might get more of this than I do.
So I have a case study here.
So when I write, I think a lot of what I'm trying to do is take what you think as a researcher, which is very raw, which a researcher is trying to encapsulate an idea at the frontier of their understanding.
And they're trying to put what is a feeling into words.
Showing 381–400 of 1,814 · page 20 of 91
← Previous
Next →