Arvind Narayanan

speaker
255 appearances 3 recordings 3 series first heard Aug 2024 last heard 25 Jan

Arvind Narayanan’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.

Appearances

newest first · ▶ plays the moment
These models are already trained on essentially all of the data that companies can get their hands on. So while data is becoming a bottleneck, I think more compute still helps, but maybe not as much as it used to. And the reason for that is that perhaps ironically, more compute allows one to build smaller models with the same capability level.
And that's actually the trend we've been seeing over the last year or so. As you know, you know, the models today have gotten somewhat smaller and cheaper than when GPT-4 initially came out, but with the same capability level. So I think that's probably going to continue. Are we going to see a GPT-5 that's as big a leap over GPT-4 as GPT-4 was over GPT-3? I'm frankly skeptical.
Right. So there are a lot of sources that haven't been mined yet. But when we start to look at the volume of that data, how many tokens is that? I think the picture is a little bit different. 150 billion hours of video sounds really impressive.
But when you put that video through a speech recognizer and actually extract the text tokens out of it and deduplicate it and so forth, it's actually not that much. It's an order of magnitude smaller than than what some of the largest models today have already been trained with.
Now training on video itself, instead of text extracted from the video, that could lead to some new capabilities, but not in the same fundamental way that we've had before, where you have the emergence of new capabilities, right? Models being able to do things that people just weren't anticipating.
So like the kind of shock that the AI community had when I think back in the day, I think it was GPT-2,
was trained primarily on English text, and they had actually tried to filter out text in other languages to keep it clean, but a tiny amount of text from other languages had gotten into it, and it turned out that that was enough for the model to pick up a reasonable level of competence for conversing in various other languages.
So these are the kinds of emergent capabilities that really spooked people, that has led to both a lot of hype and a lot of fears about what bigger and bigger models are going to be able to do. But I think that has pretty much run out because we're training on all of the capabilities that humans have expressed, like translating between languages, and have already put out there in the form of text.
So if you make the data set a little bit more diverse with YouTube video, I don't think that's fundamentally going to change. Multimodal capabilities, yes, there's a lot of room there. But new, emergent text capabilities, I'm not sure. MARK BLYTH What about synthetic data?
Yeah, let's talk about synthetic data. So there's two ways to look at this, right? So one is the way in which synthetic data is being used today, which is not to increase the volume of training data, but it's actually to overcome limitations in the quality of the training data that we do have.
So for instance, if in a particular language, there's too little data, you can try to augment that, or you can try to have a model, you know, solve a bunch of mathematical equations, throw that into the training data. And so for the next training run, that's going to be part of the pre training. And so the model will get better at doing that.
And the other way to look at synthetic data is, okay, you take 1 trillion tokens, you train a model on it, and then you output 10 trillion tokens, so you get to the next bigger model, and then you use that to output 100 trillion tokens. I'll bet that that's just not going to happen. That's just a snake eating its own tail, and...
What we've learned in the last two years is that the quality of data matters a lot more than the quantity of data. So if you're using synthetic data to try to augment the quantity, I think it's just coming at the expense of quality. You're not learning new things from the data. You're only learning things that are already there.
Yeah, I think that's really spot on. I think one way in which people's intuitions have been kind of misguided by the rapid improvements in LLMs is that all of this has been in the paradigm of learning from data on the web that's already there. And once that runs out, you have to switch to new kinds of learning, analog of riding a bike. That's just kind of tacit knowledge.
It's not something that's been written down. So a lot of what happens in organizations is the cognitive equivalent of I think what happens in the physical skill of riding a bike.
And I think for models to learn a lot of these diverse kinds of tasks that they're not going to pick up from the web, you have to have the cycle of actually using the AI system in your organization and for it to learn from that back and forth experience instead of just passively ingesting.
It's got to be more than passive observation. You have to actually deploy AI to be able to get to certain types of learning. And I think that's going to be very slow. And I think a good analogy is self-driving cars, of which we had prototypes two or three decades ago.
But for these things to actually be deployed, you have to roll it out on slightly larger and larger scales while you collect data, while you make sure you get to the next nine of reliability, four nines of reliability to five nines of reliability. So it's that very slow rollout process. It's a very slow feedback loop.
And I think that's going to happen with a lot of AI deployment and organizations as well.
Yeah, thank you for asking that. That's not obvious at all. My view is that in a lot of cases, the adoption of these models is not bottlenecked by capability. If these models were actually deployed today to do all the tasks that they're capable of, it would truly be a striking economic transformation. The bottlenecks are things other than capability. And one of the big ones is cost.
Showing 141–160 of 255 · page 8 of 13 ← Previous Next →