Yann LeCun

speaker
384 appearances 5 recordings 4 series first heard Mar 2024 last heard 20 Jun

Yann LeCun’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
2 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 4 in all, peaking in Jan 2026 with 2.

Appearances

newest first · ▶ plays the moment
So there was a surprise many years ago with what's called decoder-only LLMs. So, you know, systems of this type that are just trying to produce words from the previous one. And the fact that when you scale them up, they tend to really kind of understand more about language. When you train them on lots of data, you make them really big. That was kind of a surprise.
And that surprise occurred quite a while back, like, you know, with work from Google Meta, OpenAI, et cetera, going back to the GPT kind of work, general pre-trained transformers.
Yeah, I mean, there were work from various places, but if you want to kind of place it in the GPT timeline, that would be around GPT-2, yeah.
Well, we're fooled by their fluency, right? We just assume that if a system is fluent in manipulating language, then it has all the characteristics of human intelligence. But that impression is false. We're really fooled by it.
Alan Turing would decide that the Turing test is a really bad test. Okay. This is what the AI community has decided many years ago, that the Turing test was a really bad test of intelligence.
Hans Moravec would say the Moravec paradox still applies. Okay. Okay. Okay, we can pass.
No, of course, everybody would be impressed. But, you know, it's not a question of being impressed or not. It's the question of knowing what the limit of those systems can do. Again, they are impressive. They can do a lot of useful things. There's a whole industry that is being built around them. They're going to make progress.
But there is a lot of things they cannot do and we have to realize what they cannot do and then figure out how we get there. I'm seeing this from basically 10 years of research on the idea of self-supervised learning.
Actually, that's going back more than 10 years, but the idea of self-supervised learning, so basically capturing the internal structure of a set of inputs without training the system for any particular task, learning representations. You know, the conference I co-founded 14 years ago is called International Conference on Learning Representations.
That's the entire issue that deep learning is dealing with, right? And it's been my obsession for almost 40 years now. So learning representation is really the thing. For the longest time, you could only do this with supervised learning.
And then we started working on what we used to call unsupervised learning and sort of revived the idea of unsupervised learning in the early 2000s with Yoshua Bengio and Jeff Hinton. Then discovered that supervised learning actually works pretty well if you can collect enough data. And so the whole idea of unsupervised self-supervised learning kind of took a backseat for a bit.
And then I kind of tried to revive it in a big way, starting in 2014, basically when we started FAIR. and really pushing for finding new methods to do self-supervised learning, both for text and for images and for video and audio. And some of that work has been incredibly successful.
I mean, the reason why we have a multilingual translation system, you know, things to do content moderation on Meta, for example, on Facebook, that are multilingual, that understand whether a piece of text is hate speech or not or something. is due to that progress using self-supervised learning for NLP, combining this with transformer architectures and blah, blah, blah.
But that's the big success of self-supervised learning. We had similar success in speech recognition, a system called Wave2Vec, which is also a joint embedding architecture, by the way, trained with contrastive learning. And that system also can produce speech recognition systems that are multilingual,
with mostly unlabeled data and only need a few minutes of labeled data to actually do speech recognition. That's amazing. We have systems now based on those combination of ideas that can do real-time translation of hundreds of languages into each other.
That's right. We don't go through text. It goes directly from speech-to-speech using an internal representation of kind of speech units that are discrete. But it's called textless NLP. We used to call it this way. But yeah, so that, I mean, incredible success there. And then, you know, for 10 years, we tried to apply this idea
to learning representations of images by training a system to predict videos, learning intuitive physics by training a system to predict what's going to happen in a video, and tried and tried and failed and failed with generative models, with models that predict pixels. We could not get them to learn good representations of images. We could not get them to learn good representations of videos.
And we tried many times. We published lots of papers on it. They kind of sort of worked, but not really great. It started working. We abandoned this idea of predicting every pixel and basically just doing digital embedding and predicting in representation space. That works.
So there's ample evidence that we're not going to be able to learn good representations of the real world using generative model. So I'm telling people, everybody's talking about generative AI. If you're really interested in human-level AI, abandon the idea of generative AI.
Right. Well, there's a lot of situations that might be difficult for a purely language-based system to know. Like, okay, you can probably learn from reading texts, the entirety of the publicly available texts in the world, that I cannot get from New York to Paris by snapping my fingers. That's not going to work, right? Yes.
Showing 141–160 of 384 · page 8 of 20 ← Previous Next →