Elan Barenholtz

speaker
2,007 appearances 4 recordings 1 series first heard Feb 2025 last heard Jun 2025

Elan Barenholtz’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Uh we're gonna call tokens a word.
Token is a more technical term about how you chop up uh and and encode uh the information in in the
So guess the next word um based on that sequence.
So uh and then what what you do is uh in in these models is you uh you train them to to guess simply that.
Are you serious?
Given sequence, it can be a sentence, it can be a paragraph, it can be frankly an entire uh book, uh depending on uh how big how big your model is, how much it can handle.
And then guess just
The very next word.
And what we've discovered, and I say I I really want to use that word uh i in particular uh because it is was s by no means a given that this could ever this that this would work, what we've discovered is if you train a model to do that, uh to just simply guess the next word, then take that word, tag it on to the next uh tag it onto the sequence and feed it back in.
This is sufficient to generate human level language.
Now, the reason I I uh believe that this demonstrates something not about uh our our engineering or even about them the models themselves, 'cause there's different ways you might build a model that can do this, is because this very simple trick, this simple recipe of simply guessing the next word is
turns out to be sufficient to generate language uh at at human levels to the point where uh there really are no benchmarks, no standard benchmarks that these models aren't able to do.
And so what that suggests to me is uh just by learning the the predictive structure of language, you're able to completely solve language.
That means that that is likely to be the actual fundamental principle that's built into language in order to generate it.
If we had to come up with a very complex scheme, for example, you know, syntax syntax trees, um uh complex grammar, long-range dependencies that we had to take it into account.
And through enough compute, uh we were able to kind of master that, then I might argue, well, you know, what we're doing is uh possibly figuring out a roundabout way to capture all this complexity.
But it's the simplistic.
itself, uh that simply being able to predict the next token, uh, the next word, is sufficient to do all of this long-range thinking, uh, to be able to take an extremely long sequence and then produce an extremely long sequence on the basis of that, that suggests to me that
we discovered a principle that's actually already latent in language, that we just had to throw enough firepower at it, but with an extremely simple algorithmic trick, uh, and then language revealed its its secrets.
Um so to me this really suggests that uh there is of course, you know, the the there there's still a lot of science that needs to be done and this this kind of thing, uh I mean this kind of work that I'm doing in my lab.
Showing 41–60 of 2,007 · page 3 of 101 ← Previous Next →