Zico Colter

speaker
161 appearances 1 recordings 1 series first heard Sep 2024 last heard Sep 2024

Zico Colter’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
sort of spatio-temporal data, this is hugely important to our conception of intelligence, right? This is hugely important to the way that we interact with the world, the way that we sort of think about our own intelligence. And so I can't fathom that there is not a value to many, many more modalities of data, be it video, be it audio, be it
other time series and things like this that we sort of don't quite, that are not audio, but the sort of other sensory signals, stuff like this. There are massive amounts of data available. And I think we have not yet figured out how to properly leverage those due to either limitations of compute.
I mean, you have to process all that data and it does take, we don't have current models to do this very well, or just to do the limitations and sort of how we transfer and generalize across these modalities here. I think there has to be a use for it.
There are a few different sort of notions here. One is just the fact that we still seem to be in a world where you can increase model size and get better performance, even with the same data. So obviously, the real value of bigger models is they can suck up more data. They're able to ingest more and more data.
But it is also true that if you just take a fixed data set and run over it multiple times, if you use a bigger model, it will often work better. We have not really reached the plateau there.
The other thing, though, I don't think anyone would argue, or most people would not argue, that the current models in some sense extract the maximum information possible out of the data that is presented to them. And a very simple example of this is if you train a classifier, just to classify images of cats versus dogs on a bunch of images, you get a certain level of performance.
If you train a generative model on those exact same images, generate more synthetic data from that generative model, and then train on that more synthetic data, you don't do that much better, but you do a little bit better. And that's just wild. What that means is our current algorithms, we are not yet maximally extracting the information from data we have.
And there are way more deductions and inferences and other processes that we can apply to our current data to provide more value. And as models get bigger and better, they arguably can kind of do this themselves, either through synthetic data or through different mechanisms by which we train these models.
I don't really know, to be honest. I think this is a major open question right now in research. How do we extract the maximal information content from the data that we have? But again, as I said, I don't think we're close to even extracting all the data that's available.
When I look at this landscape, and we know that we aren't close to extracting the maximal information in the closure set of all the data that we have available... And we have not come close to processing all the data that's available to us. The idea that somehow this is a recipe for models plateauing in performance just doesn't jive to me with the reality of what we see.
We have not yet reached an equilibrium point where we have a good sense of sort of what the steady state of model size and for what application and how it's being used and is it being used as a general purpose system or for a very specific reason. This is all still being figured out right now. What I will say is that I use these models very regularly for my daily work.
I work almost exclusively with the largest models that are available to me because it just works better. And when I don't have a given task that I'm doing over and over, when I want to have that generality, I want to work with the larger models that are available.
The notion of sort of small language models and this kind of stuff, and again, I think this might be very much a possibility in the future. It kind of comes after we reach this point of generality, right? So once we've done something enough and we realize, okay, there is still a small task we want to do many, many times, maybe before we would have used...
custom-trained machine learning model for this. But the idea is that once you have a task, a rote task that you're repeating again and again enough times, and you know a small model can do it, it probably does become valuable to specialize a small model for that task only.
I think that actually has much more to do with our benchmarks and the way people typically are used to using these models than the models themselves. If you look at some of the hardest problems that models sort of face, we are still seeing gains to larger models of different techniques, things like this. Part of the problem here actually is sort of these models are a victim of their own success.
People have started to use these very regularly in their daily lives and they probably have a suite of questions that they ask these models. You know, when they first interact with a model, they'll probably ask it to write a biography of history of your school or a biography of yourself or something like this, but you have a suite of questions you sort of ask these models.
And on a lot of these pre-formatted questions that you know models already kind of do well on, the newer models don't do notably better, right? So, if I say, write a history of Carnegie Mellon University, Lama's 7 billion can do that just fine, right? Or 8 billion now can do that just fine, right? There's no need. I mean, maybe it'll be a little bit better with the largest closed source models, but
These aren't the kind of questions that are relevant. The domain I use models for most probably is coding and also doing things like transcribing lectures and stuff like this. On those tasks, I am absolutely not seeing plateauing gains. The latest models, they are notably better than the previous iteration and just make my life easier.
Let me move up to sort of higher and higher levels of abstraction when I give them instructions, when I interact with them, and when I work with them. So This perception has more to do with people's limited imagination of what they can do with these models and less to do with the models themselves. But that will evolve over time.
People will start figuring out you can use them for better and better things.
Showing 21–40 of 161 · page 2 of 9 ← Previous Next →