Zico Colter

speaker
161 appearances 1 recordings 1 series first heard Sep 2024 last heard Sep 2024

Zico Colter’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
The real negative outcome is that people are not going to believe anything that they see anymore. Arguably, we are already well along this way where people basically don't believe anything that they read or that they see or anything else. It doesn't already conform to their current beliefs. It didn't even need AI to get there, but AI is absolutely an accelerant for this process.
It is a relatively new phenomenon that we have sort of a record of objective fact in the world. I mean, things like video didn't exist more than 100 years ago. Humans evolved at a time during an environment where all we could do was trust our close associates. That's how we believed things.
Great. Thanks. Wonderful to be here.
Sure. Absolutely. So I seem to have be collecting jobs here. I have a number of different roles. I'm first and foremost, a professor and the head of the machine learning department. at Carnegie Mellon. I've been here for about 12 years. And here, the machine learning department is really kind of unique because it's a whole department just for machine learning.
And I've been heading it up actually, as of quite recently, and get to immerse myself in the business and the thought of machine learning all day, every day. Also, I am recently on the board of OpenAI, which I joined at this point a couple of weeks ago, and it's been extremely exciting as well.
Right. So let's talk about AI as LLMs, but with, of course, the context that AI is a much, much broader topic than this. LLMs are amazing. The way they work at the most basic level, you take a lot of data from the internet, you train a model. And I know that's a very sort of colloquial term that we use here. But basically, what you do is you build a great big
set of kind of mathematical equations that will learn to predict the words in the sequence that is given to them. If you see the quick brown fox as your starting phrase of a sentence, it will predict the word jumped. We train a big model on predicting words on the internet.
And then when it comes time to actually speak with an AI system, all we do is we use that model to predict what's the next word in a response. This is, to put it bluntly, a little bit absurd that this works. So there's sort of two philosophies of thought here. People often use this sort of mechanism of how these models work as a way to dismiss them.
Oftentimes, I know people say, oh, well, AI is, it's just predicting words. That's all it's doing. Therefore, it can't be intelligent. It can't be. And I think that's just demonstrably wrong. What I think is amazing, though, is the scientific fact
that when you build a model like this, when you build a model that predicts words, and then just turn this model loose, have it predict words one after the other, and then chain them all together, what comes out of that process is intelligent. And I think it's demonstrably intelligent, right? I really believe these systems are intelligent, definitely. And I would say that this fact is
You can train word predictors and they produce intelligent, coherent, long form responses. This is one of the most notable, if not the most notable scientific discovery of the past 10, 20 years, maybe much longer than that, right? Maybe it was much deeper than that, in fact. And so this is not oftentimes given its due as a scientific discovery because it is a scientific discovery.
two kind of answers to this question, which are diametrically opposed, as with many questions, right? Because you're exactly right. The thought is, because these models are built to basically predict text on the internet, if you run out of text, that would imply that they're kind of plateauing. I don't think this is actually true for several reasons, which I can get into.
But just from a raw standpoint of training these models, I mean, there's two ways in which this is sort of maybe true, maybe false. It is true that a lot of the easily available data, sort of the highest quality data that's out there on the internet has been consumed by these models. We have used this data. There is not another Wikipedia and things like this, right?
There's only so much really high quality, good text that's available out there. On the flip side, and this is the point I often make, first of all, we're only talking about text there. We're only talking about publicly available text.
If you start talking about internally available text, stuff like this, from a very straightforward standpoint, we have not gotten close to using all the data that's available. Public models that train on the order of 30 terabytes of data or something like this, right? So 30 terabytes of text data. This sounds like a lot, But this is a tiny, tiny amount of data.
And there is so much more data that's available that we are not using right now to build these models. And of course, I'm thinking about things like multimodal data, stuff like this video data, audio data, all these things, we have massive amounts available. I mean, just just a few tens of terabytes is not the amount of data these large companies that index the internet are storing.
There is so much more data than this, and we have not really come close to tapping that whole reserve. Now, whether or not we can use that data well, right, because text data in some sense is the most distilled form of a lot of this, and a lot of this is not textual data, that remains to be seen. But we are nowhere close to hitting the limits of available data in these models, right?
Arguably, we're unable to process it, because we don't have enough compute and things like this. But we're nowhere close to data limits in other senses. MARK MANDELBACHER- What are the challenges of using these new forms of multimodal data well? PAUL BAKAUSKI- I think the biggest challenge is simply compute.
If you have something like video data, just think about the size of a video file versus a text file. So if we transcribed this podcast, it would be a few kilobytes. If you take the dump of video from it, it'll be on the order of, I don't even know, I do. It would be about six and a half gigabytes. Gigabytes, exactly, right?
So tens of thousands of magnitudes of difference, orders of magnitude of difference, right? Now, arguably, depending on people's opinion, maybe the entirety of the actual valuable information is not in the audio of my voice and the video. You could argue that there's not as much usable content there. When we think about what kind of data humans use, I would argue that visual data
Showing 1–20 of 161 · page 1 of 9 Next →