Zico Colter

speaker
161 appearances 1 recordings 1 series first heard Sep 2024 last heard Sep 2024

Zico Colter’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
I think that it's been evolving so quickly between recent releases of open source models, continued progress in the closed source models. There was also this proliferation early on of a lot of open source models that none were better than the other, and they just involved a lot of training for companies, for lack of a better word, to just demonstrate that they could do it too.
It's not clear that's a valuable thing. Why would you want to train your own language model from scratch if there are very good open source ones now? Will that continue? Maybe, maybe not. I think there will be most likely consolidation, but I'm not quite sure how it will play out.
I do think that there are a lot of companies that are right now thinking about training their own models and things like this. And it's just sort of the default that, of course, you would do this, that this won't be an economically viable thing to do in the future. And so it won't happen anymore.
I'm not really sure what the rationale is for saying that we've plateaued in the compute sense. Most scaling laws that I've seen certainly suggest it can keep going. It's more expensive. You could argue that it's by far just scaling may not be the most efficient way. to achieve better results, and I actually think that's very likely true.
There are other better ways, you could argue, to achieve the same level of improvement than compute, but compute still does seem to be both a major factor and b still seems to improve things. It's more a calculus about also the monetary trade-offs of how much models will cost at inference time and how much they cost to train and all this kind of stuff.
These are much more, I would say, kind of becoming more practical concerns than a concern about the actual limits of scaling.
This question is, I think, super interesting. And those are not actually mutually exclusive, to be clear. One thing I'll also say is that the term AGI is thrown around a whole lot. I define AGI as a system that acts functionally equivalent to a close collaborator of yours over the course of about a year-long project.
So this is something that, you know, you would value as much as a close collaborator, you know, a student of mine or a colleague of mine over working on a project for a year. Let's think about me. AGI would be a system that could automate everything that I do for the most part over a year. That's a pretty high bar.
I am massively uncertain as to when this will happen, but a massive shift that I've undergone is I think this will probably happen in my lifetime. I think the answer to AGI has always been in academia, not in my lifetime. And the timeframe I give this right now is I think this is between four and 50 years or something like this, right? Which really captures my massive uncertainty.
But I have a hard time dismissing it also, given the rate of progress and the things that I sort of see evolving here. We have to take that possibility very seriously.
I actually also agree with you that we will adapt to it. I don't want to downplay the extent of transformation that might be necessary here. But I also think the companies that say survive and thrive and become dominant in this new world, the ones that succeed best will not be the ones that fire all their workers to have an AI that does the exact same thing as their old workers.
They'll be the ones that understand, okay, what's changing and what are the things that people can best do now in terms of steering these systems, in terms of sort of providing the overall guidance and framework about where we want to go with with all this intelligence.
The companies that survive, I think, will be the ones that best leverage their workforce to make the best use of this new technology.
This is actually a very nuanced question. Do we have AI products that are able to be maximally used by workforces? The answer to this right now is no. Clearly, there is a gap between what people could use these things for and what they're using them for right now.
I mean, I find this kind of interesting in a way because enterprises are all very happy to put their data in the cloud. They all use cloud services to store their data. But then, oh, train on this there? No, no, no. Can't do that. I think a lot of it comes, honestly, from kind of a misunderstanding about how this process works.
Also, frankly speaking, I think it has to do with the fact that if you think about the model of just taking all your internal data and dumping it into a large language model, this is not tenable. You can't do this for a number of reasons. The most obvious one being the data has access rights, right? Not everyone gets access to all the data.
And the default mode of language models is that if you train on some data, you can probably get it back out of the system if you want to enough. And so this doesn't work with the sort of the access controls people have in traditional data. I think these are kind of the concerns. Now, to be clear, there are very easy ways around this, right?
So this is probably why RAG-based systems are so common here and probably will remain, even with the advent of fine-tuning availability, they're going to remain a useful paradigm. RAG is, for those that maybe haven't heard the term, it's retrieval augmented generation. It basically means that you just
You go out and fetch the data you can access, that you have access rights to, that is relevant to your question. You inject it all into the context of the model, and then you answer the question based upon this data here too. These RAG-based techniques are going to remain popular precisely because they respect normal data access procedures.
I sort of feel like a lot of this hesitancy actually comes from a fundamental misunderstanding of how these models are working. People think that if you have ChatGPT answer a question about any of your data, that data is somehow being trained upon and merged into the model, whether it's an API call or whether it's a RAG-based call or anything else. And it's just not true.
Showing 41–60 of 161 · page 3 of 9 ← Previous Next →