Gaurav Misra
speaker
155 appearances
1 recordings
1 series
first heard Jan 2025
last heard Jan 2025
Gaurav Misra’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Getting into it, I think the first question behind the question that comes to mind is what exactly did we actually achieve with this AI revolution? What is actually the difference? AI existed before and it exists today. Obviously, there's something magical about what is there today. I think when you get into it, you realize that it's really about the ability to train larger and larger models.
Yes, that's actually a combination of we have better hardware to do it. We have better ML architectures, like there's transformers, there's diffusion models, there's all these new types of architectural unlocks that we've created. And then there's other techniques that we've created too, which just allow us to train larger and larger models.
And turns out the larger and larger you make these models, the more problems they can solve, the better they can be at solving like text generation or like towards AGI or video generation or media generation in general. I think when you realize that, what actually you get to is that what really matters is the data at the end of the day.
A lot of companies are like scraping the internet and the internet is also limited in some ways. There's only so much information on the internet even, and that's growing every day. But I think at the end of the day, beyond that, we're going to have to find what are those sustainable sources of data that can continue to grow bigger and bigger models.
And I think that's going to be the fundamental question behind who actually ends up winning in a lot of these different areas that AI is excelling in today. I think for us, being on the video generation, video editing side, it comes down to like video data, which is actually much heavier, much rarer to find, not as common as text or even audio.
and potentially much more expensive to train on as well, much more limited in terms of being created in the world. And so that tends to be like a big challenge.
One of the big things that we're thinking about is how do we actually create a flywheel where we can ingest data on a continuous basis and a growing basis, and that data can actually create bigger and bigger models for us and keep us at the forefront. I also want to call out here, like there's a pretty fundamental difference between different types of AI companies that are out there.
I think if you look at a lot of the text generation companies, they're not solving text generation. Like we don't call it text generation. They're actually kind of solving a totally different problem, which is intelligence. Intelligence is an unsolved problem. No one's figured that out yet. And yes, we're achieving some levels of intelligence in these models. And there's a long way to go.
It may not end at human intelligence. There's people in the world who are really smart. There's people in the world who are not so smart. They both exist. And clearly, so there's a range of intelligence as possible. There's not one value for like you're intelligent or not. So yeah, is there a chance that there's the ability to go smarter than the smartest human? It's possible.
But that's a frontier that we've never reached. And so it's kind of solving this unsolved problem. But I think if you think about audio generation or video generation or music generation or these types of things, right, it's I think a little bit less of solving an unbounded intelligence problem and a little bit more of solving actually rendering a solved problem.
And video, for example, like CGI exists. We can make fake things. We can make fake humans. We can make fake sceneries and dragons. And so this is a solved problem. We know that there are solutions to these. And with AI, we're actually just making it easier to solve these problems.
Not just a little bit, but like a hundred times easier, which in the end, that means more accessible, larger market, more people can use these types of technologies.
So I think that's one of the fundamental differences there is if you look at business models for like artificial intelligence companies that are really working on AGI, then you kind of have to think about this unbounded problem of like, okay, we put in a bunch of capital into it.
We create a model only for that model to be beat by the next model and that model becoming essentially useless and obsolete. And then there's the next model after that. And how long does this go on for? Actually, we don't know. It may go on forever. There may be like no end to this intelligence race. Whereas if you look at the media generation companies, it actually is creating an asset.
And there might be very soon a point where, oh, wow, it's just really good. It's just perfect or close to perfect. And we've kind of solved it. And then it's an asset. And then after that, it's just software company. And the asset's really expensive to create. But once it exists, it just generates value. And it doesn't lose value that easily.
So what is going to make those models better and better? I think it's going to be like fine tuning with more data, fine tuning for specific use cases, different types of things you want to generate, different types of visuals, whatever it might be. Use cases like, oh, it's going to be used in ads or movies or social media or something else.
But there may be a point where it's like, wow, yeah, this is pretty good. It's realistic. I think that's a pretty important thing we're thinking about right now. How do we bootstrap that data flywheel to be able to reach that level?
I mean, at the rate at which video models are going, I mean, you probably remember seeing like the Will Smith spaghetti thing. Everyone's seen this meme, right? And it went from like really horrible to like, wow, this is actually good. And I think really, really good is probably around a year, year and a half away.
I only say this because if you compare, for example, like text models to like video models, text models are already like in the 400 billion parameter range. People understand better how to scale LLM technology today just because more money has been put into it, more time has been put into it. Like diffusion models, still in the tens of billions. It's still early, not even close to the text models.
So as that grows, there's just no doubt it's going to get better and better. And like the experts kind of know that this is all possible. It's just that very few companies in the world have the funding and the expertise to actually go after this. So like it just takes some time. Like it's not like some unsolved problem. People know what needs to be done. It's just we're all getting there.
Showing 1–20 of 155 · page 1 of 8
Next →