Gaurav Misra

speaker
155 appearances 1 recordings 1 series first heard Jan 2025 last heard Jan 2025

Gaurav Misra’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
We're all moving towards it. And we'll see those models getting better and better, especially on the video side. I could easily see within a year and a half or so something coming pretty close to like indistinguishable, essentially from like a real recording, maybe even sooner. That's not like the worst case.
It has been a pretty interesting journey and like we've been through some interesting twists and turns through it. But I think if you like connect the dots end to end, it's interesting. When we started the company, the first app that we made was captions. We launched it. And why did we make it? The goal was to get content creators to create content on a video creation platform of some sort.
Not easy. I was at Snap before this and Snap had tried this many times. They launched apps and I mean, video is kind of a commodity. Video editors are commodities. A lot of these companies are actually foreign. And that's because we're just trying to minimize costs at this point and really difficult to compete in.
Our thought was the way we're going to crack this is we're going to use AI to help create video somehow. That's going to be our differentiator. That's why people are going to come to us. And so we saw that there was a need around speech to text. It was a technology, by the way, at that point, that was pretty good.
In text circles, people were like, of course, speech to text, we understand that it's pretty good at this point. But I think the average person actually didn't understand how good the tech had gotten and how accurate it was with names and like obscure terminology and all kinds of stuff. So when we built the first product where it was just like, hey, it's just literally put text on the videos.
And by the way, this was built in like two days on a weekend, really just bandaid together. And we put it on the app store, went to sleep. The next morning it was top of the app store. There's no explanation. We didn't do anything to make that happen. Somebody saw it. They posted it on something. It blew up and then woke up and I text Dwight.
I'm like, hey, I think there's like 600 videos per minute being created on the app, by the way. And so that was kind of like an instant success. But even in that two days of work, we had already instrumented the app in such a way that we would be able to continue training better and better models so that we can deliver better value to the user.
So the idea was like, the app is an AI app where people come in, they use the app, we use the data to make the model better and deliver even better experiences the next time the person comes back. That was done from day one, literally. That was the original plan. Now, post the launch of the app, we've added so many more features over time, expanded the offering so much more.
And we cover now the entire space of everything from like script writing to recording to video editing, distribution as well. and how AI can like transform each of these different areas, because there's applications in all of them. And there's data that can be collected across all of those that can improve those models.
And that's what makes our offering really unique, because all the other companies are not really thinking about the data collection side and just generating outputs. And that's why they have to kind of scrape the internet to make their models better. And for us, really, it's more about growing a user base so that the data can actually power better and better models.
And a lot of that comes through like video. So video being funneled directly into video generation models. That gives a significant advantage. That's potentially a possible way in which a future sort of business model could be set up. It actually is kind of familiar, by the way.
It seems to me similar to the Facebook or Google business model where you have a mass consumer free product, basically, and the data is used to power essentially like a B2B paid product.
It's interesting to think about because for the models that we train, they're diffusion models. So they actually work by starting from noise. It starts from literal noise, like static you see on TV. At every step, based on text that's provided, it looks at the noise and it tries to like predict a layer of clarity in that noise. It says man wearing blue shirt.
So it starts to like draw a little bit of man wearing blue shirt out of noise. And then every pass it's taking through it, it's discovering a little bit more of the man wearing blue shirt. So that's the text conditioning that's helping it decide how to reach the destination of what man wearing blue shirt looks like.
So that's how the diffusion models work, which is slightly different from like how a next token prediction model like GPT works, which is kind of just as you might think about it, just predicting the next word based on all the previous words that have been spoken, which are considered the context. So these models are different. We are still earlier on in the diffusion model training path.
We're still in that 10 billion, 20 billion, 30 billion. Meta's movie gen was, I believe, 30 billion parameters. People haven't really scaled this up. We actually don't know how big OpenAI Sora is. They didn't, I think, release that information. But a lot of the work is going to go into scaling up these things. Video obviously is really heavy. That's what makes it different from text.
Consumes a ton of space, a ton of processing. For us, even if we were to download, just download all of our training videos online, it will cost us a million dollars to download the training videos. That's a whole different regime than like text. It brings different types of challenges to training these models, basically.
You never know. But honestly, I think what will save us on the video model side is actually the fact that it is an easier problem than the text problem. The text problem is intelligence, as we're talking about. And the video problem is more rendering. We already know how much rendering costs. We already know, yeah, it's GPU intensive.
If you were to like literally CGI render a scene out, like, yeah, it will spend some time on the GPU. There's no doubt. Can we be more efficient than that? It's possible. It may not be the most efficient today. Maybe there's better ways of doing it. Maybe AI will be cheaper and faster than regular rendering. And I think if that's the case, then that's a good thing.
But I think we know that it shouldn't be worse than that. We should be able to solve it with fewer resources than that, potentially, or at least the same. We generally understand where it's going to fall. It's still early. Just like on the training side, we're still scaling up these models and it's still, oh, it's 10 billion parameters, 20 billion parameters, whatever.
Showing 21–40 of 155 · page 2 of 8 ← Previous Next →