Appu (Apu) Shaji
speaker
101 appearances
1 recordings
1 series
first heard Jul 2026
last heard 14 Jul
Appu (Apu) Shaji’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
Working Smarter · Building AI that can search inside videos (and photos and audio too) · 14 Jul 2026
podcast
But now we are able to go and free up a lot of their time to actually do the work that matters.
rather than this groundwork involving folder arrangement and sort of uh adding metadata attacks, et cetera.
Minimizing that time is the main sort of the motivation of the work that we do so that people can tell better stories, faster stories, um, and clearly make uh cooler narratives uh moving forward much easier.
So that uh
maybe two or three aspects to it.
One is what does this bunch of pixel constitute to be a Berlin?
Because these models have seen a lot of previous Berlin f uh photographs or visuals uh previously, inside the network it has a idea about what Berlin looks like.
That means this new certain neurons in these uh transformer networks will fire for these pixels which resemble these facets of Berlin or in idiosyncrasies of Berlin.
Similarly, nighttime also has a certain pattern to it.
Again, these because we learn these networks with large amounts of data, there is again a certain activation pattern for this.
Together, Berlin in nighttime, there is even a specialized activation uh patterns that happens in this particular network.
And this is actually quite remarkable to think about this because we are not explicitly telling, hey, this is Berlin or this is nighttime, because it has seen a lot of data.
It's able to do that.
It's a cr a very important and profound uh problem when it comes to multimodality where you are not only looking at visuals, you have to correlate with what someone is speaking, for example.
In Jurassic Park, there is a scene where it was explained how the park was being made because an insect got uh sort of sucked in sapling and et cetera, and then they could extract out the DNA, et cetera.
So, prime, the main information there was in the audio stream, plus the visuals connected to the sapling, et cetera.
Mm-hmm.
So, this is what we call as multimodality.
The principle remains the same.
It has seen a lot of this multimodal context previously.
Showing 61–80 of 101 · page 4 of 6
← Previous
Next →