Ethan He

speaker
722 appearances 1 recordings 1 series first heard Jun 2026 last heard 1 Jun

Ethan He’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Yes.
Yeah, for the generation model training, there's also really like a small percentage of
unlabeled data.
So the model is instructed to generate a video without any text instruction.
That can also help the model generalize.
So after this stage of generating the synthetic pair, one important common step is to train a compressor or a tokenizer of the image or videos.
So because if your trend, if you can technically, theoretically trend image or video models on pure pixels, but the problem is that it's a lot of tokens.
So like one image, like it's a thousand by a thousand, it's like one million tokens, one million pixels.
It's impossible to trend transform around that.
So you need to train a tokenizer which can go from image to latent space and latent space back to image.
That's why we named the podcast.
So like a million is impossible?
In generative models, the vocab is continuous.
It's a continuous space.
You can think about, like, you map an image to a vector.
It's a fixed-length vector.
It's like 16 or 48, something like that.
And then you map that vector back to the image space.
And the mapping has...
The mapping is patch-based.
Showing 101–120 of 722 · page 6 of 37 ← Previous Next →