Ethan He

speaker
722 appearances 1 recordings 1 series first heard Jun 2026 last heard 1 Jun

Ethan He’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
So you say you have a 16 by 16 patch, then you match, you map that patch of pixels into this latent space.
Yes, yes.
Yeah, actually in VAEs, there are both convolution networks and transformers.
You can actually do both.
Yeah.
After this VAE, so what you've got is you've got latent space tokens and you've got the language tokens.
So now the training of the diffusion transformer, yearly generated models use diffusion transformers.
It's actually quite standard.
It's very similar to how you train language transformer models.
It's not that much difference.
It's just the tokens, the visual tokens in, visual tokens out.
The only difference is there's a denoising process.
So you train the model to unmask some of the noise.
So you add random noise to the visual tokens.
then you train the model to remove those noise to generate the clean tokens.
And in inference, the model can iteratively remove noise from 100% noise.
Yeah.
And then there's also...
After you train such model, such image model, the reason it's a foundation for video models is that image models are cheaper to train and they have much denser connection between language and text.
sorry, language and images.
Showing 121–140 of 722 · page 7 of 37 ← Previous Next →