Ethan He

speaker
722 appearances 1 recordings 1 series first heard Jun 2026 last heard 1 Jun

Ethan He’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
So the model just needs to learn one of the distribution, not the full distribution.
Because in training, the model is asked to reconstruct the ground truth image from the internet, which is extremely hard.
And when you're training GAN, it's a one-step process.
It's just, hey, you generate image.
Does this image look as real as the image from the internet, which is a much simpler task.
And, yeah, combining a lot of these approaches together, people typically do that, like consistency model and distribution matching.
And again, we can get these few-step models.
Okay, then there's one step I wanted to add, which is audio.
Yes.
And video.
Yeah, so...
Grok Imagine 0.9, I believe, is the first audio-video joint model deployed at a large scale.
And that was your first model?
Yes, that was Grok Imagine's first model.
It's audio-video joint generation.
I think the hard part is the modality alignment.
Because before this trend model, we have text-to-video alignment.
We have this correspondence between text and video.
Typically, most of the VLMs, they understand images and videos very well, and they don't understand audio mostly.
And if you look at the audio generation on the OEM side, you can talk to them perfectly fine, but if you ask them to sing a song or something, it typically is not very good.
Showing 301–320 of 722 · page 16 of 37 ← Previous Next →