Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
So the model just needs to learn one of the distribution, not the full distribution.
Because in training, the model is asked to reconstruct the ground truth image from the internet, which is extremely hard.
And when you're training GAN, it's a one-step process.
It's just, hey, you generate image.
Does this image look as real as the image from the internet, which is a much simpler task.
And, yeah, combining a lot of these approaches together, people typically do that, like consistency model and distribution matching.
And again, we can get these few-step models.
Okay, then there's one step I wanted to add, which is audio.
Yes.
And video.
Yeah, so...
Grok Imagine 0.9, I believe, is the first audio-video joint model deployed at a large scale.
And that was your first model?
Yes, that was Grok Imagine's first model.
It's audio-video joint generation.
I think the hard part is the modality alignment.
Because before this trend model, we have text-to-video alignment.
We have this correspondence between text and video.
Typically, most of the VLMs, they understand images and videos very well, and they don't understand audio mostly.
And if you look at the audio generation on the OEM side, you can talk to them perfectly fine, but if you ask them to sing a song or something, it typically is not very good.
Showing 301–320 of 722 · page 16 of 37
← Previous
Next →