Ethan He

speaker
722 appearances 1 recordings 1 series first heard Jun 2026 last heard 1 Jun

Ethan He’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Speaking of the training, actually training the models that GPU costs, if you look up like the open source model, how big these video models are, I think like LTX has 19B parameters.
That's a dense model.
And people are also exploring MOEs.
So it might be like a 20B active and like 100B total.
So that's similar size as medium-sized LLM models.
And if you look at number of tokens,
We disclosed that in Cosmos.
It's also like tens of trillions of tokens on the virtual tokens.
So putting this together, the cost of training these video models, it's actually comparable with OLMs, not to mention the infra is slightly different from OLMs.
So it might be less efficient to train these models.
Yeah, so the inference side is a completely different story.
I think for the training side, it might be a little bit hard to reduce that cost.
And for the inference side, the biggest gain is from the dissolution of these models.
It's called step dissolution, slightly different from knowledge dissolution in all of them.
So you typically for flow matching models, you need like a hundred steps or something.
Like a distortion model even need even more like a thousand steps to generate a good image or video.
A step distillation is try to learn to generate fewer steps from the model itself.
It's kind of like...
Now, you use the full model to generate in 100 steps, and then you take a model that only generates 10 steps, and let that model learn from the perfect one.
Why does this work?
Showing 261–280 of 722 · page 14 of 37 ← Previous Next →