Ethan He

speaker
722 appearances 1 recordings 1 series first heard Jun 2026 last heard 1 Jun

Ethan He’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Like not only just as thinking, it can also be a agentic model.
For example, say you wanted to generate the image of today's news.
So it's likely they'll go to fetch today's news online and then like process and diagnose them, then organize the layout and generate it.
For example, like ChemNet Omni.
Since they said it's Omni, I believe it's a single model.
Maybe it's something like... It's a language model with a diffusion head or something.
It's a language model...
do the thinking, do the agentic tool calling, and then it would use the diffusion head to generate the image in the end.
There were also approaches like Cosmos, where you have a separate language model and separate diffusion models.
And there are also like a purely language model, like you discretize the images and then
you generate an image as discrete tokens.
So there are different approaches.
Yeah, I'm not sure if they have that process.
I guess it's definitely possible in the Omni paradigm.
So if you think about like traditional multi-model language model, they would have a VIT encoder that can encode the image.
So if they have a deficient hash, they can generate the image and then put that back into the VIT encoder, encode that, and then do the iterative refinement of the result.
I cannot comment on that.
Yes, they're different.
I feel like the common part is the image part.
So it's quite surprising that a lot of the improvement came from the language side, the thinking, the tool calling.
Showing 561–580 of 722 · page 29 of 37 ← Previous Next →