Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Like not only just as thinking, it can also be a agentic model.
For example, say you wanted to generate the image of today's news.
So it's likely they'll go to fetch today's news online and then like process and diagnose them, then organize the layout and generate it.
For example, like ChemNet Omni.
Since they said it's Omni, I believe it's a single model.
Maybe it's something like... It's a language model with a diffusion head or something.
It's a language model...
do the thinking, do the agentic tool calling, and then it would use the diffusion head to generate the image in the end.
There were also approaches like Cosmos, where you have a separate language model and separate diffusion models.
And there are also like a purely language model, like you discretize the images and then
you generate an image as discrete tokens.
So there are different approaches.
Yeah, I'm not sure if they have that process.
I guess it's definitely possible in the Omni paradigm.
So if you think about like traditional multi-model language model, they would have a VIT encoder that can encode the image.
So if they have a deficient hash, they can generate the image and then put that back into the VIT encoder, encode that, and then do the iterative refinement of the result.
I cannot comment on that.
Yes, they're different.
I feel like the common part is the image part.
So it's quite surprising that a lot of the improvement came from the language side, the thinking, the tool calling.
Showing 561–580 of 722 · page 29 of 37
← Previous
Next →