Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
You generate a video, that's it.
That's just one time done.
And some creators would try to use the last frame as the first frame for the second video.
Sometimes it works, but if you do it a few times, it's a call to a degree.
It doesn't have that context for the full video.
Exactly.
Yeah, yeah, yeah.
And for example, like a view, I remember view three has like a one second context of the last video.
It is slightly better than using the last frame, but it has the same problem, similar problems that it, the quality would degree, like if you extend a few times to like one minute, the video quality would look much worse than the first video.
Second, another problem is that the model doesn't have long-range knowledge of what's happening before.
If they generate some dialogue to people speaking, their voice might change over some time, especially if the one-second conditioning does not cover the previous context.
So these are the core challenges.
So the Grok Imagine video extension, it has historical context of all of the previous generated videos.
It has a context of who is speaking and what objects have appeared and everything, having that to generate the next video.
So if we naively do this, you can imagine, like, just put all of the previous history video tokens into the context.
The context lens will easily explode.
Especially for video models, that can be like a few million context, I would imagine, context lens.
Yes.
What's wrong with that?
Yeah, for example, like in Cosmos, I think just five seconds of video is like a 50k or 60k number of tokens.
Showing 401–420 of 722 · page 21 of 37
← Previous
Next →