Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Yeah, so the LLMs themselves, they don't have a sense of time there?
It came from the corpus on the internet.
People have an estimate.
Yeah, so disclaimer, I'm not going to debate what is word model.
There are many definitions, so I'll just talk about my definition.
Since I came from the multi-model domain, so mainly talking from video, so word model is like...
real-time, interactive, long-horizon videos.
So there are three parts.
So let's talk about them one by one.
So interaction.
So we just look at Facebook and the neural computer.
So the interaction part of it, so your work model can allow you to interact with them through keyboard, mouse, and maybe also voice.
So all of these modalities, you can interact with the model, and the model should respond reasonably.
Second part is real-time.
So once you move your mouse, say the world model generates a game, how fast can that game respond?
So if you're a professional CSGO player, you might say, oh, you have to respond in sub...
sub 10 million seconds or even less so that's not i guess most of the oh 60 fps let's go 300 fps 500 fps wait uh okay yeah i didn't do the math but yeah okay yeah 300 fps that's a 3 million seconds so you have to respond oh shit most of the video models cannot do that yeah
But if you have a video model that is, say, like a digital human, the response time might be more generous.
Typically, for real-time voice interaction, it's like 200 milliseconds.
So that's much more generous.
Showing 361–380 of 722 · page 19 of 37
← Previous
Next →