Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
They can give some general prediction of which song is this, but it's very hard to describe the details of the music.
Like we mentioned in image generation, you have to describe image as detailed as possible so that someone blind can reconstruct that.
So here it's like someone deaf can reconstruct how the music sounds like.
without actually listening to it.
Maybe like, you can think of, it needs to have the, what they call the... Subtitles, yeah.
You gotta have all the details of the music and the dialogue.
So one important thing is the alignment.
So the model has to know the video and audio, it has to have a time-based alignment, like at which time step the video and the audio token correspond to each other.
We actually don't have this kind of alignment for most of the other modalities.
If you think about text and image, text and video, they are loosely aligned.
So you can have a description of...
what's going on in the video, but you don't have to exactly, you typically don't have exact description.
Oh, at time step one second, like what happened?
It's very coarse.
So that comes down to how you design the model for the model to be aware of as a time modality.
So the model is like a time-aware.
And that's something pretty unique if you think about OLMs.
So if you ask OLM to complete a task,
So you ask them, and they will say, oh, this task will probably take 12 hours to complete.
And they come back in one hour, say, I've already spent two days on this, and I've exhausted everything.
Showing 341–360 of 722 · page 18 of 37
← Previous
Next →