Ethan He
speaker
722 appearances
1 recordings
1 series
first heard Jun 2026
last heard 1 Jun
Ethan He’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Also, they don't have music either.
The hard part is that
Actually, audio has two components.
It has like a discrete component, a continuous component.
The discrete component is like there's a language.
So when we speak, it's just some... It's an ASR issue.
Yeah.
Yeah, it's text token with some characteristics, I would say.
But music... I think the speech guys will disagree.
It's like disfluencies and then, you know... I'd say largely.
But the music is completely different.
It's...
It's very continuous and you cannot model them like discrete tokens in language models.
This is like the hard part for models is not to mention we have to align text, video and audio together.
Yeah.
How?
So some significant challenges are like... So first, we talk about as VLMs, they cannot understand... Most of them cannot understand audio.
So you have to have some way to do the synthetic data generation for...
for audio, you have to caption the model and that involves, that involves data and human data effort a lot.
And not just surprisingly, most of the RLMs are very bad at recognizing, um, like the, the beat tone and the details of the music.
Showing 321–340 of 722 · page 17 of 37
← Previous
Next →