Pengchuan Zhang
speaker
270 appearances
1 recordings
1 series
first heard Dec 2025
last heard 18 Dec
Pengchuan Zhang’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Dec 2025 with 1.
Appearances
Can SFT can really imitation learning, really do the job, get to very good performance.
But if you only do SFT and the SFT data is annotated by human, then your performance is buried by humans.
cannot get kind of superhuman performance just by kind of uh this kind of data angle approach to use human annotate data then learn from that.
You need to go to this RL HF domain that human really just tell which two point which one is better.
This is exactly kinda the philosophy that
to kinda to tell which one is better is easier to run it kinda to construct the data point from scratch.
So you can get kinda higher performance.
can can get better performance from human tool from scratch.
I would say that kinda I hope that after science three we can see kinda new research emerge from kinda in computer vision, which is okay, how we go beyond human performance.
Science three is close to that, but I would say that new learning paradigm is needed to go beyond human performance for science three tasks.
Wonderful computer vision.
Yeah, yeah.
I would say that's
running kinda good video, not language kind of video mode, not multimodal model.
So when we do sensory is kinda earlier this year or kind of last year, you can see that image not multimodal model is very good, but video not multimodal model, I think running kinda it gets become good or practical later this year.
Like kinda Queen Sway this kind of model gets kind of roughly kinda okay in that
stage.
So we have a good kind of base model to function our data and to get human performance for this recognition or verification task.
I would say that you can see that we need definite kind of sensory like efforts in the perception side, but we also need kind of this kind of m multimodal not language model kind of effort, kind of good foundation model on the kind of vision language side.
I think it's
Showing 121–140 of 270 · page 7 of 14
← Previous
Next →