Joseph Nelson
speaker
239 appearances
1 recordings
1 series
first heard Dec 2025
last heard 18 Dec
Joseph Nelson’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Dec 2025 with 1.
Appearances
The case.
For positive and negative examples, I don't know that I have seen like a golden ratio that that works well or not works well, but I can't offer anecdotally that a single negative example goes a long way.
A common place where fine-tuning is really helpful is like data that's out of distribution that might might have been impossibly in distribution.
Like one of my favorite fine-tuned examples is like counting Waymo's.
There's not that much data that have like Waymo's labeled throughout the streets of San Francisco.
But Sam does a really good job to identify Waymo as like a vehicle.
If you um
Prompt with Waymo, it doesn't find anything, you find vehicle, it find it labels a Waymo as a vehicle, which is valid, but a Waymo is a specific type of vehicle, right?
Usually from even just like a 10-second video clip, you can actually start to have SAM3 learn what should have been seen versus as a Waymo versus what should have been seen as a vehicle.
And even on a single image example, we see that like SAM3 starts to adapt because it takes the text and image prompt into
Into account when it makes a subsequent inference, from like three to five negative examples alongside positive examples, you start to see the model.
Update its priors, if you will, for where it would predict things from what the user provided.
All this is written with caveats, right?
Because like when you talk about visual world, the negative example and the positive examples could have been a very different perspective or a very different type of object.
Um like maybe you're like labeling dog breeds and suddenly a new dog breed appears, or maybe you have a perspective where it's overhead and then suddenly you have a side-by-side view.
So usually the best way is to like have these things meet the
Real world data and try, but I'll offer maybe the note that a small number of negative examples was a really long way, like small, like three to five, not like hundreds.
And to connect to something we just discussed earlier, you mentioned that at video component, each object needs to be tracked independently.
That's why the compute scales linearly with the number of classes, right?
Because each of those instance types needs to be maintained.
Showing 61–80 of 239 · page 4 of 12
← Previous
Next →