Joseph Nelson

speaker
239 appearances 1 recordings 1 series first heard Dec 2025 last heard 18 Dec

Joseph Nelson’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Dec OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Dec 2025 with 1.

Appearances

newest first · ▶ plays the moment
In your example of SAM3 as a tool call or SAM three as natively a part of the multimodal models,
Which future do you think is more likely?
So for very simple system one questions, things like how many fingers on a hand, that should be native.
But for maybe more complex things that are maybe long running tasks and long running reasoning, then maybe there's a bit more of like a tool call approach.
Maybe slightly different question.
SAM3 is an incredibly powerful piece of work.
Um and it's open source as a part of uh now MSL.
Open source critical to achieving AGI?
Uh re-where things are going based on like community questions.
Similar to how Nikila mentioned after Sam 1, the like almost obvious thing that people wanted was like open concept prompting.
Because people are like, great, this model can see things, but I want to tell it what I want it to see.
And now with the introduction of SAM 3, you have this stepwise component, which feels like a key component of
You know, the Chat GPT era for vision is is arriving as a result.
What's going to happen is now you've provided people with an open text box and media.
And so you're going to get all sorts of queries from people that maybe the model isn't primed to be able to perform particularly well on yet.
For example, earlier we were talking about document understanding and document reasoning being a place where there's known improvements to be made.
And so you'll have people that will probably prompt to try to OCR things.
Or you'll have people that will want to do work with spatial reasoning.
Reasoning, like give me the object to the left of this other object, or give me a sense of where things are in relation to one another, which is critical for robotics, like we were discussing, because that's how you navigate throughout the real world.
You'll also have, I think, people will want action recognition and vision language action models, VLAs, like the same things that where you have these tasks where people are used to providing open text prompts and getting here's this part of the scene where the player kicked the
Showing 161–180 of 239 · page 9 of 12 ← Previous Next →