Moritz Sudhof

speaker
399 appearances 1 recordings 1 series first heard Jul 2026 last heard 8 Jul

Moritz Sudhof’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
A lot of failures that are happening that aren't visible.
Yeah, I I would describe it broadly as behavior.
You know, I have a young son and sometimes it's helpful for me to to think about it also from that aspect of obviously I want my son to learn a lot, to practice critical thinking, to be really smart, all those kind of be competent.
But the things that I feel like I pay the most attention to in parenting is how he behaves, how he makes decisions to step in or sit back and listen, how he makes decisions when to advocate for himself or when to clarify.
These are all things that have nothing to do with raw IQ.
They're about how we actually behave when we're in the room with somebody, when we have something that we're trying to achieve.
And that's not something that magically gets fixed if you make a model get a higher score on an LSAT.
Um and that's precisely what matters when there's a real user interaction with a model, like the behavior part is so important.
And so we we ran a study on this.
We did an experiment just to look at how does model improvement from a capability perspective
How does it change how many failures users are exposed to and how many these AI interactions are failing?
So we looked at GPT-4 and even 3.5, and we compared that to the the most recent kind of frontier models, GPT-5 and also uh Opus.
you can clearly track these kind of factual errors and capability errors going down.
So the models are in these interactions, they're getting better at kind of understanding what the user roughly wants.
They're making fewer factual errors that they made a couple of years ago.
The capability has improved.
But they are making the same number, if not more, of these behavioral moves that lead that are not in themselves necessarily bad, but that lead to failures in the AI interaction.
So the the biggest one, just to make this super concrete, the biggest one is models that
give you an answer rather than than clarifying.
So they have an ambiguous prompt from a user, an ambiguous request, or maybe there's there's like something that's not clear or there's some tension of like, wait, do you think this pr this requirement is more important?
Showing 61–80 of 399 · page 4 of 20 ← Previous Next →