Moritz Sudhof

speaker
399 appearances 1 recordings 1 series first heard Jul 2026 last heard 8 Jul

Moritz Sudhof’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And so the really important thing for for product builders is not to rely on or or put the burden on the user to catch when the AI is confidently wrong and dressing that up with a lot of specific.
specifics, it's to look specifically for those confidence signatures in their AI response.
Well first with evals.
The most important gap in my lived experience is they only catch what you know to look for.
AI, particularly as you deploy to users who don't look like the people in your organization who wrote the evals.
The set of things that they actually do and the emergent behavior that engenders from the model are not things that you could have a priori thought about or enumerated or translated into evals.
And also, you know, evals are very hard to write.
They're very time consuming.
And in our research, the kinds of failures that are happening are not the kinds of failures that are caught by, oh, I I need an eval for a specific guardrail.
They're not failures often that happen with one specific turn, which often evals are just checking, but is a single response accurate or does it contain these markers?
These are failure modes that happen over the interaction.
They happen over the course of a couple turns.
And so to catch them, to have a system that can identify and detect that they're happening and surface that and measure that, it needs to be deployed online, looking at the actual interactions and the full arc of the interaction between the user and the model.
And it uh needs to be attuned to very
very specific patterns and trends in the user experience and and what the user is actually getting out of the interaction.
And so evals are a necessary step, but they catch only the things you know to think for, and they only catch them on generally on single turns in kind of a walled setting of a small set of test scenarios.
Yeah, so the silent mismatch is when the AI response does not
Answer exactly what the user intended or wanted.
But it doesn't surface that gap.
It just plows ahead.
Showing 161–180 of 399 · page 9 of 20 ← Previous Next →