Wendy Wang
speaker
80 appearances
1 recordings
1 series
first heard Jul 2026
last heard 20 Jul
Wendy Wang’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.
Appearances
So we had OCV data and thumbs down data.
So we are able to also see, okay, is this loss pattern actually happening at scale in production?
So when we detect that is the case, it makes it into the official loss pattern taxonomy.
I think what we also noticed in running these evals for over a year is that users' perception of what a great response is evolving very quickly.
So we've been able to also leverage these user evals as a way to set the standard of what great means and define great responses in its different
variables and metrics and then be able to
train the LM and also teach the LLM and human judges on what
great means and the dimensions of great responses, so that we are updating our human judges and our LLM judges on what is current, currently perceived as a great response by users.
Because we've seen that when a user gets a very detailed, personalized, you know, anticipatory response, they expect that going forward.
There's never a regression of expectations of what a great response looks like.
So the definition of a great response just continues to evolve.
And being able to leverage these user evals as basically a thermometer for the LM and the human judges.
Yeah, I think to add to that, um, I think it's we found that it's been very helpful to also do log inspection to actually
Once we've done these, you know, UX evals, have done the interviews to inform the attributes for the evals, we have a good instinct about, okay, what are the users actually saying about what great a great response looks like?
And what are the gaps that might be that we might see in our product?
Then actually do some log inspection and actually look at in production, are there signals of these behaviors that we're seeing in the
studies because and then be able to maybe even measure that at stale because what that does is getting the whole product and engineering team to understand not just these insights, but the prevalence of the insights in product and then be able to un build a process in which you can take these UX evals and actually inform what the product team does and how to do that and build the bridges to get there.
Yeah, I think what's been interesting is that a lot of the evals that teams are used to seeing are basically just metrics of, okay, we are, you know, this this number apart from our competitors.
It's almost like uh just a scorecard.
Um, but what we found's been really helpful in these UX evals is actually bringing the humans into the room and bringing the conversation.
Showing 41–60 of 80 · page 3 of 4
← Previous
Next →