James Reggio
speaker
667 appearances
1 recordings
1 series
first heard Jan 2026
last heard 17 Jan
James Reggio’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.
Appearances
Latent Space: The AI Engineer Podcast · Brex’s AI Hail Mary — With CTO James Reggio · 17 Jan 2026
podcast
On the product AI side, that's where it starts getting a little bit more challenging because the multi multi agent network um is quite challenging to evaluate.
And so what we do there is we try to adopt some of the state of the art for multi turn evals where we will um will basically have a an agent embody the user and like you know have
But basically the the end user
agent is given an objective and then we basically have it run a multi turn um
uh conversation and then use L Ms.
Judge at the end to do all of the different uh facet assessment.
The one other thing that we do technique wise that is interesting is sometimes you want to
you don't want to do like, you know, I think these multi-turn evals are kind of like integration tests.
They um they sometimes test more than than you what you want to to assess.
And so sometimes what we'll do is we'll also pre-can like an initial preamble to a conversation or maybe a couple turns will be handwritten and we'll basically set set the the um eval to start and we'll see if uh we're able to like isolate certain um certain behaviors
So uh
It's it's still like a work in progress.
And I say like at the end of the day, a lot of the just um periodic rev human review and and like looking at um at cases where uh we've detected as we go to like summarize um like what we'll do is we'll reflect on a conversation after a certain amount of time has has passed where we'll summarize it, like extract
Facets like, did it seem like the user accomplished their their objective?
And we'll just manually when uh uh a lot of the cases when that's that's failed and decide to write another eval for it.
Yeah, it's interesting.
Um, I don't know if we have any that that are like, oh
Someday I hope it'll be good enough to do this, but it's like there there are the evals that are are blocking because they would indicate like a uh a regression, an unacceptable regression.
So these tend to be just accuracy related um
evals, but then there are others that are more about like tone and uh coherency and these types of things where they're they're more subjective and we were just looking at those over time as a as a metric.
Showing 481–500 of 667 · page 25 of 34
← Previous
Next →