Jongmin (Jeongmin) Baek

speaker
142 appearances 1 recordings 1 series first heard Jun 2026 last heard 30 Jun

Jongmin (Jeongmin) Baek’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
And once they converge, then you have a slightly stronger conviction that the large-scale automated evaluation it does on your unlabeled conversations that the agent has will be informative in terms of its quality.
I just mentioned, okay, we're going to automate the automation by using LLMs to evaluate the agents.
And then we try to match that to human performance, but that process itself can be automated.
Right.
So we have LLMs that automatically tune the judges.
until they converge to human performance.
Yes.
So this, once it's working, provides an automated large-scale evaluation of your agent's sessions.
And the nice thing about it is you can apply it to sessions that happen, or you can evaluate a new agent on past data that you had before, that you have collected as a data set, in order to validate new behavior from a new agent.
So imagine you added a new tool.
Maybe you have a better tool than the previous tool you had.
Then you can backtest your new agent that's equipped in your tool against past data and then have the same LLM that was acting as a judge decide or evaluate whether the sessions were proper.
So once you have it set up and you trust the setup, it becomes sort of a flywheel.
But of course, it's important to have humans in the loop to decide that the overall systems are working as intended.
Generally, yes.
But you have to keep in mind that we had this automated loop to tune the agents to the given setup, which includes the choice of model.
So sometimes you see that after you swap out the model, the performance is lower because everything was optimized for the previous model that we had.
But then the exercise becomes, well, can we apply the same techniques, same optimizations, same automation to the new model, right?
And then retune the agent to play nice with this given model.
And generally, you know, that improves things over the baseline.
Showing 61–80 of 142 · page 4 of 8 ← Previous Next →