Jongmin (Jeongmin) Baek
speaker
142 appearances
1 recordings
1 series
first heard Jun 2026
last heard 30 Jun
Jongmin (Jeongmin) Baek’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Working Smarter · How agentic AI works behind the scenes to find the answers you need · 30 Jun 2026
podcast
And once they converge, then you have a slightly stronger conviction that the large-scale automated evaluation it does on your unlabeled conversations that the agent has will be informative in terms of its quality.
I just mentioned, okay, we're going to automate the automation by using LLMs to evaluate the agents.
And then we try to match that to human performance, but that process itself can be automated.
Right.
So we have LLMs that automatically tune the judges.
until they converge to human performance.
Yes.
So this, once it's working, provides an automated large-scale evaluation of your agent's sessions.
And the nice thing about it is you can apply it to sessions that happen, or you can evaluate a new agent on past data that you had before, that you have collected as a data set, in order to validate new behavior from a new agent.
So imagine you added a new tool.
Maybe you have a better tool than the previous tool you had.
Then you can backtest your new agent that's equipped in your tool against past data and then have the same LLM that was acting as a judge decide or evaluate whether the sessions were proper.
So once you have it set up and you trust the setup, it becomes sort of a flywheel.
But of course, it's important to have humans in the loop to decide that the overall systems are working as intended.
Generally, yes.
But you have to keep in mind that we had this automated loop to tune the agents to the given setup, which includes the choice of model.
So sometimes you see that after you swap out the model, the performance is lower because everything was optimized for the previous model that we had.
But then the exercise becomes, well, can we apply the same techniques, same optimizations, same automation to the new model, right?
And then retune the agent to play nice with this given model.
And generally, you know, that improves things over the baseline.
Showing 61–80 of 142 · page 4 of 8
← Previous
Next →