Jongmin (Jeongmin) Baek
speaker
142 appearances
1 recordings
1 series
first heard Jun 2026
last heard 30 Jun
Jongmin (Jeongmin) Baek’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.
Appearances
Working Smarter · How agentic AI works behind the scenes to find the answers you need · 30 Jun 2026
podcast
Is this correct?
We rely on this pattern for a lot of, you know, sanity checks and evaluations on the generic things that agents can perform.
The place where it gets tricky is a lot of the tasks that we want agents to do for Dropbox and Dash are very contextual.
So if you want the agent to come and fetch the latest updates in your project,
you're really the only person that can decide whether the update was proper or not.
You could automate this to some extent and evaluate certain aspects, like was the agent's updates grammatically correct and had the right tone?
So all these things are things that you can ask an LLM to evaluate.
And these days, LLMs do a good job.
And that's the basis for a lot of automated evaluations that run on all the executions that we have.
But the tricky part is, you know, how do you...
build data as an evaluation where you hope to mimic the contextual nature of the things that the agents are supposed to do.
You want to converge your automated evaluation with your manual evaluation.
So the former is much cheaper and scalable, but less accurate.
The latter is accurate, but more expensive.
So now there is a lot of meta-evaluation work that happens across industry where you try to converge the two in some automated way.
So this is something that we do a lot at Dropbox to tune the automated evaluation we use to mimic human evaluation as much as possible.
In order to scale the automated evaluation we have and make sure that the quality at least converges towards that of human evaluation, we have...
you know, continuous inputs to the, you know, human evaluation as a form of labels.
So, you know, people say this is right or this is wrong.
And then you can do this, you know, meta prompting where you try to see if the element that you use to automate evaluation gives the same judgment as humans did on the data set that humans actually annotated.
Showing 41–60 of 142 · page 3 of 8
← Previous
Next →