Micah-Hill Smith
speaker
601 appearances
1 recordings
1 series
first heard Jan 2026
last heard 8 Jan
Micah-Hill Smith’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsRecordings per month over the last 12 months — 1 in all, peaking in Jan 2026 with 1.
Appearances
Cause some of that, and MTV was a great project that was a good example of some of this a while ago, was about judging conversations and like a lot of style type stuff.
Here we've got the task that the gradering grading model is doing.
Is quite different to the task of taking the test.
When you're taking the test, you've got all of the agentic tools you're working with the code interpreter and web search, the file system to go through many, many turns to try to create the documents.
Then on the other side, when we're greening it, we're running it through a pipeline to extract visual and text versions of the files and be able to provide that to Gemini.
I
And we're providing the criteria for the task and getting it to pick which one more effectively meets the criteria.
Of the task out of two potential outcomes.
It turns out that we proved that it's just very, very good at getting that right, matched with human preference a lot of the time, because it's I think it's got the raw intelligence, but it's combined with the correct representation of the outputs, the fact that the outputs were created with an agentic task that is quite different to the way the grading model works, and we're comparing it against criteria, not just
I'm
kind of zero shot trying to ask the model to pick which one is better.
Oh uh what?
Like model has to go find clips on the internet and try to put it together.
The models are not that good at doing that one for now, to be clear.
It's pretty it's pretty hard to do that with a code adopter.
Um and the computer use stuff doesn't work quite well enough and so on and so on.
But um
So we like haven't grounded this score in that exact way.
I agree that it can be helpful, but we wanted to generalize this to a very large number of models.
That's one of the reasons that presenting it as ELO is quite helpful and allows us to add models and it'll stay relevant for quite a long time.
Showing 341–360 of 601 · page 18 of 31
← Previous
Next →