Jaeden Schafer

speaker
4,189 appearances 35 recordings 4 series first heard Dec 2025 last heard 17 Jun

Jaeden Schafer’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
16 · Jan OctJan 26AprJulnow

Recordings per month over the last 12 months — 35 in all, peaking in Jan 2026 with 16.

Appearances

newest first · ▶ plays the moment
So that's what I'm excited about.
They were sharing a bunch of the results from some independent evaluations, a bunch of the benchmarks, especially humanity's last exam.
This is kind of, it feels like the AI models cooked a lot of the benchmarks and just kind of beat them.
And they weren't basically hard enough or built well enough for the AI models.
And so now we've kind of come up with some more challenging ones.
One of the more challenging ones is humanity's last exam.
Gemini 3.1 Pro outperformed Gemini 3.
I mean, obviously, if it didn't, I don't think they'd be releasing it to us, but it did it by a huge margin.
So the model's also coming up, climbing on a bunch of real world performance leaderboards.
This is what I think is actually the most important.
A lot of the leaderboards where people are like, look, we...
Basically, anytime these AI companies can test their own model on a benchmark, it feels like they are cheating, they're being scammy in some way.
And I mean, I don't mean to call the kettle black, but I feel like Anthropic, Google, and OpenAI have all been caught doing some form of this over the last few years.
So I don't really put as much stock in, you know, because basically those really screenshots where they're like, we scored, you know, 72% on this exam and have a screenshot where it's like they skipped out a couple of the questions that it probably didn't do good on.
Anyways, I'm not saying this is Google, but there is an AI company that has done this.
And so when it comes to these companies testing themselves, I trust them a lot less than the real world leaderboards.
So some of those examples are.
are when essentially they have side-by-side comparisons of their model versus another model, and they just have people give blind, they vote blindly on which response they like better.
And when a new model really starts crushing it on those type of leaderboards, I take stock in that because this is actual people blind testing saying that their model is better.
So that's great.
Showing 1121–1140 of 4,189 · page 57 of 210 ← Previous Next →