Josh McGrath

speaker
251 appearances 1 recordings 1 series first heard Dec 2025 last heard 31 Dec

Josh McGrath’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Dec OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Dec 2025 with 1.

Appearances

newest first · ▶ plays the moment
Like see there's there's something strained that I think we haven't like looked at that axis of okay, well how like sort of clean is this signal, how much do I trust it?
And like I totally agree that you know, you don't necessarily trust the RLH signal as much as like is this the solution to this polynomial?
But I think there's a whole spectrum of like how high quality
the signal, what's gonna happen when I like do a lot of optimization against it.
And that's very different than I think worrying about like the variance of different gradients, which I think is what you end up seeing in a lot of the the papers that are currently coming out.
Um, rather than being like very data centric, they're pretty optimization centric, even though I think the the innovation really is is where the data's coming from.
Uh well, anthropic and deep mind we're all saying I'm working on stuff and things, you know.
Uh I I think like it's more so uh talking a lot more broadly with my my friends there.
Or or we're just talking about man, the the infra's so hard to keep up.
We're not not necessarily talking too much about methods directly.
Um because on on one level it kind of doesn't matter.
Yeah.
And I think also like there's there's something that's very different about academic work where like the what really matters is how narrativizable it is.
And I think that's one of the reasons you see like a lot of optimization papers come out is a lot of the data work, there's a less clear narrative around it.
But it doesn't have like necessarily the same narrative that you get out of like some of the papers that you see here.
And so like there becomes more of a like
Given a a specific vertical, how do I like understand that?
Um and I I wish there was actually more papers on it here, but I think it can sometimes be harder to wrap up into a a clean story.
Yeah, definitely.
I mean like yeah, the as you said, it came out in the Deep Seek math paper and like it's an interesting optimization method, but is like the more interesting thing that they have a new reward signal that they sort of like re that we can really, really trust.
Showing 101–120 of 251 · page 6 of 13 ← Previous Next →