Rohin Shah

speaker
1,071 appearances 1 recordings 1 series first heard Jun 2026 last heard 2 Jun

Rohin Shah’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jun 2026 with 1.

Appearances

newest first · ▶ plays the moment
Most of the time, research will look at, did it actually solve the problem that it set out to solve?
Definitely important.
You've got to do that evaluation.
That is the most important one.
But then there's also things like, how much cost does this add in terms of compute?
How much latency will it add?
If you're imagining a step that runs after the AI system has produced a response, and then you do lots of additional things before you then have to send the response to the user, probably a non-starter.
Not obviously, but it's a big cost.
There's implementation complexity or organizational complexity.
If your solution involves doing something at the scaffold level and also doing something that involves reading the internals of the AI system and connecting these up together, that just spans so many different teams and so many different abstraction layers in the stack that it's going to be much, much harder to implement.
than something that, for example, it just happens at inference time and involves running a monitor and then alerting some teams about it.
Yeah, I think there are several.
So for example, if their data sets can be a useful contribution, it's pretty easy to try to add a data set to post-training.
Sometimes you try to add it to post-training, and then for some reason, it makes the model worse on some totally unrelated thing, and then you can't use that dataset, which is a bit unfortunate.
So ideally, when you're building datasets, you would like to evaluate whether it could have some sort of negative effect on something else, but this is a bit hard to do.
But if someone came to me with a data set and said, look, training on this data set improves this metric by a decent amount and doesn't seem to have any bad effects on these three obvious other metrics that it might have had bad effects on, I'd find that fairly compelling and think, yeah, maybe we should do it, assuming it was solving a real problem.
So I think data sets, metrics, evals, these are usually relatively easy things to do.
Monitors, I think all of those are fairly easy to implement if we think it's worth implementing.
Yeah, any other advice?
Oh, one is just like,
Showing 961–980 of 1,071 · page 49 of 54 ← Previous Next →