Illia Polosukhin

speaker
714 appearances 3 recordings 1 series first heard Oct 2024 last heard Sep 2025

Illia Polosukhin’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
And so you can actually evaluate how good that like system is.
uh by training you know a few sizes of the models and so people can actually compete on whose curriculum builder is better, right?
And so then when you pick the best curriculum builder, then you run a larger run with m you know more parameters and more data.
So so uh
I would say like i it's it's gonna be a kind of a a very iterative process, right?
And the same way it is right now in in in inside Frontier Labs because indeed we are like we're kind of experimenting with all those things as we go.
Uh now with pri specifically private user data, it's a bit more sensitive and complicated.
And but again, the b biggest benefit we can do is like as you contribute your data, first of all you can get a reward from the the model uh kind of outcome.
But you can also like we can have effective as a filtering procedure.
Because one of the things is like there's gonna be a lot of PII information, your phone numbers, SSN, etcetera, that you don't want in the model.
And it's also like not actually helpful for the model for the most part.
Right.
Um and so but this filtering procedure will be like open source and everybody can affecting, you know, not everybody will inspect, but some developers will inspect like and say, yes, this is actually valid uh like way to filter data.
So so there's like a few different components again that that we can do to make sure that both data is cleaner and then how it's get filtered and and kind of processed into the curriculum of the training.
And probably, you know, a bunch of your, you know, personal chats are probably not that actually useful for training these models.
There's like other types of data that's more useful uh for those things.
And we see like in general we're moving to more synthetic data anyway.
So there are gonna be a lot more of that.
You can still and importantly, we still have, you know, the crowdsourcing and data labeling that's run through this decentralized network as well.
Right.
Showing 521–540 of 714 · page 27 of 36 ← Previous Next →