Tina Eliassi-Rad

speaker
154 appearances 1 recordings 1 series first heard Jan 2025 last heard Jan 2025

Tina Eliassi-Rad’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
You're trying to understand what is going on, what is the underlying process that is happening in this network and why these links exist. Now, the one thing that makes studying of graphs and networks really interesting is that it is not a closed world. So just because you didn't see a link between me and Jennifer doesn't mean that we're not friends.
And so for machine learning where you need both positive examples and both negative examples, which negative examples do you pick becomes difficult because the edges or the links or the friendships that don't exist may because like they don't want to be friends or for other reasons. And so this what are the negative examples becomes an important aspect of things.
Indeed, indeed. So there are lots of assumptions being made, obviously, in terms of like how the network is being observed. And in fact, this is one of the big differences between computer scientists and
that study graphs and network scientists that are typically physicists or social scientists, where, for example, they're like, well, there's a distribution and this graph fell from it versus like the machine learning graph mining folks typically don't question where the graph came from. They're like, oh, here's data and they run with it. Right.
And it's just it boggles the mind that like you should think about where this data came from, how it was collected, What were maybe the errors in collecting it? And in fact, this touches on a sore point for me because what happens is they don't question the data, right? They just like feed it into their machine learning AI models. And then on the other end, they don't measure any uncertainty.
So like if you have something like, let's say, a social network that you've observed, there's all this stuff about like representation learning, right? Where basically I take Tina in the social network and I represent her as a vector in a Euclidean space, right? Like maybe with 60,000, a vector with 16,000 elements in it. So the cardinality is 16,000 and there's no uncertainty.
They're like, no, Tina falls exactly here and it just doesn't make sense at all, right? And so then those kinds of models, given that, You didn't start with, okay, well, my data could have some noise in it, some uncertainty in it. And then you don't even capture the uncertainty of the model at the end.
It just, there are lots of problems that can occur, including, for example, adversarial attacks or like your model is not just going to be, your model is not going to be robust. Let's just put it that way.
Yeah, I think in part, one of the reasons that folks, at least in the CS side, the computer science and the machine learning side, aren't too bothered by it these days is because we are going through this era where prediction is everything. Prediction and accuracy is everything. And so, you know, there are these benchmarks and it's basically benchmark hacking or state of the art hacking, right?
And that's basically what is going on. You know, that's the reality of it, you know. And and so so there's a lot of that kind of engineering going on as opposed to like really thinking about what is the phenomena that I'm interested in? How is the data coming to me? What are the sources of noise? Should I how should I take them into account? Should I even take them into account?
And what are the uncertainties in terms of the predictions that I am outputting?
Yeah, so basically you create a bunch of data and you get a buy-in from the community that these are good data sets to test a machine learning or an AI model on. And then there's a leaderboard and you want to be number one. Right. And so you hack the systems that exist or you hack your own system. You create your own to be number one, you know, as as much as possible.
And that's basically what is going on. And I like that.
this metaphor so my colleague um barabaschi said it's like there are two camps there's like a toolbox it's a finite toolbox right and the machine learning the ai people the engineers put tools into that toolbox and because it's finite it's very competitive that is my tool beats your tool even if it's like one percent by one percent that it's not clear if it's statistically significant or not and i may be king for only 30 seconds
because another tool comes in, right? And then there's like the scientists on the other end that just open the toolbox and say, okay, well, what is good for whatever, you know, whatever prediction task I want to do. And then they pick a tool out of that.
And so a lot of this like benchmark hacking or state of the art hacking happens on the engineering, on the AI machine learning side, the computer science side, because you want your tool in that finite toolbox.
It is a very big problem. I mean, there are multiple angles to this. So one is, for example, because of all the hype, oftentimes people on the engineering side don't talk about the assumptions that they have made or the technical limitations of their system. Because of that, we have this reproducibility problem.
So not even a replicability problem, but a reproducibility problem, which is just a code. Can I just reproduce your code as you have it? Right. And even with your training data, even with like how you broke it up with these different like folds or whatever, you know, and so which is like very, very, very low bar to pass.
But that doesn't happen because there are lots of assumptions that are being made, etc. Then there's this notion of we are living through this era of big models. I want a model that has many, many, many parameters, even if I don't need all those many parameters. Or for example, maybe I do care about interpretability. That is, I want to know what the model is actually doing.
But because, again, for that one or two percentage point on the prediction side, you let go of it and you go with the bigger models. But yes, it's a big, big problem. For me, the lowest bar would be that we require, at least with federal funding, and in some of the service that I do for the federal government, I've been pushing this.
Showing 21–40 of 154 · page 2 of 8 ← Previous Next →