Roman Yampolsky
speaker
257 appearances
1 recordings
1 series
first heard Jun 2024
last heard Jun 2024
Roman Yampolsky’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
So when I wrote a paper, Artificial Intelligence Safety Engineering, which kind of coins the term AI safety, that was 2011. We had 2012 conference, 2013 journal paper. One of the things I proposed, let's just do formal verifications on it. Let's do mathematical formal proofs. In the follow-up work, I basically realized it will still not get us 100%. We can get 99.9.
We can put more resources exponentially and get closer, but we never get to 100%. If a system makes a billion decisions a second and you use it for 100 years, you're still going to deal with a problem. This is wonderful research. I'm so happy they're doing it. This is great, but it is not going to be a permanent solution to that problem.
There are many, many levels. So first you're verifying the hardware in which it is run. You need to verify communication channel with the human. Every aspect of that whole world model needs to be verified. Somehow it needs to map the world into the world model. Map and territory differences. So how do I know internal states of humans? Are you happy or sad? I can't tell.
So how do I make proofs about real physical world? Yeah, I can verify that deterministic algorithm follows certain properties. That can be done. Some people argue that maybe just maybe two plus two is not four. I'm not that extreme. But once you have sufficiently large proof over sufficiently complex environment, the probability that it has zero bugs in it is greatly reduced.
If you keep deploying this a lot, eventually you're going to have a bug anyways.
There is always a bug. And the fundamental difference is what I mentioned. We're not dealing with cybersecurity. We're not going to get a new credit card, new humanity.
You can improve the rate at which you are learning. You can become more efficient meta-optimizer.
So if you have fixed code, for example, you can verify that code, static verification at the time. But if it will continue modifying it, you have a much harder time guaranteeing that important properties of that system have not been modified, then the code changed.
It can always cheat. It can store parts of its code outside in the environment. It can have kind of extended mind situation. So this is exactly the type of problems I'm trying to bring up.
So I like Oracle types where you kind of just know that it's right. Turing likes Oracle machines. They know the right answer. How? Who knows? But they pull it out from somewhere, so you have to trust them. And that's a concern I have about humans in a world with very smart machines. We experiment with them.
We see after a while, okay, they've always been right before, and we start trusting them without any verification of what they're saying.
We remove ourselves from that process. We are not scientists who understand the world. We are humans who get new data presented to us.
preserved portion of it can be done. But in terms of mathematical verification, it's kind of useless. You're saying you are the greatest guy in the world because you are saying it. It's circular and not very helpful, but it's consistent. We know that within that world, you have verified that system. In a paper, I try to kind of brute force all possible verifiers.
It doesn't mean that this one is particularly important to us.
Any smart system would have doubt about everything, right? You're not sure if what information you are given is true, if you are subject to manipulation. You have this safety and security mindset.
I may be wrong, but I think Stuart Russell's ideas are all about machines which are uncertain about what humans want and trying to learn better and better what we want. The problem, of course, is we don't know what we want and we don't agree on it.
It could also backfire. Maybe you're uncertain about completing your mission. Like I am paranoid about your cameras not recording right now. So I would feel much better if you had a secondary camera, but I also would feel even better if you had a third. And eventually I would turn this whole world into cameras pointing at us, making sure we're capturing this.
So it's a multi-objective optimization. It depends how much I value capturing this versus not destroying the universe.
You might be scared to do anything.
Mess things up.
Showing 121–140 of 257 · page 7 of 13
← Previous
Next →