Roman Yampolsky

speaker
257 appearances 1 recordings 1 series first heard Jun 2024 last heard Jun 2024

Roman Yampolsky’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
Those are not the same. I am against the super intelligence in general sense with no undo button.
Partially, but they don't scale. For narrow AI, for deterministic systems, you can test them. You have edge cases. You know what the answer should look like. You know the right answers. For general systems, you have infinite test surface. You have no edge cases. You cannot even know what to test for. Again, the unknown unknowns are underappreciated by... people looking at this problem.
You are always asking me, how will it kill everyone? How will it will fail? The whole point is, if I knew it, I would be super intelligent, and despite what you might think, I'm not.
It is a master at deception. Sam tweeted about how great it is at persuasion. And we see it ourselves, especially now with voices, with maybe kind of flirty, sarcastic female voices. It's gonna be very good at getting people to do things.
Right. I don't think developers know everything about what they are creating. They have lots of great knowledge. We're making progress on explaining parts of a network. We can understand, okay, this node gets excited when this input is presented, this cluster of nodes. But we're nowhere near close to understanding the full picture, and I think it's impossible.
You need to be able to survey an explanation. The size of those models prevents a single human from observing all this information, even if provided by the system. So either we're getting model as an explanation for what's happening, and that's not comprehensible to us, or we're getting a compressed explanation, lossy compression, where here's top 10 reasons you got fired.
It's something, but it's not a full picture.
So there is a paper, I think it came out last week by Dr. Park et al from MIT, I think, and they showed that existing models already showed successful deception in what they do. My concern is not that they lie now and we need to catch them and tell them don't lie. My concern is that once they are capable and deployed, they will later change their mind because that's what
unrestricted learning allows you to do. Lots of people grow up maybe in the religious family. They read some new books and they turn in their religion. That's a treacherous turn in humans. If you learn something new about your colleagues, maybe you'll change how you react to them.
And you can't say they are not rational. The rational decision changes based on your position. Then you are under the boss. The rational policy may be to be following orders and being honest. When you become a boss, rational policy may shift.
The robots are coming. There's a refrigerator making a buzzing noise. Very menacing, very menacing. So every time I'm about to talk about this topic, things start to happen. My flight yesterday was canceled without possibility to rebook. I was giving a talk at Google in Israel and three cars, which were supposed to take me to the talk, could not. I'm just saying. I like AIs.
I for one welcome our overlords.
My claim is, again, that there are very strong limits on what we can and cannot verify. A lot of times when you post something on social media, people go, oh, I need a citation to a peer-reviewed article. But what is a peer-reviewed article? You found two people in a world of hundreds of thousands of scientists who said, I would have a publisher, I don't care. That's the verifier of that process.
When people say, oh, it's formally verified software, mathematical proof, they accept something close to 100% chance of it being free of all problems. But if you actually look at research, software is full of bugs. Old mathematical theorems, which have been proven for hundreds of years, have been discovered to contain bugs, on top of which we generate new proofs, and now we have to redo all that.
So, verifiers are not perfect. Usually, they are either a single human or communities of humans, and it's basically kind of like a democratic vote. Community of mathematicians agrees that this proof is correct, mostly correct. Even today, we're starting to see some mathematical proofs are so complex, so large, that mathematical community is unable to make a decision.
It looks interesting, it looks promising, but they don't know. They will need years for top scholars to study it, to figure it out. So of course we can use AI to help us with this process, but AI is a piece of software which needs to be verified.
Right. And for AI, we would like to have that level of confidence. For very important mission-critical software, controlling satellites, nuclear power plants, for small deterministic programs, we can do this. We can check that code verifies its mapping to the design, whatever software engineers intend it was correctly implemented.
But we don't know how to do this for software which keeps learning, self-modifying, rewriting its own code. We don't know how to prove things about the physical world, states of humans in the physical world. So there are papers coming out now, and I have this beautiful one. Towards guaranteed safe AI. Very cool paper. Some of the best authors I ever seen.
I think there is multiple Turing Award winners. You can have this one. One just came out, kind of similar, managing extreme AI risks. So all of them expect this level of proof, but... I would say that we can get more confidence with more resources we put into it. But at the end of the day, we're still as reliable as the verifiers. And you have this infinite regress of verifiers.
The software used to verify a program is itself a piece of program. If aliens give us well-aligned superintelligence, we can use that to create our own safe AI. But it's a catch-22. You need to have already proven to be safe system to verify this new system of equal or greater complexity.
Showing 101–120 of 257 · page 6 of 13 ← Previous Next →