Roman Yampolsky

speaker
257 appearances 1 recordings 1 series first heard Jun 2024 last heard Jun 2024

Roman Yampolsky’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
We are in a situation where people making more capable systems just need more resources. They don't need to invent anything, in my opinion. Some will disagree, but so far at least I don't see diminishing returns. If you have 10x compute, you will get better performance. The same doesn't apply to safety.
If you give Miri or any other organization 10 times the money, they don't output 10 times the safety. And the gap between capabilities and safety becomes bigger and bigger all the time. So it's hard to be completely optimistic about our... results here. I can name 10 excellent breakthrough papers in machine learning. I would struggle to name equally important breakthroughs in safety.
A lot of times a safety paper will propose a toy solution and point out 10 new problems discovered as a result. It's like this fractal. You're zooming in and you see more problems and it's infinite in all directions.
So I guess we can look at related technologies with cybersecurity, right? We did manage to have banks and casinos and Bitcoin, so you can have...
secure narrow systems which are doing okay uh narrow attacks on them fail but you can always go outside outside of the box so if i can't hack your bitcoin i can hack you so there is always something if i really want it i will find a different way we talk about guardrails for ai well that's a fence I can dig a tunnel under it, I can jump over it, I can climb it, I can walk around it.
You may have a very nice guardrail, but in a real world, it's not a permanent guarantee of safety. And again, this is a fundamental difference. We are not saying we need to be 90% safe to get those trillions of dollars of benefit. We need to be 100% indefinitely, or we might lose the principle.
I think we can generalize it to just prisoner's dilemma in general, personal self-interest versus group interest. The incentives are such that everyone wants what's best for them. Capitalism obviously has that tendency to maximize your personal gain, which does create this race to the bottom.
I don't have to be a lot better than you, but if I'm 1% better than you, I'll capture more of a profit, so it's worth for me personally to take the risk, even if society as a whole will suffer as a result.
Right. Look at the governance structures. Then you have someone with complete power. They're extremely dangerous. So the solution we came up with is break it up. You have judicial, legislative, executive. Same here. Have narrow AI systems work on important problems. Solve immortality.
It's a biological problem we can solve similar to how progress was made with protein folding using a system which doesn't also play chess. There is no reason to create super intelligent system to get most of the benefits we want from much safer narrow systems.
Like- But the bragging rights. But being first, that is the same humans who are in charge of the systems, right?
the condition would be not time, but capabilities. Pause until you can do X, Y, Z. And if I'm right and you cannot, it's impossible, then it becomes a permanent ban. But if you're right and it's possible, so as soon as you have the safety capabilities, go ahead.
So then I think about this problem. I think about having a toolbox I would need, capabilities such as explaining everything about that system's design and workings, predicting not just terminal goal, but all the intermediate steps of a system. control in terms of either direct control, some sort of a hybrid option, ideal advisor.
It doesn't matter which one you pick, but you have to be able to achieve it. In a book, we talk about others. Verification is another very important tool. communication without ambiguity, human language is ambiguous, that's another source of danger.
So basically, there is a paper we published in ACM surveys, which looks at about 50 different impossibility results, which may or may not be relevant to this problem, but we don't have enough human resources to investigate all of them for relevance to AI safety. The ones I mentioned to you I definitely think would be handy, and that's what we see AI safety researchers working on.
Explainability is a huge one. The problem is that it's very hard to separate capabilities work from safety work. If you make good progress in explainability, now the system itself can engage in self-improvement much easier, increasing capability greatly. So it's not obvious that there is any research which is pure safety work without disproportionate increase in capability and danger.
Right now, it's comprised of weights on a neural network. If it can convert it to manipulatable code, like software, it's a lot easier to work in self-improvement.
You can do intelligent design instead of evolutionary gradual descent.
Not completely. So if they're sufficiently large... you simply don't have the capacity to comprehend what all the trillions of connections represent. Again, you can obviously get a very useful explanation which talks about top, most important features which contribute to the decision, but the only true explanation is the model itself.
Absolutely, and you can probably have targeted deception where different individuals will understand explanation in different ways based on their cognitive capability. So while what you're saying may be the same and true in some situations, ours will be deceived by it.
Showing 141–160 of 257 · page 8 of 13 ← Previous Next →