Menu
Sign In Search Podcasts Charts People & Topics Add Podcast API Pricing
Podcast Image

Astral Codex Ten Podcast

Deceptively Aligned Mesa-Optimizers: It's Not Funny If I Have To Explain It

12 Apr 2022

Description

https://astralcodexten.substack.com/p/deceptively-aligned-mesa-optimizers A Machine Alignment Monday post, 4/11/22 I. Our goal here is to popularize obscure and hard-to-understand areas of AI alignment, and surely this meme (retweeted by Eliezer last week) qualifies: So let's try to understand the incomprehensible meme! Our main source will be Hubinger et al 2019, Risks From Learned Optimization In Advanced Machine Learning Systems. Mesa- is a Greek prefix which means the opposite of meta-. To "go meta" is to go one level up; to "go mesa" is to go one level down (nobody has ever actually used this expression, sorry). So a mesa-optimizer is an optimizer one level down from you. Consider evolution, optimizing the fitness of animals. For a long time, it did so very mechanically, inserting behaviors like "use this cell to detect light, then grow toward the light" or "if something has a red dot on its back, it might be a female of your species, you should mate with it". As animals became more complicated, they started to do some of the work themselves. Evolution gave them drives, like hunger and lust, and the animals figured out ways to achieve those drives in their current situation. Evolution didn't mechanically instill the behavior of opening my fridge and eating a Swiss Cheese slice. It instilled the hunger drive, and I figured out that the best way to satisfy it was to open my fridge and eat cheese.

Audio
Featured in this Episode

No persons identified in this episode.

Transcription

This episode hasn't been transcribed yet

Help us prioritize this episode for transcription by upvoting it.

0 upvotes
šŸ—³ļø Sign in to Upvote

Popular episodes get transcribed faster

Comments

There are no comments yet.

Please log in to write the first comment.