11 - Attainable Utility and Power with Alex Turner - AXRP - the AI X-risk Research Podcast | Transcription & Insights

Description

Many scary stories about AI involve an AI system deceiving and subjugating humans in order to gain the ability to achieve its goals without us stopping it. This episode's guest, Alex Turner, will tell us about his research analyzing the notions of "attainable utility" and "power" that underlie these stories, so that we can better evaluate how likely they are and how to prevent them. Topics we discuss: - Side effects minimization - Attainable Utility Preservation (AUP) - AUP and alignment - Power-seeking - Power-seeking and alignment - Future work and about Alex The transcript: axrp.net/episode/2021/09/25/episode-11-attainable-utility-power-alex-turner.html Alex on the AI Alignment Forum: alignmentforum.org/users/turntrout Alex's Google Scholar page: scholar.google.com/citations?user=thAHiVcAAAAJ&hl=en&oi=ao Conservative Agency via Attainable Utility Preservation: arxiv.org/abs/1902.09725 Optimal Policies Tend to Seek Power: arxiv.org/abs/1912.01683 Other works discussed: - Avoiding Side Effects by Considering Future Tasks: arxiv.org/abs/2010.07877 - The "Reframing Impact" Sequence: alignmentforum.org/s/7CdoznhJaLEKHwvJW - The "Risks from Learned Optimization" Sequence: alignmentforum.org/s/7CdoznhJaLEKHwvJW - Concrete Approval-Directed Agents: ai-alignment.com/concrete-approval-directed-agents-89e247df7f1b - Seeking Power is Convergently Instrumental in a Broad Class of Environments: alignmentforum.org/s/fSMbebQyR4wheRrvk/p/hzeLSQ9nwDkPc4KNt - Formalizing Convergent Instrumental Goals: intelligence.org/files/FormalizingConvergentGoals.pdf - The More Power at Stake, the Stronger Instumental Convergence Gets for Optimal Policies: alignmentforum.org/posts/Yc5QSSZCQ9qdyxZF6/the-more-power-at-stake-the-stronger-instrumental - Problem Relaxation as a Tactic: alignmentforum.org/posts/JcpwEKbmNHdwhpq5n/problem-relaxation-as-a-tactic - How I do Research: lesswrong.com/posts/e3Db4w52hz3NSyYqt/how-i-do-research - Math that Clicks: Look for Two-way Correspondences: lesswrong.com/posts/Lotih2o2pkR2aeusW/math-that-clicks-look-for-two-way-correspondences - Testing the Natural Abstraction Hypothesis: alignmentforum.org/posts/cy3BhHrGinZCp3LXE/testing-the-natural-abstraction-hypothesis-project-intro

Audio

Featured in this Episode

No persons identified in this episode.

Transcription

This episode hasn't been transcribed yet

Help us prioritize this episode for transcription by upvoting it.

0 upvotes

🗳️ Sign in to Upvote

Popular episodes get transcribed faster

Other recent transcribed episodes

Transcribed and ready to explore now

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

01 Jan 1970

El Partidazo de COPE

Buchladen: Tipps für Weihnachten

20 Dec 2025

eat.READ.sleep. Bücher für dich

LVST 19 de diciembre de 2025

19 Dec 2025

La Venganza Será Terrible (oficial)

Christmas Party, Debris & Ping-Pong

19 Dec 2025

My Therapist Ghosted Me

Episode 1320: Becoming 'The Monk': Rex Ryan on playing Gerry Hutch on stage (Part 1)

19 Dec 2025

Crime World

Friends Thru A Lens: The Holidays with Ella Risbridger

19 Dec 2025

Sentimental Garbage

Comments

There are no comments yet.

Please log in to write the first comment.

AXRP - the AI X-risk Research Podcast

11 - Attainable Utility and Power with Alex Turner

This episode hasn't been transcribed yet

Other recent transcribed episodes

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

Buchladen: Tipps für Weihnachten

LVST 19 de diciembre de 2025

Christmas Party, Debris & Ping-Pong

Episode 1320: Becoming 'The Monk': Rex Ryan on playing Gerry Hutch on stage (Part 1)

Friends Thru A Lens: The Holidays with Ella Risbridger

Sign in to Audioscrape

Share this moment