The Easy Goal Inference Problem Is Still Hard

Description

One approach to the AI control problem goes like this:Observe what the user of the system says and does.Infer the user’s preferences.Try to make the world better according to the user’s preference, perhaps while working alongside the user and asking clarifying questions.This approach has the major advantage that we can begin empirical work today — we can actually build systems which observe user behavior, try to figure out what the user wants, and then help with that. There are many applications that people care about already, and we can set to work on making rich toy models.It seems great to develop these capabilities in parallel with other AI progress, and to address whatever difficulties actually arise, as they arise. That is, in each domain where AI can act effectively, we’d like to ensure that AI can also act effectively in the service of goals inferred from users (and that this inference is good enough to support foreseeable applications).This approach gives us a nice, concrete model of each difficulty we are trying to address. It also provides a relatively clear indicator of whether our ability to control AI lags behind our ability to build it. And by being technically interesting and economically meaningful now, it can help actually integrate AI control with AI practice.Overall I think that this is a particularly promising angle on the AI safety problem.Original article:https://www.alignmentforum.org/posts/h9DesGT3WT9u2k7Hr/the-easy-goal-inference-problem-is-still-hardAuthors:Paul ChristianoA podcast by BlueDot Impact.Learn more on the AI Safety Fundamentals website.

Audio

Featured in this Episode

No persons identified in this episode.

Transcription

This episode hasn't been transcribed yet

Help us prioritize this episode for transcription by upvoting it.

0 upvotes

🗳️ Sign in to Upvote

Popular episodes get transcribed faster

Other recent transcribed episodes

Transcribed and ready to explore now

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

01 Jan 1970

El Partidazo de COPE

Buchladen: Tipps für Weihnachten

20 Dec 2025

eat.READ.sleep. Bücher für dich

BOJ alza 25pb decennale sopra 2%, Oracle vola con accordo Tik Tok, 90 mld eurobond per Ucraina | Morning Finance 

19 Dec 2025

Black Box - La scatola nera della finanza

365. The BEST advice for managing ADHD in your 20s ft. Chris Wang

19 Dec 2025

The Psychology of your 20s

LVST 19 de diciembre de 2025

19 Dec 2025

La Venganza Será Terrible (oficial)

Cuando la Ciencia Ficción Explicó el Mundo que Hoy Vivimos

19 Dec 2025

El Podcast de Marc Vidal

Comments

There are no comments yet.

Please log in to write the first comment.

BlueDot Narrated

This episode hasn't been transcribed yet

Other recent transcribed episodes

3ª PARTE | 17 DIC 2025 | EL PARTIDAZO DE COPE

Buchladen: Tipps für Weihnachten

BOJ alza 25pb decennale sopra 2%, Oracle vola con accordo Tik Tok, 90 mld eurobond per Ucraina | Morning Finance

365. The BEST advice for managing ADHD in your 20s ft. Chris Wang

LVST 19 de diciembre de 2025

Cuando la Ciencia Ficción Explicó el Mundo que Hoy Vivimos

Sign in to Audioscrape

Share this moment