https://astralcodexten.substack.com/p/constitutional-ai-rlhf-on-steroids A Machine Alignment Monday post, 5/8/23 What Is Constitutional AI? AIs like GPT-4 go through several different 1 types of training. First, they train on giant text corpuses in order to work at all. Later, they go through a process called "reinforcement learning through human feedback" (RLHF) which trains them to be "nice". RLHF is why they (usually) won't make up fake answers to your questions, tell you how to make a bomb, or rank all human races from best to worst. RLHF is hard. The usual method is to make human crowdworkers rate thousands of AI responses as good or bad, then train the AI towards the good answers and away from the bad answers. But having thousands of crowdworkers rate thousands of answers is expensive and time-consuming. And it puts the AI's ethics in the hands of random crowdworkers. Companies train these crowdworkers in what responses they want, but they're limited by the crowdworkers' ability to follow their rules.f
No persons identified in this episode.
This episode hasn't been transcribed yet
Help us prioritize this episode for transcription by upvoting it.
Popular episodes get transcribed faster
Other episodes from Astral Codex Ten Podcast
Transcribed and ready to explore now
Your Review: Joan of Arc
07 Aug 2025
Astral Codex Ten Podcast
Book Review: Selfish Reasons To Have More Kids
03 Jun 2025
Astral Codex Ten Podcast
Links For February 2025
11 Mar 2025
Astral Codex Ten Podcast
The Emotional Support Animal Racket
28 May 2024
Astral Codex Ten Podcast
The Psychopolitics Of Trauma
27 Jan 2024
Astral Codex Ten Podcast
Book Review: A Clinical Introduction To Lacanian Psychoanalysis
27 Apr 2022
Astral Codex Ten Podcast