Noam Brown
speaker
1,199 appearances
1 recordings
1 series
first heard Dec 2022
last heard Dec 2022
Noam Brown’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
It's kind of like if you're developing a self-driving car, you don't want to measure that car on the road with a bunch of expert stunt drivers.
You want to put it on a road of like an actual American city and see, is this car crashing less often than an expert driver would?
So that's the metric that we've used.
We're saying like, we're going to stick this game, we're going to stick this bot in games with a wide variety of skill levels.
And then are we doing better than a strong or expert human player would in the same situation?
Yeah.
And honestly, when you look at what makes a good human diplomacy player, obviously they're able to handle themselves in games with other expert humans, but where they really shine is when they're playing with these weak players and they know how to take advantage of the fact that they're a weak player, that they won't be able to like pull off a stab as well, or that they have certain tendencies and they can take them under their wing and persuade them to do things that might not even be in their interest.
Yeah.
The really good diplomacy players are able to take advantage of the fact that there are some weak players in the game.
Yeah, so that's really the crux of the problem.
How do we leverage the benefits of self-play that have been so successful in all these other previous games while keeping the strategy as human compatible as possible?
And so what we did is we first trained a language model and then we made that language model controllable
on a set of intents, what we call intents, which are basically like an action that we want to play and an action that we would like the other player to play.
And so this gives us a way to generate dialogue that's not just trying to imitate the human style, whatever a human would say in this situation, but to actually give it an intent, a purpose in its communication.
We can talk about a specific move or we can make a specific request and the determination of what that move is that we're discussing
comes from a strategic reasoning model that uses reinforcement learning and planning.
It's a combination of reinforcement learning and planning.
Actually very similar to how we approached poker and how people approached chess and Go as well.
We're using self-play and search to try to figure out what is an optimal move for us and what is a desirable move that we would like this other player to play.
Now, the difference between the way that we approached reinforcement learning and search in this game versus those previous games is that we have to keep it human compatible.
Showing 841–860 of 1,199 · page 43 of 60
← Previous
Next →