Noam Brown
speaker
1,199 appearances
1 recordings
1 series
first heard Dec 2022
last heard Dec 2022
Noam Brown’s voice in public audio — every appearance, attributed to the second.
Trend
recordings per month · last 12 monthsNo recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.
Appearances
Um, I think there's a few different directions to take this work.
Um,
I think really what it's showing us is the potential that language models have.
I mean, I think a lot of people didn't think that this kind of result was possible even today, despite all the progress that's been made in language models.
And so it shows us how we can leverage the power of things like self-play on top of language models to get increasingly better performance.
And the ceiling is really much higher than what we have right now.
I think one of the things that's really useful about diplomacy is that we have a well-defined value function.
There is a well-defined score that the bot is trying to optimize.
And in a setting like a general chatbot setting, it would need that kind of objective in order to fully leverage the techniques that we've developed.
The way that we've approached AI and diplomacy is you condition the language on an intent.
Now, that intent in diplomacy is an action, but it doesn't have to be.
And you can imagine, you know, you could have NPCs in video games or the metaverse or whatever, where there's some intent or there's some objective that they're trying to maximize, and you can specify what that is.
And then the language can correspond to that intent.
Now, I'm not saying that this is happening imminently, but I'm saying that this is like a future application potentially of this direction of research.
The way that we've approached self-play in diplomacy is we're trying to come up with good intents to condition the language model on.
And the space of intents is actions that can be played in the game.
Now, there is the potential to have a broader set of intents.
Things like long-term cooperation or long-term objectives or gossip about what another player was saying.
These are things that we're currently not conditioning the language model on.
And so we're not able to control it to say like, oh, you should be talking about this thing right now.
Showing 1021–1040 of 1,199 · page 52 of 60
← Previous
Next →