Noam Brown

speaker
1,199 appearances 1 recordings 1 series first heard Dec 2022 last heard Dec 2022

Noam Brown’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
But it's quite possible that you could expand the scope of intents to be able to allow it to talk about those things.
Now, in the process of doing that, the self-play would become much more complicated.
And so that is a potential for future work.
So we train in neural nets to imitate the human data as closely as possible.
And that's what we call the anchor policy.
And now when we're doing self-play,
The problem with the anchor policy is that it's not a perfect approximation of how humans actually play.
Because we don't have infinite data, because we don't have unlimited neural network capacity, it's actually a relatively suboptimal approximation of how humans actually play.
And we can improve that approximation by adding planning and RL.
And so what we do is we get a better approximation, a better model of human play by adding
During the self-play process, we say you can deviate from this human anchor policy if there is an action that has particularly high expected value.
But it would have to be a really high expected value in order to deviate from this human-like policy.
So you basically say try to maximize your expected value while at the same time stay as close as possible to the human policy.
And there is a parameter that controls the relative weighting of those competing objectives.
Well, it really comes down to how much human data you have.
I think the more human data you have, the better.
And I think that that's going to be the major bottleneck in scaling to more complicated domains.
But that said, there might be the potential, just like in the language model where we leveraged tons of data on the internet and then specialized it for diplomacy, there is the future potential that you can leverage huge amounts of data across the board and then specialize it in the data set that you have for diplomacy.
And that way you're essentially augmenting the amount of data that you have.
Well, like I said, the original motivation for the game of diplomacy was the failures of World War I, the diplomatic failures that led to war.
Showing 1041–1060 of 1,199 · page 53 of 60 ← Previous Next →