Noam Brown

speaker
1,199 appearances 1 recordings 1 series first heard Dec 2022 last heard Dec 2022

Noam Brown’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
No recordings in the last 12 months.Older appearances are listed below; set an alert to hear about the next one.

Appearances

newest first · ▶ plays the moment
The approach that we took in 2017 was much more search-based.
It was trying to say, okay, well, let me in real time try to compute a much better strategy than what I had pre-computed by playing against myself during self-play.
So in a game like chess, the search is like, okay, I'm in this chess position and I can move these different pieces and see where things end up.
In poker, what you're searching over is the actions that you can take for your hand, the probabilities that you take those actions, and then also the probabilities that you take other actions with other hands that you might have.
And that's kind of hard to wrap your head around.
Why are you searching over these other hands that you might have and trying to figure out what you would do with those hands?
And the idea is, again, you want to always be balanced and unpredictable.
If your search algorithm is saying, oh, I want to raise with this hand, well, in order to know whether that's a good action, let's say it's a bluff.
Let's say you have a bad hand and you're saying, oh, I think I should be betting here with this really bad hand and bluffing.
Well, that's only a good action if you're also betting with a strong hand.
Otherwise, it's an obvious bluff.
Basically, what you want to do is put your opponent into a tough spot.
So you want them to always have some doubt, like, should I call here?
Should I fold here?
And if you are raising in the appropriate balance between bluffs and good hands, then you're putting them into that tough spot.
And so that's what we're trying to do.
We're always trying to search for a strategy that would put the opponent into a difficult position.
Yeah, ultimately what you're trying to maximize is your expected winnings, like your expected value, the amount of money that you're going to walk away from, assuming that your opponent was playing optimally in response.
So you're going to assume that your opponent is also playing as well as possible a Nash equilibrium approach, because if they're not, then you're just going to make more money, right?
Anything that deviates, like by definition, the Nash equilibrium is the strategy that does the best in expectation.
Showing 261–280 of 1,199 · page 14 of 60 ← Previous Next →