Nufar Gaspar

speaker
2,488 appearances 11 recordings 1 series first heard Oct 2025 last heard 3 Sep

Nufar Gaspar’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 11 in all, peaking in Jun 2026 with 5.

Appearances

newest first · ▶ plays the moment
So in a simple word, token is a chunk of a text that the model reads and writes.
It's typically bigger than one character, and it's usually smaller than a word.
And if you have never, ever seen a tokenizer or how tokens look in action, OpenAI has a very good page that is open to everyone that you can just take a look at how tokens actually look.
So it looks something like that.
You can paste the text and then you will see how words are being chanced.
So you can see
that some words are stayed staying like as one token while others might be separated into multiple tokens and interestingly numbers often are being chopped in the middle and so on so we'll put it in the show notes but a very interesting experiment if you have never seen how your text looks and by the way if you paste a non-english or non-latin language you will see that typically the amount of tokens is much larger than an english language
So that's the OpenAI tokenizer.
And a few things just to lend it home.
In general, like the ratio in English is around three quarters of a word to token, meaning that if you have a page of text, it's roughly 1000 tokens.
And some languages that are like Hindi, Thai, Greek, other languages like that might get two to five X more tokens for the same content.
And because billing is per token, then some questions, if you ask them in other languages, might cost you much more.
And that's sometimes referred to as
language tax.
With AI, with code also, it's different and it has its own way.
Indentation and brackets and white spaces, they all become tokens.
There are some newer ways to tokenize text that are more code friendly in order to do that, but still numbers is a huge problem.
So you've seen the one, two, three, four, five being chopped in the middle.
And by the way, that's also why whenever everybody's doing like the strawberry test for AI and it's
very badly fails in trying to count how many R's are in the word strawberry.
Showing 661–680 of 2,488 · page 34 of 125 ← Previous Next →