Nufar Gaspar

speaker
2,488 appearances 11 recordings 1 series first heard Oct 2025 last heard 3 Sep

Nufar Gaspar’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
5 · Jun OctJan 26AprJulnow

Recordings per month over the last 12 months — 11 in all, peaking in Jun 2026 with 5.

Appearances

newest first · ▶ plays the moment
that can easily get to one million or more tokens per task.
And heavy coding can also be very aggressive, similarly with many agentic working flows.
So just to give you a sense of where it is, and to be a little bit more concrete here, the everyday stuff, as you've seen, like the emails and so on, is almost free.
So it's around half a cent, and nobody should ration emails.
It's not where their money goes.
Search and research can multiply very quietly.
So that can be a place to look for efficiency.
And the top of the ladder, that's a completely different sport.
So if you compare like email to agent decoding task, it can be a factor of a thousand or even more.
And another thing that you need to pay attention is that every conversation compounds.
So the model doesn't remember your previous messages.
And as such, it sends all of the previous conversations within the same session back to the model.
So by, let's say, turn number 10, it may be processing so much of the earlier exchange alongside your new message that the total goes much faster than the number of turns suggests.
That even happened before the system prompt, and we'll talk about strategies later on.
But this is one of the things that can very easily, just having very long sessions, can very easily amount to a ton of tokens being consumed.
So you want them to do deep research where the deep research is required, but you don't want them to accidentally do a deep research on a yes or no question that they can Google in a second.
Good.
So speaking of the agentic or the more advanced capabilities,
those can significantly grow the amount of tokens because agents work autonomously in loops and as such they consume by very widely cited industry estimates five to 30 times the tokens of a simple chat and poorly designed agentic loops or agentic harnesses can be even worse than that because a typical task involves between 10 to 20 model calls carrying instructions and history and tool definition and previous results.
And I think according to McKinsey, they estimate that roughly around 60% of an agentic task's cost is tied to the checking and refining and the regeneration of the answers after the first response.
Showing 701–720 of 2,488 · page 36 of 125 ← Previous Next →