Akshat Bubna

speaker
405 appearances 1 recordings 1 series first heard Jul 2026 last heard 8 Jul

Akshat Bubna’s voice in public audio — every appearance, attributed to the second.

Trend

recordings per month · last 12 months
1 · Jul OctJan 26AprJulnow

Recordings per month over the last 12 months — 1 in all, peaking in Jul 2026 with 1.

Appearances

newest first · ▶ plays the moment
And they are taking the primitives we have and trying to use them in very interesting ways like continual learning.
It's possible as the stuff gets better, some of that will be part of our offering as well if more people need it.
But we're just waiting to see how it all shakes out.
I guess we've been going much deeper into LLM inference because we realized that some of the advantages we have with auto-scaling, again, especially in different regions and whatnot, are not present elsewhere.
And the place where we had a gap was we weren't working on the model there itself.
We were a black box.
And we realized that we actually...
can get to frontier level model performance by having great people who work on all of this.
And we've actually been open sourcing a lot for work in terms of recently we shared our work on D Flash, which is a block based speculator and we've open sourced all of it.
So you can get by using open source D Flash, you can get the same performance as you would with one of the proprietary providers.
Yeah, absolutely.
I mean, the high-level summaries, would it help to describe what speculative decoding is?
The speculative decoding is you have a smaller model called a draft model, predict tokens ahead of the bigger model, and then you have the bigger model verify all of this, all the tokens are predicted.
And the reason it's faster is if you're predicting one token at once, you're kind of bound by memory bandwidth.
But if you can batch the verification of the draft model, then you're much more efficient using compute and it's faster.
And as long as your draft model is producing a lot of tokens that can get accepted, which is called the accept length,
you can get speed up that's multiple times of, you know, the original model speed.
And that's what we highlight here.
It's like people talk a lot about we made these kernels faster and whatnot, but improving kernel only give you like a few percentage points of improvement.
And increasing except length literally is a multiplicative decrease.
Showing 101–120 of 405 · page 6 of 21 ← Previous Next →